• It's recommended to check out the latest log via: the Stats page >> View log >> Tail

PROJECT (scraping_t), SPIDER (auditor_bot)

2026-10-10 18:57:57 [scrapy.utils.log] INFO: Scrapy 2.19.0 started (bot: audit_engine)
2026-10-10 18:57:57 [scrapy.utils.log] INFO: Versions:
{'lxml': '6.1.3',
 'libxml2': '2.14.6',
 'cssselect': '1.5.0',
 'parsel': '1.12.1',
 'w3lib': '2.5.0',
 'Twisted': '26.4.0',
 'Python': '3.13.5 (main, Aug 10 2026, 12:06:59) [GCC 14.2.0]',
 'pyOpenSSL': '26.4.0 (OpenSSL 4.0.3 29 Sep 2026)',
 'cryptography': '50.0.2',
 'Platform': 'Linux-6.8.0-138-generic-x86_64-with-glibc2.41'}
2026-10-10 18:57:57 [scrapy.crawler] DEBUG: Using AsyncCrawlerProcess
2026-10-10 18:57:57 [asyncio] DEBUG: Using selector: EpollSelector
2026-10-10 18:57:57 [scrapy.addons] INFO: Enabled addons:
[]
2026-10-10 18:57:57 [scrapy.utils.log] DEBUG: Using reactor: twisted.internet.asyncioreactor.AsyncioSelectorReactor
2026-10-10 18:57:57 [scrapy.utils.log] DEBUG: Using asyncio event loop: asyncio.unix_events._UnixSelectorEventLoop
2026-10-10 18:57:57 [scrapy.extensions.telnet] INFO: Telnet Password: 12a5f47870143918
2026-10-10 18:57:57 [scrapy.middleware] INFO: Enabled extensions:
['scrapy.extensions.corestats.CoreStats',
 'scrapy.extensions.logcount.LogCount',
 'scrapy.extensions.telnet.TelnetConsole',
 'scrapy.extensions.memusage.MemoryUsage',
 'scrapy.extensions.feedexport.FeedExporter',
 'scrapy.extensions.logstats.LogStats',
 'scrapy.extensions.throttle.AutoThrottle',
 'scrapy.extensions.remote_control.RemoteControl']
2026-10-10 18:57:57 [scrapy.crawler] INFO: Overridden settings:
{'AUTOTHROTTLE_ENABLED': True,
 'AUTOTHROTTLE_MAX_DELAY': 10.0,
 'AUTOTHROTTLE_START_DELAY': 1.0,
 'BOT_NAME': 'audit_engine',
 'CONCURRENT_REQUESTS_PER_DOMAIN': 2,
 'COOKIES_ENABLED': False,
 'DOWNLOAD_DELAY': 1.0,
 'LOG_FILE': '/var/lib/scrapyd/logs/scraping_t/auditor_bot/2026-10-10T18_57_53.log',
 'NEWSPIDER_MODULE': 'audit_engine.spiders',
 'SPIDER_MODULES': ['audit_engine.spiders']}
2026-10-10 18:57:57 [scrapy.middleware] INFO: Enabled downloader middlewares:
['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware',
 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware',
 'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware',
 'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware',
 'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware',
 'scrapy.downloadermiddlewares.retry.RetryMiddleware',
 'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware',
 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware',
 'scrapy.downloadermiddlewares.redirect.RedirectMiddleware',
 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware',
 'scrapy.downloadermiddlewares.stats.DownloaderStats']
2026-10-10 18:57:57 [scrapy.middleware] INFO: Enabled spider middlewares:
['scrapy.spidermiddlewares.start.StartSpiderMiddleware',
 'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware',
 'scrapy.spidermiddlewares.referer.RefererMiddleware',
 'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware',
 'scrapy.spidermiddlewares.depth.DepthMiddleware',
 'scrapy.spidermiddlewares.metacopy.MetaCopyDetectionMiddleware']
2026-10-10 18:57:57 [scrapy.middleware] INFO: Enabled item pipelines:
['audit_engine.pipelines.AuditMariaDBPipeline']
2026-10-10 18:57:57 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:60: ScrapyDeprecationWarning: AuditMariaDBPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
  self._check_mw_method_spider_arg(mw.open_spider)

2026-10-10 18:57:57 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:63: ScrapyDeprecationWarning: AuditMariaDBPipeline.close_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
  self._check_mw_method_spider_arg(mw.close_spider)

2026-10-10 18:57:57 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:66: ScrapyDeprecationWarning: AuditMariaDBPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
  self._check_mw_method_spider_arg(mw.process_item)

2026-10-10 18:57:57 [scrapy.core.engine] INFO: Spider opened
2026-10-10 18:58:07 [scrapy.core.engine] INFO: Closing spider (shutdown)
2026-10-10 18:58:07 [scrapy.utils.signal] ERROR: Error caught on signal handler: <bound method CoreStats.spider_closed of <scrapy.extensions.corestats.CoreStats object at 0x77e29cfe5940>>
Traceback (most recent call last):
  File "/usr/local/lib/python3.13/dist-packages/scrapy/utils/signal.py", line 192, in handler
    robustApply(
    ~~~~~~~~~~~^
        receiver, *arguments, signal=signal, sender=sender, **named
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    ),
    ^
  File "/usr/local/lib/python3.13/dist-packages/pydispatch/robustapply.py", line 55, in robustApply
    return receiver(*arguments, **named)
  File "/usr/local/lib/python3.13/dist-packages/scrapy/extensions/corestats.py", line 43, in spider_closed
    assert self.start_time is not None
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionError
2026-10-10 18:58:08 [scrapy.statscollectors] INFO: Dumping Scrapy stats:
{'items_per_minute': None, 'responses_per_minute': None}
2026-10-10 18:58:08 [scrapy.core.engine] INFO: Spider closed (shutdown)
2026-10-10 18:58:08 [asyncio] ERROR: Task exception was never retrieved
future: <Task finished name='Task-1' coro=<AsyncCrawlerRunner._crawl_and_track() done, defined at /usr/local/lib/python3.13/dist-packages/scrapy/crawler.py:671> exception=OperationalError(2003, "Can't connect to MySQL server on 'host.docker.internal' (timed out)")>
Traceback (most recent call last):
  File "/usr/local/lib/python3.13/dist-packages/pymysql/connections.py", line 681, in connect
    sock = socket.create_connection(
        (self.host, self.port), self.connect_timeout, **kwargs
    )
  File "/usr/lib/python3.13/socket.py", line 864, in create_connection
    raise exceptions[0]
  File "/usr/lib/python3.13/socket.py", line 849, in create_connection
    sock.connect(sa)
    ~~~~~~~~~~~~^^^^
TimeoutError: timed out

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/usr/local/lib/python3.13/dist-packages/scrapy/crawler.py", line 675, in _crawl_and_track
    await crawler.crawl_async(*args, **kwargs)
  File "/usr/local/lib/python3.13/dist-packages/scrapy/crawler.py", line 315, in crawl_async
    await self.engine.open_spider_async()
  File "/usr/local/lib/python3.13/dist-packages/scrapy/core/engine.py", line 569, in open_spider_async
    await self.scraper.open_spider_async()
  File "/usr/local/lib/python3.13/dist-packages/scrapy/core/scraper.py", line 177, in open_spider_async
    await self.itemproc.open_spider_async()
  File "/usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py", line 147, in open_spider_async
    await self._process_parallel("open_spider")
  File "/usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py", line 134, in _process_parallel
    return await self._process_parallel_asyncio(methodname)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py", line 128, in _process_parallel_asyncio
    awaitables = [self.get_awaitable(m) for m in methods]
                  ~~~~~~~~~~~~~~~~~~^^^
  File "/usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py", line 115, in get_awaitable
    result = method(self._spider)
  File "/var/lib/scrapyd/eggs/scraping_t/2026-10-10T14_56_21.egg/audit_engine/pipelines.py", line 9, in open_spider
    self.conn = MySQLdb.connect(
                ~~~~~~~~~~~~~~~^
        host="host.docker.internal", # Routes cleanly out of the container to your host machine
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    ...<4 lines>...
        charset="utf8mb4"
        ^^^^^^^^^^^^^^^^^
    )
    ^
  File "/usr/local/lib/python3.13/dist-packages/pymysql/connections.py", line 373, in __init__
    self.connect()
    ~~~~~~~~~~~~^^
  File "/usr/local/lib/python3.13/dist-packages/pymysql/connections.py", line 744, in connect
    raise exc
pymysql.err.OperationalError: (2003, "Can't connect to MySQL server on 'host.docker.internal' (timed out)")

PROJECT (scraping_t), SPIDER (auditor_bot)