• 'pip install logparser' on host 'scrapyd:6800' and run command 'logparser'. Or wait until LogParser parses the log.

PROJECT (AuditCrawl), SPIDER (auditor_bot)

  • Log analysis
  • Log categorization
  • View log
  • Crawler.stats
  • projectAuditCrawl
    spiderauditor_bot
    job2026-10-10T20_04_31
    first_log_time2026-10-10 20:04:37
    latest_log_time2026-10-10 20:04:39
    runtime0:00:02
    crawled_pages 1
    scraped_items 0
    shutdown_reasonN/A
    finish_reasonfinished
    log_critical_count0
    log_error_count0
    log_warning_count3
    log_redirect_count0
    log_retry_count0
    log_ignore_count0
    latest_crawl
    latest_scrape
    latest_log
    current_time
    latest_itemN/A
    • WARNING+

    • warning_logs
      3 in total

      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:60: ScrapyDeprecationWarning: AuditMariaDBPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.open_spider)
      
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:63: ScrapyDeprecationWarning: AuditMariaDBPipeline.close_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.close_spider)
      
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:66: ScrapyDeprecationWarning: AuditMariaDBPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.process_item)
      

      INFO

      DEBUG

    • scrapy_version

      2.19.0
    • telnet_console

      127.0.0.1:6023
    • telnet_password

      8c29a2a0dd3323c3
    • latest_crawl

      2026-10-10 20:04:39 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://toscrape.com> (referer: None)
    • latest_stat

      2026-10-10 20:04:38 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
    • Head

      2026-10-10 20:04:37 [scrapy.utils.log] INFO: Scrapy 2.19.0 started (bot: audit_engine)
      2026-10-10 20:04:37 [scrapy.utils.log] INFO: Versions:
      {'lxml': '6.1.3',
       'libxml2': '2.14.6',
       'cssselect': '1.5.0',
       'parsel': '1.12.1',
       'w3lib': '2.5.0',
       'Twisted': '26.4.0',
       'Python': '3.13.5 (main, Aug 10 2026, 12:06:59) [GCC 14.2.0]',
       'pyOpenSSL': '26.4.0 (OpenSSL 4.0.3 29 Sep 2026)',
       'cryptography': '50.0.2',
       'Platform': 'Linux-6.8.0-138-generic-x86_64-with-glibc2.41'}
      2026-10-10 20:04:37 [scrapy.crawler] DEBUG: Using AsyncCrawlerProcess
      2026-10-10 20:04:37 [asyncio] DEBUG: Using selector: EpollSelector
      2026-10-10 20:04:37 [scrapy.addons] INFO: Enabled addons:
      []
      2026-10-10 20:04:37 [scrapy.utils.log] DEBUG: Using reactor: twisted.internet.asyncioreactor.AsyncioSelectorReactor
      2026-10-10 20:04:37 [scrapy.utils.log] DEBUG: Using asyncio event loop: asyncio.unix_events._UnixSelectorEventLoop
      2026-10-10 20:04:37 [scrapy.extensions.telnet] INFO: Telnet Password: 8c29a2a0dd3323c3
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled extensions:
      ['scrapy.extensions.corestats.CoreStats',
       'scrapy.extensions.logcount.LogCount',
       'scrapy.extensions.telnet.TelnetConsole',
       'scrapy.extensions.memusage.MemoryUsage',
       'scrapy.extensions.feedexport.FeedExporter',
       'scrapy.extensions.logstats.LogStats',
       'scrapy.extensions.throttle.AutoThrottle',
       'scrapy.extensions.remote_control.RemoteControl']
      2026-10-10 20:04:38 [scrapy.crawler] INFO: Overridden settings:
      {'AUTOTHROTTLE_ENABLED': True,
       'AUTOTHROTTLE_MAX_DELAY': 10.0,
       'AUTOTHROTTLE_START_DELAY': 1.0,
       'BOT_NAME': 'audit_engine',
       'CONCURRENT_REQUESTS_PER_DOMAIN': 2,
       'COOKIES_ENABLED': False,
       'DOWNLOAD_DELAY': 1.0,
       'LOG_FILE': '/var/lib/scrapyd/logs/AuditCrawl/auditor_bot/2026-10-10T20_04_31.log',
       'NEWSPIDER_MODULE': 'audit_engine.spiders',
       'SPIDER_MODULES': ['audit_engine.spiders']}
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled downloader middlewares:
      ['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware',
       'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware',
       'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware',
       'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware',
       'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware',
       'scrapy.downloadermiddlewares.retry.RetryMiddleware',
       'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware',
       'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware',
       'scrapy.downloadermiddlewares.redirect.RedirectMiddleware',
       'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware',
       'scrapy.downloadermiddlewares.stats.DownloaderStats']
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled spider middlewares:
      ['scrapy.spidermiddlewares.start.StartSpiderMiddleware',
       'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware',
       'scrapy.spidermiddlewares.referer.RefererMiddleware',
       'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware',
       'scrapy.spidermiddlewares.depth.DepthMiddleware',
       'scrapy.spidermiddlewares.metacopy.MetaCopyDetectionMiddleware']
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled item pipelines:
      ['audit_engine.pipelines.AuditMariaDBPipeline']
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:60: ScrapyDeprecationWarning: AuditMariaDBPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.open_spider)
      
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:63: ScrapyDeprecationWarning: AuditMariaDBPipeline.close_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.close_spider)
      
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:66: ScrapyDeprecationWarning: AuditMariaDBPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.process_item)
      
      2026-10-10 20:04:38 [scrapy.core.engine] INFO: Spider opened
      2026-10-10 20:04:38 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
      2026-10-10 20:04:38 [scrapy.extensions.telnet] INFO: Telnet console listening on 127.0.0.1:6023
      2026-10-10 20:04:38 [scrapy.extensions.remote_control] INFO: Remote control HTTP server listening on port 40365 (job 44-5d4aa5bba4b34fa4afd04e5e010b9fc8)
      2026-10-10 20:04:39 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://toscrape.com> (referer: None)
      2026-10-10 20:04:39 [scrapy.core.engine] INFO: Closing spider (finished)
      2026-10-10 20:04:39 [scrapy.extensions.feedexport] INFO: Stored jsonlines feed (0 items) in: file:///var/lib/scrapyd/items/AuditCrawl/auditor_bot/2026-10-10T20_04_31.jl
      2026-10-10 20:04:39 [scrapy.statscollectors] INFO: Dumping Scrapy stats:
      {'downloader/request_bytes': 223,
       'downloader/request_count': 1,
       'downloader/request_method_count/GET': 1,
       'downloader/response_bytes': 1213,
       'downloader/response_count': 1,
       'downloader/response_status_count/200': 1,
       'elapsed_time_seconds': 0.2642331342212856,
       'feedexport/success_count/FileFeedStorage': 1,
       'finish_reason': 'finished',
       'finish_time': datetime.datetime(2026, 10, 10, 20, 4, 39, 46408, tzinfo=datetime.timezone.utc),
       'httpcompression/response_bytes': 3939,
       'httpcompression/response_count': 1,
       'items_per_minute': None,
       'log_count/DEBUG': 1,
       'log_count/INFO': 4,
       'memusage/max': 88129536,
       'memusage/startup': 88129536,
       'response_received_count': 1,
       'responses_per_minute': None,
       'scheduler/dequeued': 1,
       'scheduler/dequeued/memory': 1,
       'scheduler/enqueued': 1,
       'scheduler/enqueued/memory': 1,
    • Tail

      2026-10-10 20:04:37 [scrapy.utils.log] INFO: Scrapy 2.19.0 started (bot: audit_engine)
      2026-10-10 20:04:37 [scrapy.utils.log] INFO: Versions:
      {'lxml': '6.1.3',
       'libxml2': '2.14.6',
       'cssselect': '1.5.0',
       'parsel': '1.12.1',
       'w3lib': '2.5.0',
       'Twisted': '26.4.0',
       'Python': '3.13.5 (main, Aug 10 2026, 12:06:59) [GCC 14.2.0]',
       'pyOpenSSL': '26.4.0 (OpenSSL 4.0.3 29 Sep 2026)',
       'cryptography': '50.0.2',
       'Platform': 'Linux-6.8.0-138-generic-x86_64-with-glibc2.41'}
      2026-10-10 20:04:37 [scrapy.crawler] DEBUG: Using AsyncCrawlerProcess
      2026-10-10 20:04:37 [asyncio] DEBUG: Using selector: EpollSelector
      2026-10-10 20:04:37 [scrapy.addons] INFO: Enabled addons:
      []
      2026-10-10 20:04:37 [scrapy.utils.log] DEBUG: Using reactor: twisted.internet.asyncioreactor.AsyncioSelectorReactor
      2026-10-10 20:04:37 [scrapy.utils.log] DEBUG: Using asyncio event loop: asyncio.unix_events._UnixSelectorEventLoop
      2026-10-10 20:04:37 [scrapy.extensions.telnet] INFO: Telnet Password: 8c29a2a0dd3323c3
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled extensions:
      ['scrapy.extensions.corestats.CoreStats',
       'scrapy.extensions.logcount.LogCount',
       'scrapy.extensions.telnet.TelnetConsole',
       'scrapy.extensions.memusage.MemoryUsage',
       'scrapy.extensions.feedexport.FeedExporter',
       'scrapy.extensions.logstats.LogStats',
       'scrapy.extensions.throttle.AutoThrottle',
       'scrapy.extensions.remote_control.RemoteControl']
      2026-10-10 20:04:38 [scrapy.crawler] INFO: Overridden settings:
      {'AUTOTHROTTLE_ENABLED': True,
       'AUTOTHROTTLE_MAX_DELAY': 10.0,
       'AUTOTHROTTLE_START_DELAY': 1.0,
       'BOT_NAME': 'audit_engine',
       'CONCURRENT_REQUESTS_PER_DOMAIN': 2,
       'COOKIES_ENABLED': False,
       'DOWNLOAD_DELAY': 1.0,
       'LOG_FILE': '/var/lib/scrapyd/logs/AuditCrawl/auditor_bot/2026-10-10T20_04_31.log',
       'NEWSPIDER_MODULE': 'audit_engine.spiders',
       'SPIDER_MODULES': ['audit_engine.spiders']}
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled downloader middlewares:
      ['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware',
       'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware',
       'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware',
       'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware',
       'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware',
       'scrapy.downloadermiddlewares.retry.RetryMiddleware',
       'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware',
       'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware',
       'scrapy.downloadermiddlewares.redirect.RedirectMiddleware',
       'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware',
       'scrapy.downloadermiddlewares.stats.DownloaderStats']
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled spider middlewares:
      ['scrapy.spidermiddlewares.start.StartSpiderMiddleware',
       'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware',
       'scrapy.spidermiddlewares.referer.RefererMiddleware',
       'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware',
       'scrapy.spidermiddlewares.depth.DepthMiddleware',
       'scrapy.spidermiddlewares.metacopy.MetaCopyDetectionMiddleware']
      2026-10-10 20:04:38 [scrapy.middleware] INFO: Enabled item pipelines:
      ['audit_engine.pipelines.AuditMariaDBPipeline']
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:60: ScrapyDeprecationWarning: AuditMariaDBPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.open_spider)
      
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:63: ScrapyDeprecationWarning: AuditMariaDBPipeline.close_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.close_spider)
      
      2026-10-10 20:04:38 [py.warnings] WARNING: /usr/local/lib/python3.13/dist-packages/scrapy/pipelines/__init__.py:66: ScrapyDeprecationWarning: AuditMariaDBPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute.
        self._check_mw_method_spider_arg(mw.process_item)
      
      2026-10-10 20:04:38 [scrapy.core.engine] INFO: Spider opened
      2026-10-10 20:04:38 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
      2026-10-10 20:04:38 [scrapy.extensions.telnet] INFO: Telnet console listening on 127.0.0.1:6023
      2026-10-10 20:04:38 [scrapy.extensions.remote_control] INFO: Remote control HTTP server listening on port 40365 (job 44-5d4aa5bba4b34fa4afd04e5e010b9fc8)
      2026-10-10 20:04:39 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://toscrape.com> (referer: None)
      2026-10-10 20:04:39 [scrapy.core.engine] INFO: Closing spider (finished)
      2026-10-10 20:04:39 [scrapy.extensions.feedexport] INFO: Stored jsonlines feed (0 items) in: file:///var/lib/scrapyd/items/AuditCrawl/auditor_bot/2026-10-10T20_04_31.jl
      2026-10-10 20:04:39 [scrapy.statscollectors] INFO: Dumping Scrapy stats:
      {'downloader/request_bytes': 223,
       'downloader/request_count': 1,
       'downloader/request_method_count/GET': 1,
       'downloader/response_bytes': 1213,
       'downloader/response_count': 1,
       'downloader/response_status_count/200': 1,
       'elapsed_time_seconds': 0.2642331342212856,
       'feedexport/success_count/FileFeedStorage': 1,
       'finish_reason': 'finished',
       'finish_time': datetime.datetime(2026, 10, 10, 20, 4, 39, 46408, tzinfo=datetime.timezone.utc),
       'httpcompression/response_bytes': 3939,
       'httpcompression/response_count': 1,
       'items_per_minute': None,
       'log_count/DEBUG': 1,
       'log_count/INFO': 4,
       'memusage/max': 88129536,
       'memusage/startup': 88129536,
       'response_received_count': 1,
       'responses_per_minute': None,
       'scheduler/dequeued': 1,
       'scheduler/dequeued/memory': 1,
       'scheduler/enqueued': 1,
       'scheduler/enqueued/memory': 1,
       'start_time': datetime.datetime(2026, 10, 10, 20, 4, 38, 782171, tzinfo=datetime.timezone.utc)}
      2026-10-10 20:04:39 [scrapy.core.engine] INFO: Spider closed (finished)
    • Log

      /1/log/utf8/AuditCrawl/auditor_bot/2026-10-10T20_04_31/?job_finished=True

    • Source

      http://scrapyd:6800/logs/AuditCrawl/auditor_bot/2026-10-10T20_04_31.log

  • sourcelog
    last_update_time2026-10-10 20:04:39
    last_update_timestamp1791662679
    downloader/request_bytes223
    downloader/request_count1
    downloader/request_method_count/GET1
    downloader/response_bytes1213
    downloader/response_count1
    downloader/response_status_count/2001
    elapsed_time_seconds0.2642331342212856
    feedexport/success_count/FileFeedStorage1
    finish_reasonfinished
    finish_timedatetime.datetime(2026, 10, 10, 20, 4, 39, 46408, tzinfo=datetime.timezone.utc)
    httpcompression/response_bytes3939
    httpcompression/response_count1
    items_per_minuteNone
    log_count/DEBUG1
    log_count/INFO4
    memusage/max88129536
    memusage/startup88129536
    response_received_count1
    responses_per_minuteNone
    scheduler/dequeued1
    scheduler/dequeued/memory1
    scheduler/enqueued1
    scheduler/enqueued/memory1
    start_timedatetime.datetime(2026, 10, 10, 20, 4, 38, 782171, tzinfo=datetime.timezone.utc)