Problem
Scrapy’s S3DownloadHandler sends signed S3 requests over plaintext HTTP by default.
A normal request like s3://bucket/key is converted into http://bucket.s3.amazonaws.com/key unless request.meta["is_secure"] is explicitly set. The generated request is then signed with configured AWS credentials, so AWS authorization material can be sent without TLS.
Vulnerable code in scrapy/core/downloader/handlers/s3.py:
scheme = "https" if request.meta.get("is_secure") else "http"
url = f"{scheme}://{bucket}.s3.amazonaws.com{path}"
The request is then signed and dispatched:
self._signer.add_auth(awsrequest)
request = request.replace(url=url, headers=awsrequest.headers.items())
Impact
Users making Scrapy s3:// requests with AWS credentials are impacted.
A network attacker able to observe traffic between Scrapy and S3, such as a public Wi-Fi attacker, compromised router, ISP/corporate network observer, or local network attacker using ARP spoofing, can read:
bucket/key path
AWS Authorization header
X-Amz-Security-Token, if temporary credentials are used
S3 object contents
S3 response headers
An active MITM attacker can also modify the plaintext S3 response body, status code, and headers before Scrapy processes it. This can cause scraped data poisoning, poisoned exports, HTTP cache poisoning when cache is enabled, and influence over later crawl targets through forged redirects or attacker-controlled links.
Suggested classification: CWE-319: Cleartext Transmission of Sensitive Information.
PoC
A minimal PoC creates an s3:// request, enables fake AWS settings, captures the request produced by S3DownloadHandler, and prints the rewritten URL and auth headers.
References
Problem
Scrapy’s
S3DownloadHandlersends signed S3 requests over plaintext HTTP by default.A normal request like
s3://bucket/keyis converted intohttp://bucket.s3.amazonaws.com/keyunlessrequest.meta["is_secure"]is explicitly set. The generated request is then signed with configured AWS credentials, so AWS authorization material can be sent without TLS.Vulnerable code in
scrapy/core/downloader/handlers/s3.py:The request is then signed and dispatched:
Impact
Users making Scrapy
s3://requests with AWS credentials are impacted.A network attacker able to observe traffic between Scrapy and S3, such as a public Wi-Fi attacker, compromised router, ISP/corporate network observer, or local network attacker using ARP spoofing, can read:
An active MITM attacker can also modify the plaintext S3 response body, status code, and headers before Scrapy processes it. This can cause scraped data poisoning, poisoned exports, HTTP cache poisoning when cache is enabled, and influence over later crawl targets through forged redirects or attacker-controlled links.
Suggested classification: CWE-319: Cleartext Transmission of Sensitive Information.
PoC
A minimal PoC creates an
s3://request, enables fake AWS settings, captures the request produced byS3DownloadHandler, and prints the rewritten URL and auth headers.References