Conventions this project follows. They mirror the algotrade component set so the two feel like siblings, adapted for a crawler/archiver.
- Stdlib first. Fetching and parsing use
urllib,xml.etree,re,json. A third-party dependency must earn its place. Current runtime deps:psycopg(Postgres),websockets(control surface),playwright(last-resort JS rendering only). - Headless browsing is a last resort. Anything obtainable with a plain GET
goes through
fetch.py.render.py(Playwright) is used only by adapters for journals whose article body or comments exist only after client-side JS. - Point-in-time honesty. Store the source's own timestamps (publication /
post time) separately from
fetched_at(poll time). Ids never depend on mutable data — only on identity. - Idempotent by construction. Journals/articles/comments have stable
hash-based ids; a new snapshot row is written only when
content_hashchanges. Re-polling is safe and cheap. - Graceful degradation. If Postgres is down the daemon keeps running and mirrors every record to the JSONL store for later replay.
- Secrets out of the repo. Only non-secret values live in
config.toml. Postgres credentials come fromsecret_postgre.envkept outside the tree (gitignored), read bydb.py. - Be a polite crawler. Per-host minimum delay, descriptive User-Agent, robots.txt respected by default. This is private research, not a scraper farm.
Articles (/story/…) are robots-allowed and fetched normally. The comment API
(api.<domain>/comment/v1/comments) is on a host whose robots.txt is
Disallow: /. The user explicitly opted to collect comments (2026-08-23), so
comment fetches — and only those — pass force_allow=True to bypass the robots
check; the per-host politeness delay still applies. Everything else stays fully
robots-compliant. If this posture changes, unset it in tamedia.fetch_comments.
proximity scores a writer as thirteen aggregate rates. lexicon scores the
character 4-grams they repeat. Held out by time over 744 Le Matin profiles, a
probe of 1300 characters finds its own author at the top of the list:
| reading | top-1 | top-5 | median rank |
|---|---|---|---|
| aggregate rates | 8% | 46% | 8 |
| rare words, idf-weighted | 40% | 60% | 4 |
| character 4-grams | 49% | 71% | 2 |
| turns of phrase (2-3 words) | 17% | 37% | 12 |
Both are reported, neither is blended into the other: no weighting this corpus
can justify exists, and an unjustified one would read as precision. Two
artefacts are removed inside lexicon rather than left to callers — @mentions
and URLs (the strongest match ever produced rested entirely on fragments of a
third party's handle both had replied to) and document size (uncapped, the
longest profile headed the ranking for five unrelated probes out of five).
Neither reading can say the true author is present at all: best score 0.175
with them in the population, 0.154 with them removed. Every ranking therefore
reports standout — how far its top sits above what the best of a field that
size is worth by chance, which is sqrt(2 ln n), near 3.6 for seven hundred.
Quote the excess, never the raw score.
Le Matin renders its comment list client-side, so www.lematin.ch/comment/<id>
returns a page with a comment count and no comments — not one nickname from a
150-comment thread appears in the 126 KB the server sends. Driving a real
browser was benchmarked against the plain-HTTP route on the same articles:
| requests | transferred | wall | comments | |
|---|---|---|---|---|
| HTTP (what we do) | 3 | 242 KiB | 3.2 s | 149 of 150 |
| Headless Chrome | 96 | 7,073 KiB | 7–10 s | 10 of 150 |
The browser makes the same API call we do — captured verbatim as
api.lematin.ch/comment/v1/comments?tenantId=4&contentId=…&limit=10 — with a
page size of 10 against our 100, then needs a click per further ten behind a
consent overlay. It is thirty times the traffic to obtain a fifteenth of the
data per request, and offers no field the API lacks, because the DOM it builds
is built from that JSON. The dependency is available on this machine and was
tested; it is not used.
- One lowercase package (
mediatracker/) inside the repo root, withserver.py/__main__.py/config.py/protocol.py/db.py/store.py, plus one module per external concern. sources/holds one adapter per journal;sources/__init__.pyis the registry. Adapters only parse into theParsed*shapes — the pipeline does all persistence, so every journal is handled uniformly.
defaults (config.py) → config.toml → MEDIATRACKER_* env → CLI flags.
Self-migrating at startup (db.ensure_schema, CREATE TABLE IF NOT EXISTS +
ADD COLUMN IF NOT EXISTS). Additive changes need no migration. Reserve numbered
migration scripts under a future migrations/ dir for breaking changes only.
Stdlib logging; log = logging.getLogger(__name__) per module;
basicConfig once in main(). WARNING/ERROR are the alerting channel and are
counted by health.py.
Scope is private, local-only research (confirmed with Cedric, 2026-08-23). Comment data — including pseudonyms and stylometric author-linkage — is retained locally for analysis and is never republished. Keep it that way: any feature that would export or expose individuals needs an explicit decision first.
Real names are uninteresting, not forbidden (Cedric, 2026-08-27). The database already holds journalists' names, and a commenter's name that turns up in a note or a handle is not a problem. The line is one of purpose: this is not a system for tracking real people down. So nothing guards the free-text fields, and nothing should — but no feature should set out to resolve a pseudonym to a person either, and the profiling pass still records what places a writer rather than who they are, because that is what the sociology needs.
article.origin now takes three values, and they are not interchangeable:
live— we asked the site, on a schedule, and it answered. Absence means the thread was not there.pdf— Cedric printed the page. Selected by his own participation: the threads are ones he commented on (confirmed 2026-08-29), so every other commenter in them is present for having written alongside him. This is not a thin sample of the paper; it is a complete sample of a different thing, and no share computed onpdfrows describes a year. Measured against the archive's unselected capture of the same year, the name-shaped share differs by fourteen points.wayback— a crawler happened to visit once. Absence means nothing at all, and presence is one moment: a thread caught at noon is missing its afternoon, and the same thread caught twice is two different lengths.
Every query that compares volumes across time has to say which origins it is
standing on, the way the community split already has to. The 2012-2016 material
arriving from the Internet Archive is wayback, sits beside pdf captures of
some of the same threads, and will double-count a comment only if the two
sources disagree on its id — which they do, because the printed archive has no
comment ids and uses a synthetic one. build_subjects already de-duplicates
identical text before measuring; that is the thing keeping this honest, and it
must keep being true.
Only the archive has comments. Every paid press database — Swissdox, SMD,
Factiva — indexes what the newsroom published, not what the public wrote
underneath. See SOURCES-BACKLOG.md.
Obvious in hindsight, and it produced a wrong answer anyway. Asked whether 2007 could be covered, I enumerated the archive for 2007 captures of the article grammar, found zero, and reported that there was nothing to fetch. Then a night of fetching 2008 captures returned 531 comments dated 2007 — July 2007 onward, on articles whose threads kept growing until a crawler finally visited in 2008.
The error was treating a capture as a window onto its own date. A capture is a window onto everything the page held up to that date, and on these platforms a thread is rendered whole. So the reachable past of a comment public extends behind the earliest capture, by however long the threads had been accumulating.
Two consequences for planning:
- A year with no captures is not a year with no comments. Before writing one off, check whether the next year's captures reach back into it — they are the same fetch, at no extra cost.
- The way to get more of an early year is to fetch more of the year after. There is no separate leg to run.
The measured shape for Le Matin: a trickle from July 2007 (4, 6, 10 a month), then 246 in November and 265 in December. The comment system opened in mid-2007 and found its public that autumn — four months before the URL grammar the archive can be searched by even existed.
The recoverable past is not one era but two, and they differ in what they can be asked:
- PHP "reactions", from somewhere in 2008 to about 2009. Thread inline on
the article page, unpaginated, each comment bracketed in
<!-- BEGIN/END COMMENT HTML -->. Every comment names a numericidUser. This is the sturdiest identity in the corpus: a number does not move when the display name does, so it can confirm a rename outright. The URL carries the section before the id —/fr/actu/economie/<slug>_11-271110— and the grammar appears part-way through 2008, not in January. Pages arecharset=iso-8859-1. - Drupal, roughly 2009 to March 2012. The article page carries the whole
thread inline; the URL is a slug with the id glued on the end. Every comment
names
/users/<key>, a real account identifier, with the display form beside it —MountaiDiverposting from/users/mountaidiver. The key is derived from the display name, so a rename moves it: weaker than a number, far better than nothing. No reply threading, no like counts. - Newsnetz, March 2012 to about 2017. The thread lives at
?comments=1. Reply threading survives in an HTML comment, and likes and dislikes are counted. No user id of any kind.
The direction is one way: a numeric id, then a name-derived slug, then nothing. The further back the material, the more it can settle.
And the layers are Le Matin's alone. Asked for the same two grammars, 24 heures and the Tribune return single digits or zero for every year — measured across all thirteen legs on 2026-08-31. Their recoverable history begins with Newsnetz in 2012. So the corpus is not three parallel histories of different depth; it is one deep history and two shallow ones, and a measure that needs pre-2012 material from all three titles cannot be built at any amount of fetching.
So identity is a different question on either side of March 2012. Before it,
two comments can be tied to one account outright. After it, a nickname is all
there is, and every claim that two handles are one writer is an inference from
style — which is what alias_candidates and the proximity view exist to make,
and why they return a shortlist and never a verdict. A method calibrated on the
Newsnetz years cannot be validated on itself; the Drupal years are the only
stretch of this corpus that carries ground truth about succession, and that
makes them worth more per page than their comment counts suggest.
The corollary is a trap. author_key is present for one era and absent for the
other, so any count grouped by it silently becomes a count of 2009-2012. Group
by nickname unless the question is specifically about accounts, and say which
era the answer stands on.
Which reader parses a page is decided by the markup, never by the capture year:
the changeover was a deployment, and captures from the same week fall either
side of it. Order the tests carefully — class="commentaire is a prefix of the
PHP era's class="commentaires", so the loose Drupal test claims every 2008
page, and the wrong reader returns an empty thread rather than an error.
Encoding is part of the era, not a detail. The socket timeout measures
inactivity and errors="replace" does not raise: both are silent failures that
produce plausible-looking output. A listing that dribbles forever and a page
whose every accent became \ufffd each cost a day before being noticed. Where a
failure cannot announce itself, put a clock or a declaration in its way.
Notes and off-platform accounts are their own tables (subject_note,
subject_account), not columns on author_profile. Two reasons, and both
matter:
- A profiling run rewrites every column of an
author_profilerow. A remark typed by a reader would be erased by the next run of the machine. - Everything in
author_profileis derived from the stored comments and reproducible from them. A note is not, and mixing the two would put an observation and a measurement in the same shape — the same mistake themetrics/ inferred split exists to prevent.
Both are keyed (community, subject_kind, subject_key) like a profile. A
persona reads its own records and those written against each of its
handles, never the reverse: linking two nicknames must not bury what was
already recorded about either, while a note about the person says more than any
one handle can carry.
Off-platform accounts are structured rather than prose because the questions worth asking about them are counting ones — how many of a public carry an identity elsewhere, on which network, and whether a rename there lines up with a rename here. The platform is read off the pasted link rather than asked for, so the two cannot disagree.
pytest, tests in tests/ as test_*.py. Run with python3 -m pytest -q from
the repo root. Unit tests must not require Postgres or network.