Skip to content

Latest commit

 

History

History
268 lines (216 loc) · 14.2 KB

File metadata and controls

268 lines (216 loc) · 14.2 KB

MediaTracker doctrine

Conventions this project follows. They mirror the algotrade component set so the two feel like siblings, adapted for a crawler/archiver.

Principles

  1. Stdlib first. Fetching and parsing use urllib, xml.etree, re, json. A third-party dependency must earn its place. Current runtime deps: psycopg (Postgres), websockets (control surface), playwright (last-resort JS rendering only).
  2. Headless browsing is a last resort. Anything obtainable with a plain GET goes through fetch.py. render.py (Playwright) is used only by adapters for journals whose article body or comments exist only after client-side JS.
  3. Point-in-time honesty. Store the source's own timestamps (publication / post time) separately from fetched_at (poll time). Ids never depend on mutable data — only on identity.
  4. Idempotent by construction. Journals/articles/comments have stable hash-based ids; a new snapshot row is written only when content_hash changes. Re-polling is safe and cheap.
  5. Graceful degradation. If Postgres is down the daemon keeps running and mirrors every record to the JSONL store for later replay.
  6. Secrets out of the repo. Only non-secret values live in config.toml. Postgres credentials come from secret_postgre.env kept outside the tree (gitignored), read by db.py.
  7. Be a polite crawler. Per-host minimum delay, descriptive User-Agent, robots.txt respected by default. This is private research, not a scraper farm.

Comment collection & robots.txt

Articles (/story/…) are robots-allowed and fetched normally. The comment API (api.<domain>/comment/v1/comments) is on a host whose robots.txt is Disallow: /. The user explicitly opted to collect comments (2026-08-23), so comment fetches — and only those — pass force_allow=True to bypass the robots check; the per-host politeness delay still applies. Everything else stays fully robots-compliant. If this posture changes, unset it in tamedia.fetch_comments.

Two readings of a writer, never blended (measured 2026-08-27)

proximity scores a writer as thirteen aggregate rates. lexicon scores the character 4-grams they repeat. Held out by time over 744 Le Matin profiles, a probe of 1300 characters finds its own author at the top of the list:

reading top-1 top-5 median rank
aggregate rates 8% 46% 8
rare words, idf-weighted 40% 60% 4
character 4-grams 49% 71% 2
turns of phrase (2-3 words) 17% 37% 12

Both are reported, neither is blended into the other: no weighting this corpus can justify exists, and an unjustified one would read as precision. Two artefacts are removed inside lexicon rather than left to callers — @mentions and URLs (the strongest match ever produced rested entirely on fragments of a third party's handle both had replied to) and document size (uncapped, the longest profile headed the ranking for five unrelated probes out of five).

Neither reading can say the true author is present at all: best score 0.175 with them in the population, 0.154 with them removed. Every ranking therefore reports standout — how far its top sits above what the best of a field that size is worth by chance, which is sqrt(2 ln n), near 3.6 for seven hundred. Quote the excess, never the raw score.

No headless browser (measured 2026-08-27)

Le Matin renders its comment list client-side, so www.lematin.ch/comment/<id> returns a page with a comment count and no comments — not one nickname from a 150-comment thread appears in the 126 KB the server sends. Driving a real browser was benchmarked against the plain-HTTP route on the same articles:

requests transferred wall comments
HTTP (what we do) 3 242 KiB 3.2 s 149 of 150
Headless Chrome 96 7,073 KiB 7–10 s 10 of 150

The browser makes the same API call we do — captured verbatim as api.lematin.ch/comment/v1/comments?tenantId=4&contentId=…&limit=10 — with a page size of 10 against our 100, then needs a click per further ten behind a consent overlay. It is thirty times the traffic to obtain a fifteenth of the data per request, and offers no field the API lacks, because the DOM it builds is built from that JSON. The dependency is available on this machine and was tested; it is not used.

Layout

  • One lowercase package (mediatracker/) inside the repo root, with server.py / __main__.py / config.py / protocol.py / db.py / store.py, plus one module per external concern.
  • sources/ holds one adapter per journal; sources/__init__.py is the registry. Adapters only parse into the Parsed* shapes — the pipeline does all persistence, so every journal is handled uniformly.

Config layering

defaults (config.py) → config.toml → MEDIATRACKER_* env → CLI flags.

Schema evolution

Self-migrating at startup (db.ensure_schema, CREATE TABLE IF NOT EXISTS + ADD COLUMN IF NOT EXISTS). Additive changes need no migration. Reserve numbered migration scripts under a future migrations/ dir for breaking changes only.

Logging

Stdlib logging; log = logging.getLogger(__name__) per module; basicConfig once in main(). WARNING/ERROR are the alerting channel and are counted by health.py.

Privacy posture

Scope is private, local-only research (confirmed with Cedric, 2026-08-23). Comment data — including pseudonyms and stylometric author-linkage — is retained locally for analysis and is never republished. Keep it that way: any feature that would export or expose individuals needs an explicit decision first.

Real names are uninteresting, not forbidden (Cedric, 2026-08-27). The database already holds journalists' names, and a commenter's name that turns up in a note or a handle is not a problem. The line is one of purpose: this is not a system for tracking real people down. So nothing guards the free-text fields, and nothing should — but no feature should set out to resolve a pseudonym to a person either, and the profiling pass still records what places a writer rather than who they are, because that is what the sociology needs.

Archive capture is a third kind of evidence (2026-08-28)

article.origin now takes three values, and they are not interchangeable:

  • live — we asked the site, on a schedule, and it answered. Absence means the thread was not there.
  • pdf — Cedric printed the page. Selected by his own participation: the threads are ones he commented on (confirmed 2026-08-29), so every other commenter in them is present for having written alongside him. This is not a thin sample of the paper; it is a complete sample of a different thing, and no share computed on pdf rows describes a year. Measured against the archive's unselected capture of the same year, the name-shaped share differs by fourteen points.
  • wayback — a crawler happened to visit once. Absence means nothing at all, and presence is one moment: a thread caught at noon is missing its afternoon, and the same thread caught twice is two different lengths.

Every query that compares volumes across time has to say which origins it is standing on, the way the community split already has to. The 2012-2016 material arriving from the Internet Archive is wayback, sits beside pdf captures of some of the same threads, and will double-count a comment only if the two sources disagree on its id — which they do, because the printed archive has no comment ids and uses a synthetic one. build_subjects already de-duplicates identical text before measuring; that is the thing keeping this honest, and it must keep being true.

Only the archive has comments. Every paid press database — Swissdox, SMD, Factiva — indexes what the newsroom published, not what the public wrote underneath. See SOURCES-BACKLOG.md.

The year a capture was taken is not the year its comments were written (2026-08-31)

Obvious in hindsight, and it produced a wrong answer anyway. Asked whether 2007 could be covered, I enumerated the archive for 2007 captures of the article grammar, found zero, and reported that there was nothing to fetch. Then a night of fetching 2008 captures returned 531 comments dated 2007 — July 2007 onward, on articles whose threads kept growing until a crawler finally visited in 2008.

The error was treating a capture as a window onto its own date. A capture is a window onto everything the page held up to that date, and on these platforms a thread is rendered whole. So the reachable past of a comment public extends behind the earliest capture, by however long the threads had been accumulating.

Two consequences for planning:

  • A year with no captures is not a year with no comments. Before writing one off, check whether the next year's captures reach back into it — they are the same fetch, at no extra cost.
  • The way to get more of an early year is to fetch more of the year after. There is no separate leg to run.

The measured shape for Le Matin: a trickle from July 2007 (4, 6, 10 a month), then 246 in November and 265 in December. The comment system opened in mid-2007 and found its public that autumn — four months before the URL grammar the archive can be searched by even existed.

The corpus spans three platforms, and identity gets worse over time (2026-08-30)

The recoverable past is not one era but two, and they differ in what they can be asked:

  • PHP "reactions", from somewhere in 2008 to about 2009. Thread inline on the article page, unpaginated, each comment bracketed in <!-- BEGIN/END COMMENT HTML -->. Every comment names a numeric idUser. This is the sturdiest identity in the corpus: a number does not move when the display name does, so it can confirm a rename outright. The URL carries the section before the id — /fr/actu/economie/<slug>_11-271110 — and the grammar appears part-way through 2008, not in January. Pages are charset=iso-8859-1.
  • Drupal, roughly 2009 to March 2012. The article page carries the whole thread inline; the URL is a slug with the id glued on the end. Every comment names /users/<key>, a real account identifier, with the display form beside it — MountaiDiver posting from /users/mountaidiver. The key is derived from the display name, so a rename moves it: weaker than a number, far better than nothing. No reply threading, no like counts.
  • Newsnetz, March 2012 to about 2017. The thread lives at ?comments=1. Reply threading survives in an HTML comment, and likes and dislikes are counted. No user id of any kind.

The direction is one way: a numeric id, then a name-derived slug, then nothing. The further back the material, the more it can settle.

And the layers are Le Matin's alone. Asked for the same two grammars, 24 heures and the Tribune return single digits or zero for every year — measured across all thirteen legs on 2026-08-31. Their recoverable history begins with Newsnetz in 2012. So the corpus is not three parallel histories of different depth; it is one deep history and two shallow ones, and a measure that needs pre-2012 material from all three titles cannot be built at any amount of fetching.

So identity is a different question on either side of March 2012. Before it, two comments can be tied to one account outright. After it, a nickname is all there is, and every claim that two handles are one writer is an inference from style — which is what alias_candidates and the proximity view exist to make, and why they return a shortlist and never a verdict. A method calibrated on the Newsnetz years cannot be validated on itself; the Drupal years are the only stretch of this corpus that carries ground truth about succession, and that makes them worth more per page than their comment counts suggest.

The corollary is a trap. author_key is present for one era and absent for the other, so any count grouped by it silently becomes a count of 2009-2012. Group by nickname unless the question is specifically about accounts, and say which era the answer stands on.

Which reader parses a page is decided by the markup, never by the capture year: the changeover was a deployment, and captures from the same week fall either side of it. Order the tests carefully — class="commentaire is a prefix of the PHP era's class="commentaires", so the loose Drupal test claims every 2008 page, and the wrong reader returns an empty thread rather than an error.

Encoding is part of the era, not a detail. The socket timeout measures inactivity and errors="replace" does not raise: both are silent failures that produce plausible-looking output. A listing that dribbles forever and a page whose every accent became \ufffd each cost a day before being noticed. Where a failure cannot announce itself, put a clock or a declaration in its way.

What a hand writes lives outside what a pass rewrites (2026-08-27)

Notes and off-platform accounts are their own tables (subject_note, subject_account), not columns on author_profile. Two reasons, and both matter:

  • A profiling run rewrites every column of an author_profile row. A remark typed by a reader would be erased by the next run of the machine.
  • Everything in author_profile is derived from the stored comments and reproducible from them. A note is not, and mixing the two would put an observation and a measurement in the same shape — the same mistake the metrics / inferred split exists to prevent.

Both are keyed (community, subject_kind, subject_key) like a profile. A persona reads its own records and those written against each of its handles, never the reverse: linking two nicknames must not bury what was already recorded about either, while a note about the person says more than any one handle can carry.

Off-platform accounts are structured rather than prose because the questions worth asking about them are counting ones — how many of a public carry an identity elsewhere, on which network, and whether a rename there lines up with a rename here. The platform is read off the pasted link rather than asked for, so the two cannot disagree.

Testing

pytest, tests in tests/ as test_*.py. Run with python3 -m pytest -q from the repo root. Unit tests must not require Postgres or network.