<span class="mw-page-title-main">Syncy</span>
Fabrice P. Lauss𝕪s oomie Web

Sync𝕐

Syncy is a script, made with Claude, which synchronizes the 𝕐 and ℤ posts made online (typically in real time) to my local laussywiki (llw), so as to keep everything harmonious in one and the same branch of the multiverse.

Versions

  • v°1.4 (10 August (2026)) — a file the gadget names in lowercase (its own lw_*.jpg uploads) was compared un-normalized against the API's ucfirst'd title, so it could never be found "present": every run downloaded it from online again and re-uploaded it into fileexists-no-change, reported as FAILED for ever. Titles now compare through canon() (ucfirst, underscores) and—belt and braces, the llw2lw 2.5.0 lesson—fileexists-no-change counts as "already here", never a failure.
  • v°1.5 (11 August (2026)) — ℤ joins 𝕐. Post ids are [YZ]- and page selection covers both branches: a bare run syncs the current month of each, --all walks 𝕐 from May 2026 and ℤ (the Z/ pages) from August 2026. The merge machinery was branch-clean already, so this is mostly page selection. An explicit title of either branch is also what llw2lw 3.5.0 hands syncy, one page at a time, when its own guard finds online posts local lacks.
  • v°1.6 (11 August (2026)) — deletions travel. A post deleted ONLINE, which the gadget signs with an invisible <!--deleted:ID--> tombstone, is dropped from the merge—deletion beats the local copy—instead of being faithfully resurrected as though it were a local post awaiting its push. Reported as "N deleted online".
  • v°1.7 (18 August (2026)) — no invented ids, and deletions travel both ways. A clash used to be settled by keeping both copies and giving the second a made-up id, Z-…-155034b; such a post is broken as written, because an id is a timestamp and {{thistime}} builds its stamp link from what the composer gave it, so the stamp lands on …-155034 and no longer matches its own blockquote, the gadget cannot lift it to the head of the post, and a {{continue}} hand-off from the other branch points at an anchor that is not on the page. Local's copy is kept, each declined copy is written to ~/.syncy-backups as …CLASH-<id>.txt, and the clash is reported on stdout—stderr being invisible inside an llw2lw run. A post deleted LOCALLY is now tombstoned too (gadget v1.46): the union with online used to hand it straight back and the next push republished it, so it could not be killed.
  • v°1.8 (18 August (2026)) — nothing between the posts is thrown away. The rebuild regenerates the posts region from the posts it could parse and keeps only the preamble and the postamble, so any other text living between the first post and the last—a stray paragraph, or a post whose <blockquote> wrapper a hand edit had broken—was destroyed in silence. The Würzburg post of 3 June (2026) went that way: a section edit on the 7th left its text outside any blockquote and the sync of the 9th swept it off the page. Such a page is now refused, with the first three hundred orphaned characters printed so one can see what is at stake.
  • v°1.9 (18 August (2026)) — post ids may be Anno Fabri tags. The composer (gadget v1.49) stopped naming a post by its timestamp spelled out and names it Y-006Nj instead: the same second in five base62 characters, unreadable without the Feistel of ~/bin/AF. Both spellings are recognised, and nothing here may go on slicing an id for its date—the day a post belongs to and its order inside that day are now decoded, an Anno Fabri tag neither showing its date at [2:10] nor sorting by it. A month page holds both kinds at once and merges the same either way.
  • v°2.0 (24 August (2026)) — a clash is settled, not shrugged off: the two copies are judged against ~/.syncy-base, the texts as they stood at the last clean sync, and only a post that has moved on both wikis is a question—one that now stops the machinery instead of scrolling past. See below.
  • v°2.1 (1 September (2026)) — the chrome is merged too. Everything on a month page that is not a post—the {{MonthNav}} header above the first one, whatever sits below the last—was taken from local and online's copy dropped unread, so a header edited at laussy.org was reverted by the next sync in silence: not merged, not backed up, not printed. It goes through the same three-way test as the posts now, and both sides moving is a question that stops the machinery like any other. Two smaller ones ride with it: a page refused for orphan text between its posts exited 0, so a refusal read as a clean run and llw2lw pushed local over the very text syncy had declined to touch; and the guard that decides whether syncy is called at all lives in llw2lw, which is why 3.12.0 goes with this.
  • v°2.2 (8 September (2026)) — a scheduled post is read through its wrapper. The composer's Schedule box writes a post to its month page at once, at the place its tag gives it, inside one <!-- --> pair that ~/bin/y-publish takes off when the moment comes. Nothing here knew that: the draft has an id, so it read as a post, and its two marker lines were left over as text between the posts—which the v°1.8 refusal rightly would not rebuild around, so the first scheduled post, of 7 September (2026), stalled its whole month for four llw2lw runs. Drafts are now lifted out whole before anything reads the page and put back, each where its tag's moment puts it, once the page is rebuilt. Local's copy of a draft wins; one written online is pulled in, loudly, since it has no line in the schedule and will not publish by itself. Goes with llw2lw 3.14.0.
  • v°2.3 (19 September (2026)) — Photo repository is pulled too. The composer's Repo tick files pictures on that page instead of writing a post, and from the phone that is a page written online which nothing brought home: it is not a month page, and it has no posts and no ids for any of the machinery here to hold on to. What it has is simpler—it only ever grows, by whole day-sections and whole <gallery> blocks—so the pages named in UNION_PAGES get a three-way merge of their own, by sections, in which nothing is ever a question. The pictures a gallery names come over like any post's. See below; goes with llw2lw 3.15.0.
  • v°2.4 (19 September (2026)) — the Cloudflare secret is no longer in the script. It was a literal in EXTRA_HEADERS, this page lists the source whole, and on 24 August (2026) a refresh of the listing from the file on disk put the real value here, over the placeholder I had set by hand; it stayed twenty-six days. load_headers() now reads it from the header = line of ~/.llw2lw, the one llw2lw and the other tools already use, so that there is nothing in the source to leak and rotating the secret is one line in one file. Goes with llw2lw 3.16.0, which refuses to push any page holding a value from the credential files.

When the edit was made online

On 24 August (2026) I noticed that the sentences I had added at laussy.org to a post about a golden number in an enneadecaeteris were not there any more. They were not, and neither were eight further edits to the month pages, the oldest of them from 6 July (2026). Every one of those had been made ONLINE in the ordinary MediaWiki section editor rather than in the microblog composer, and two ordinary pages outside this tool's reach had gone the same way, which brings the edits the sync destroyed to eleven; three more that came back with them were nobody's doing but my own de-duplication of 15 August (2026).

That is the whole of the trap. Which of two copies of a post is the newer one is read off the gadget's invisible <!--lastedited:--> marker, and the ordinary editor does not touch that marker—it has never heard of it. An edit made online therefore leaves the two copies looking exactly the same age while saying different things, which is not a verdict but the absence of one. Until V2.0 syncy called it a CLASH and disposed of it by keeping local, because local comes first in the list it consults; from V1.7 the declined words went to ~/.syncy-backups as a CLASH file—before that both copies were kept, the second under an invented id—and the event was announced by a single plain line in the middle of a two-hundred-line llw2lw run. Nothing was lost on disk and everything was lost in practice. Worse, syncy was usually never called at all: llw2lw's month-page reconcile compared SETS of post ids, where a post local lacks is a phone post and a tombstoned one is a deletion, so a post present on both sides that merely said something different matched neither test and the push went straight over it—no warning, no backup, no line in the log. That half is fixed in llw2lw 3.9.0, which compares the posts' text as well as their ids.

So V2.0 asks a question the marker can answer. syncy now keeps, per page, ~/.syncy-base/August_(2026).json.gz—the local text, the online text and the moment they last agreed, spaces and slashes flattened out of the title. A clash then becomes an ordinary three-way merge: the side that has MOVED away from the base is the side somebody wrote on, and it wins, whichever wiki it is. Only when both sides have moved since their last common state is there a real question, and that one is now loud: a banner, the declined copy in a CLASH file, and the machine-readable line SYNCY-UNRESOLVED: <page> <id> <file>, with exit status 2. llw2lw 3.9.0 reads that and refuses the page with syncfirst, which lands it in the Failures and Pending-retry blocks at the end of the run, where a line cannot be missed the way one line mid-run could.

The base is deliberately NOT written while anything on the page is unresolved. Were it written, the next run would find neither side moved—both would match a base taken after the divergence—hand the post to local, and bury the question for good; leaving the base where it is, is what keeps the question being asked. The base also has to move at the other end: llw2lw advances it after a verified push, from the very bytes whose sha1 it has just checked online. Without that, the base's idea of "what online said last time" stays at the pre-push text, so the next perfectly ordinary local edit looks like both sides moving, and every month page deadlocks on a clash that is not one.

A page with no base cannot be judged and its clashes are unresolved by construction, which is what syncy --seed-base is for: it declares the two wikis, as they stand right now, to be the reference a later clash is measured against, and merges nothing. It is needed once, when the store is introduced, and again after a page has been repaired by hand—otherwise the run stops for ever on a difference that is in fact already settled.

The repair has a regression test, ~/bin/syncy-selftest, which touches neither wiki: it loads ~/bin/syncy and ~/bin/llw2lw as modules and exercises settle(), diverged_posts() and llw2lw's local_ever_held()—the sha1 test that tells a legacy sync stamp from an edit made online—on the real words of the lost post. Twenty-three checks in six groups: the 23 August loss replayed, its three mirror cases (local moved, both moved, neither moved) and the no-base case, the two-pushes-in-a-row sequence with a NEGATIVE CONTROL that shows the deadlock when the base is left un-moved, three cases where the guard must not cry wolf, six on the sha1 test that decides whether a restamp may overwrite what is online, and a post that quotes a blockquote inside itself. Run it after touching either tool.

laussy@azag:~$ syncy-selftest

The 23 August 2026 loss, replayed
  ✓ his online edit survives the merge
  ✓ nothing is left for him to sort out
  ✓ and the merge says why
  ✓ llw2lw would have reconciled the page first

The other three ways round
  ✓ local moved, online did not → local wins
  ✓ both moved → unresolved, and local kept meanwhile
  ✓ neither moved → no question asked
  ✓ no base yet → unresolved rather than silently guessed

Two pushes in a row must not deadlock
  ✓ local edit settles in local's favour
  ✓ NEGATIVE CONTROL: a base left un-moved deadlocks
  ✓ with the base moved after the push, no deadlock
  ✓ llw2lw computes the sha1 the wiki reports

The guard must not cry wolf
  ✓ identical copies
  ✓ local edited with the gadget (its marker is later)
  ✓ phone post online only: the id guards handle that one

Whose text is online? (the restamp guard)
  ✓ a sha1 on the first page of history is found
  ✓ a sha1 only on a LATER page is found too (continuation)
  ✓   and it really did have to page for it
  ✓ a sha1 the local wiki never had is NOT found — this is the refusal
  ✓ no sha1 at all → False; a guard that cannot see waves nothing through
  ✓ page gone locally → False

A post that quotes a blockquote is read whole
  ✓ llw2lw reaches the real end tag
  ✓ and reads it exactly as syncy does

All good.

Recovering the fourteen edits was possible at all only because llw2lw imports rather than overwrites, so every version a post ever had online is still in laussy.org's history—which is what syncloss reads: it walks the online history of the 𝕐 and ℤ month pages, flags every point where a post's body goes back to a state it already had (prose does not revert by itself, so the state in between is an edit somebody made and something undid), and cross-examines it against the local history and against both wikis as they stand today.

When the edit was not in a post

V2.0 taught the posts to ask which side had moved. It did not occur to me to ask it of anything else on the page, and on 1 September (2026) that came back. At laussy.org I pointed a link in the {{MonthNav}} header of September (2026) at an {{anchor}} I had just planted in the August page, doubled the ☕ of its label three minutes later, and four minutes after that wrote a post on the same page in the same editor. The sync kept the post and put the local header back over both of the other edits. The post had an id; the header had none.

A month page is rebuilt here out of the posts that can be parsed, and everything else is chrome: split_chrome() returns what stands before the first post and what stands after the last, and until V2.1 it read them from the local text and from nothing else. Its own docstring stated the rule plainly—local is authoritative for chrome; online's header (if any) is ignored—so the two were never compared, a difference could not be reported, and no CLASH file was written for words nobody had noticed were about to go. settle_chrome() now asks of the header, and of the footer, precisely what settle() asks of a post: the side that has moved away from the base is the side somebody wrote on. Tombstones are stripped before the comparison, since they are rebuilt from both texts and their coming and going is bookkeeping rather than anybody's writing.

The merge alone would not have saved those words, which is the part worth keeping. syncy only ever sees a month page because llw2lw hands it one, and that decision was made on three tests that all count posts—a post local lacks, a post tombstoned online, a post whose words moved. Had I touched only the header, syncy would not have been called at all. llw2lw 3.12.0 adds diverged_chrome() beside them, and the pair of fixes is one fix: a merge in the tool that is never called is no fix.

~/bin/syncy-selftest replays the loss and its mirror cases, and asserts on every run that llw2lw.chrome_of() and syncy.split_chrome() read the same chrome out of the same text—two readings that drift apart being two different pages. That check failed the first time it ran: llw2lw's copy ended the postamble before the closing </blockquote> and this one ended it after.

When the page has no posts

Everything above is about posts: a thing with an id, which can be found on both wikis and compared with itself. Photo repository has none. It came on 19 September (2026) with the Repo tick of the composer, which files what is in the box—pictures shared from the phone's gallery, mostly—under a heading that is the day, == {{af|0Btah|h}} ==, in a <gallery>. From the phone that is written at laussy.org, and I had built the way in without a way back: the page was in nobody's Recent Changes here, so llw2lw never looked at it, and the day I edited it locally the best that could happen was the onlineahead refusal of 3.11.0, for good.

Such a page only grows, and that is what merge_union() leans on. The two texts and the base of ~/.syncy-base are cut at their headings, and each section goes through the question V2.0 taught the posts—which side has moved?—with one difference: when both have, nobody is asked. Local's blocks stand and online's new ones follow them; a section only one side has is kept, unless the base had it too, in which case the other side deleted it and it stays deleted; a section filed online for a day local already has, under another tag (two tags of one day share no letters, so the headings are compared by decoding them), is poured into local's, since the page is one section a day; and the newest day goes first. A page that only one side touched is taken from that side to the letter, so nothing is re-laid-out that did not have to be.

The first version did none of this. It handed the three texts to git merge-file --union, which is the same idea by lines, and it passed every test but the one I nearly did not write: both wikis filing on the same day. The two new sections then sit at the same place, each ending in a </gallery> that looks exactly like the other's, and diff3 pairs the look-alikes off—one gallery came out with no closing tag and swallowed the section after it. A page made of blocks has to be merged by blocks.

The base moves as it does for a month page: here when the two wikis already agree, otherwise by llw2lw once the push that makes them agree is verified. A bare syncy and syncy --all both take the page in; syncy "Photo repository" does it alone.

Code

#!/usr/bin/env python3
# ___
#/ __|_  _ _ _  __ _|      _|
#\__ \ || | ' \/ _|  _|  _|
#|___/\_, |_||_\__|    _|
#     |__/ SyncY       _|
# V2.4 Laussy&Claude   _|
# Synchronize Y posts from lw to llw
#
# V2.4: the Cloudflare secret is no longer in this file. It was a literal in
#       EXTRA_HEADERS, this source is published whole on [[Syncy]], and on
#       24 August 2026 a refresh of that listing from the file on disk put the
#       real value online, over the placeholder Fabrice had set there by hand;
#       it stayed 26 days. load_headers() reads it from ~/.llw2lw (the line
#       llw2lw itself uses) or ~/.syncy. Nothing else changes.
#
# V2.3: [[Photo repository]] is pulled too. The composer's "Repo" tick (gadget
#       v1.79, 2026-09-19) files pictures there instead of writing a post, and
#       from the phone that is a page written ONLINE which nothing brought
#       back: not a month page, so not syncy's; never edited here, so never in
#       an llw2lw run either — and the day it was edited here, llw2lw's
#       'onlineahead' refusal was the best that could happen to it. It has no
#       posts and no ids, so none of the machinery below applies. What it has
#       is simpler: it only ever GROWS, by whole sections and whole <gallery>
#       blocks. So a page listed in UNION_PAGES is settled by a plain
#       three-way text merge against ~/.syncy-base (git merge-file --union):
#       what one side changed since the base is taken, and where both added
#       at the same place BOTH additions are kept, one after the other —
#       nothing is a question, nothing is refused. The files a <gallery>
#       names come over like any post's. The base moves as for a month page:
#       here when the two wikis already agree, otherwise by llw2lw (3.15.0)
#       once its push is verified. Named explicitly, or part of the default
#       run and of --all.
#
# V2.2: a scheduled post is read through its wrapper. The composer's Schedule
#       box (gadget v1.65, 2026-09-04) writes a post to its month page at
#       once, at the place its tag gives it, wrapped in one <!-- / --> pair
#       by draftify() — exactly the lines the splice added, day heading
#       included when the day was new — and ~/bin/y-publish takes the two
#       marker lines off when the moment comes. Nothing here knew that. The
#       draft's <blockquote> has an id, so it read as a post, and the markers
#       were left over: '<!--' as the last line of the chrome, '-->' orphaned
#       between the posts — which is what stalled the first scheduled post,
#       Y-0F8dL of 7 September 2026, for four llw2lw runs ("3 character(s)
#       sit between the posts"). The V1.8 refusal did its job: had the page
#       been rebuilt, the '-->' would have been dropped and every post down
#       to the first <!--lastedited:--> commented out of existence.
#
#       Draft blocks are now lifted out whole before anything reads a page
#       (lift_drafts) and put back once it is rebuilt, each where its id's
#       moment puts it (place_drafts) — the same rule the composer's
#       insertIntoPage() follows, applied to an opaque block whose inside
#       nothing may touch, since y-publish's whole job is to remove exactly
#       two lines from it. Local is the base of record for drafts as for
#       everything else: local's copy wins, a draft whose post is live or
#       tombstoned on either side goes, a draft the base of the last push
#       carried and local no longer has was cancelled here and stays so. One
#       thing is pulled: a draft written ONLINE (the phone runs the same
#       composer), which the base never saw — it comes in, loudly, since it
#       has no line in [[Private:Scheduled posts]] and will not publish by
#       itself. llw2lw 3.14.0 reads pages the same way (strip_drafts).
#
# V2.1: the chrome is merged, not assumed. Everything on a month page that is
#       not a post — the {{MonthNav}} header above the first one, whatever
#       sits below the last — was taken from LOCAL and online's copy dropped
#       unread, so an edit made to the header at laussy.org was reverted by
#       the next sync without a word: not merged, not backed up, not printed.
#       That is the V2.0 blind spot in its last hiding place, and it bit on
#       1 September 2026 — a link retargeted to a new {{anchor}} at 17:55 and
#       its coffee-cup label doubled at 17:58, both gone at 18:04, while the
#       post written online in the same editor minutes later came through
#       untouched because it had an id. Chrome now goes through settle()'s
#       three-way test against ~/.syncy-base, and both sides moving is a
#       question that stops the machinery like any other. Also: a 'refused'
#       page (orphan text between the posts) now exits 2 as it always should
#       have — main() looked only for 'unresolved', so a refusal read as a
#       clean run and llw2lw pushed local over the very text syncy declined
#       to touch.
#
# V2.0.1: the count above said thirteen edits from 13 June; the 13 June one
#       (Y-20260613-011637) was superseded by a local rewrite the next day, not
#       destroyed, so the oldest real loss is 6 July and there are eleven.
#
# V2.0: a clash is settled, not shrugged off. Two copies of one post with the
#       same last-edit time and different words used to be "keep local, write
#       the other out, print a line" — and the line goes by inside an llw2lw
#       run. That is how the golden-number sentences of Y-01on2 died on
#       24 August 2026, along with twelve more edits going back to 13 June:
#       every one of them made ONLINE in the ordinary MediaWiki editor, which
#       does not touch the <!--lastedited:--> marker, so the two copies looked
#       the same age and local won by being first in the list.
#
#       The marker cannot tell them apart, so stop asking it to. syncy now
#       remembers, per page, the two texts as they stood when it last ran
#       (~/.syncy-base), and a clash becomes an ordinary three-way merge:
#       whichever side has MOVED since that base is the side that wrote
#       something, and it wins. Only when BOTH have moved is there a real
#       question, and that one is now loud — a banner, a CLASH file, the line
#       "SYNCY-UNRESOLVED: <page> <id> <file>" for llw2lw to read, and exit
#       status 2. llw2lw 3.9.0 turns that into a refusal to push the page, so
#       nothing can overwrite an answer nobody has given yet.
#
#       The base is deliberately NOT written when anything is unresolved: were
#       it written, the next run would see neither side move, hand it to local
#       and bury the question for good.
#
# V1.9: post ids may be Anno Fabri tags. The composer (gadget v1.49) stopped
#       naming a post by its timestamp spelled out and names it Y-006Nj
#       instead — the same second, five base62 characters, unreadable without
#       the Feistel of ~/bin/AF. Both spellings are recognised, and nothing
#       else here may go on slicing an id for its date: the day a post belongs
#       to and the order it takes inside that day are now DECODED (read_id),
#       an Anno Fabri tag neither showing its date at [2:10] nor sorting by
#       it. Old ids are untouched, so a month page holds both kinds at once
#       and merges the same either way.
#
# V1.8: nothing between the posts is thrown away. The rebuild regenerates the
#       posts region from the posts it could PARSE and keeps the preamble and
#       postamble, so any other text living between the first post and the
#       last — a stray paragraph, or a post whose <blockquote> wrapper a hand
#       edit had broken — was silently destroyed. The 3 June 2026 Würzburg
#       post went that way: a section edit on 7 June left its text outside any
#       blockquote, and the sync of 9 June swept it off the page. syncy now
#       refuses to write such a page and prints what is at stake.
#
# V1.7: no invented ids, and deletions travel BOTH ways.
#
#       A CLASH — two copies of one post with the same last-edit time and
#       different words — used to be settled by keeping both and giving the
#       second a made-up id, Z-…-155034b. Such a post is broken as written: an
#       id IS a timestamp, and {{thistime}} builds its stamp link from what
#       the composer gave it, so the stamp lands on …-155034, never on the
#       invented …-155034b. The
#       stamp then does not match its own blockquote, so the gadget cannot
#       lift it to the head of the post and it sits at the foot; and the
#       {{continue}} hand-off from the other branch points at an anchor that
#       is not on the page, so the link from 𝕐 to its ℤ half dies. Local's
#       copy is kept, each declined copy is written to ~/.syncy-backups as
#       …CLASH-<id>.txt, and the clash is reported on stdout — stderr being
#       invisible inside an llw2lw run.
#
#       Deletions: A post deleted LOCALLY used to be
#       invisible to this merge — the union with online handed it straight
#       back, the next push republished it, and it could not be killed. The
#       gadget (v1.46) now signs a local delete with the same
#       <!--deleted:ID--> tombstone it has always written online, and a
#       tombstone in EITHER text buries the post here. The marker is kept
#       only while the other side still carries the post, and dropped once it
#       does not, so the pages do not accumulate them.
#
# V1.6: deletions travel. A post deleted ONLINE (gadget v1.36 leaves an
#       invisible <!--deleted:ID--> tombstone there) is dropped from the
#       merge — deletion beats the local copy — instead of being faithfully
#       resurrected as if it were a local post awaiting its push. Reported
#       as "N deleted online (…)". Local deletions never needed help: llw2lw
#       pushes the whole local text.
#
# V1.5: ℤ joins 𝕐. Post ids are [YZ]-, and page selection covers both
#       branches: the default run syncs the current month of each, --all
#       walks 𝕐 from May 2026 and ℤ (the "Z/" pages) from August 2026. An
#       explicit title of either branch always worked once the id regex
#       allowed it — which is also what llw2lw 3.5.0 calls per page. The
#       merge machinery was branch-clean already (ids were grouped by their
#       [2:10] date slice — V1.9 decodes it instead — backups flatten "/").
#
# V1.4: a file referenced with a lowercase first letter (the gadget's own
#       lw_*.jpg names) was compared un-normalized against the API's
#       ucfirst'd title, so it could never be "present": every run
#       re-downloaded it and re-uploaded it into fileexists-no-change,
#       reported as FAILED for ever. Titles are now compared through canon()
#       (ucfirst, underscores), and — belt and braces, the llw2lw 2.5.0
#       lesson — fileexists-no-change counts as "already here", never a
#       failure.
#
"""
syncy                 -> the CURRENT month of BOTH branches, 𝕐 and ℤ
                         (e.g. "June (2026)" and "Z/June (2026)")
syncy "May (2026)"    -> a specific month page (either branch)
syncy --all           -> every month from each branch's start (𝕐: May 2026,
                         ℤ: August 2026) to the current one
syncy --seed-base     -> take the two wikis as they stand right now to be the
                         reference a later clash is judged against, and merge
                         nothing. Needed once, and again after a page has been
                         repaired by hand.

Exit status: 0 everything settled; 2 a post changed on BOTH wikis since their
last common state and only Fabrice can say which words to keep. llw2lw 3.9.0
reads that and refuses to push the page rather than overwrite the answer.

Pull a 𝕐 or ℤ month-page from the ONLINE wiki into the LOCAL wiki:
fetch both, merge by day/time (latest edit wins), and write the result back to
the LOCAL wiki only — local is the base of record. Online is never modified.

It writes only when online actually contributed something (a post local lacks,
or a newer edit of a post local has). If local already contains everything
online has, nothing is written. A backup of the local page is saved first.

Which copy of a post is the newer one is read off the gadget's invisible
<!--lastedited:--> marker — except that an edit made online in the ORDINARY
MediaWiki editor never touches it, so the two copies read as the same age with
different words. Since V2.0 that is settled against ~/.syncy-base, the two
texts as they stood at the end of the last clean sync of the page: the side
that has MOVED away from them is the side somebody wrote on, and it wins. Only
when both have moved is there a question, and a question stops the machinery.

Setup:
  - Install on PATH:   cp syncy ~/.local/bin/syncy && chmod +x ~/.local/bin/syncy
  - Credentials: a MediaWiki bot password from the LOCAL wiki's
    Special:BotPasswords. Grants needed: "Edit existing pages" AND "Upload,
    replace, and move files" (uploads need their own grant). Provide via env
    vars  SYNCY_USER / SYNCY_PASS  or a ~/.syncy file (chmod 600):
        user = Fabrice@syncy
        pass = xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
  - Local must have $wgEnableUploads = true (it does, since the gadget uploads).
"""
import sys, os, re, html, json, gzip, datetime
import urllib.request, urllib.parse, http.cookiejar

# ---- config ---------------------------------------------------------------
LOCAL_BASE  = 'http://azag/laussywiki/index.php/'   # article path (read)
LOCAL_API   = 'http://azag/laussywiki/api.php'      # API (write)
ONLINE_BASE = 'https://laussy.org/wiki/'                     # article path (read)

USER_AGENT  = 'Mozilla/5.0 (Y-sync)'
# Secret header to get the ONLINE read past Cloudflare (matches your WAF rule).
# NOT written here (V2.4): this source is listed in full on the wiki page
# [[Syncy]], and the listing of 24 August 2026 took the value online with it.
# It is read from the "header = Name: value" line(s) of ~/.llw2lw — the file
# llw2lw and the other tools already read it from, so that rotating the secret
# is one line in one file — or of ~/.syncy.
def load_headers():
    found = {}
    for path in ('~/.llw2lw', '~/.syncy'):
        try:
            lines = open(os.path.expanduser(path), encoding='utf-8').readlines()
        except OSError:
            continue
        for line in lines:
            key, eq, val = line.strip().partition('=')
            if eq and key.strip() == 'header' and ':' in val:
                name, _, hv = val.partition(':')
                found.setdefault(name.strip(), hv.strip())
    return found


EXTRA_HEADERS = load_headers()
BACKUP_DIR = os.path.expanduser('~/.syncy-backups')
# The two texts as they stood at the end of the last clean sync of a page.
# A clash is decided by asking which side has moved away from them (V2.0).
BASE_DIR = os.path.expanduser('~/.syncy-base')
# Per-branch page prefix and first month — the inclusive lower bound used
# by --all. 𝕐 lives at the bare title, ℤ under the ASCII "Z/" prefix.
BRANCHES = [('', 2026, 5),             # 𝕐: May 2026
            ('Z/', 2026, 8)]           # ℤ: August 2026
# ---------------------------------------------------------------------------

# Pages that are not month pages but are written online all the same, and only
# ever by adding to them (V2.3). Settled by sync_union(), not sync_page().
UNION_PAGES = ['Photo repository']

MONTHS = ['January', 'February', 'March', 'April', 'May', 'June',
          'July', 'August', 'September', 'October', 'November', 'December']
# ---- Anno Fabri -----------------------------------------------------------
# Since 18 August 2026 a post is named by its Anno Fabri tag — Y-006Nj — where
# it used to be named by its timestamp spelled out, Y-20260820-145603. The tag
# holds the same second in five base62 characters (two of era, three put
# through the 6-round Feistel of ~/bin/AF and Module:AF), plus a sixth letter
# carrying the UTC offset when the post was written away from home. Nothing is
# lost — the transform is a bijection — but a date no longer reads off the id
# by eye. Ids written before that day keep their old spelling and are still
# read: read_id() below takes either.
AF_A = '0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz'
AF_ZL = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz'
AF_OFF = [-39600, -36000, -34200, -32400, -28800, -25200, -21600, -18000,
          -14400, -12600, -10800, -9000, -7200, -3600, 0, 3600, 7200, 10800,
          12600, 14400, 16200, 18000, 19800, 20700, 21600, 23400, 25200,
          28800, 31500, 32400, 34200, 36000, 37800, 39600, 43200, 45900,
          46800, 49500, 50400]
AF_BASE, AF_ERA, AF_a, AF_b, AF_SALT = 1787004000, 238328, 62, 3844, 232959
AF_K = [0x2f4b6c1d, 0x7a19e3b5, 0x5c83d271, 0x1e6fa94b, 0x3d5e8a17, 0x64c2b90f]
# A post id in either spelling. Deliberately not used to FIND ids in free
# wikitext — five loose base62 characters would match half the prose — only to
# recognise one already taken from a blockquote or a tombstone.
ID_RE = r'[YZ]-(?:\d{8}-\d{6}|[0-9A-Za-z]{5,6})'


def _af_f(x, i, tw):
    h = ((x ^ AF_K[i] ^ tw) * 0x9E3779B1) & 0xFFFFFFFF
    h = ((h ^ (h >> 15)) * 0x85EBCA6B) & 0xFFFFFFFF
    h = ((h ^ (h >> 13)) * 0xC2B2AE35) & 0xFFFFFFFF
    return h ^ (h >> 16)


def _af_home(t):
    """Europe/Madrid's offset at t. CEST between the last Sundays of March and
    October, each switching at 01:00 UTC: the whole EU rule."""
    y = datetime.datetime.fromtimestamp(t, datetime.timezone.utc).year

    def switch(mo):
        d = datetime.date(y, mo, 31)
        d -= datetime.timedelta(days=(d.weekday() + 1) % 7)     # back to Sunday
        return int(datetime.datetime(d.year, d.month, d.day,
                                     tzinfo=datetime.timezone.utc)
                   .timestamp()) + 3600
    return 7200 if switch(3) <= t < switch(10) else 3600


def af_read(tag):
    """A tag -> (the instant, the offset its wall clock is written in)."""
    if not re.match(r'^[0-9A-Za-z]{5,6}$', tag):
        return None
    hi = lo = 0
    for i in range(5):
        k = AF_A.index(tag[i])
        if i < 2:
            hi = hi * 62 + k
        else:
            lo = lo * 62 + k
    L, R, tw = lo // AF_b, lo % AF_b, hi * AF_SALT
    for i in (4, 2, 0):
        R = (R - _af_f(L, i + 1, tw)) % AF_b
        L = (L - _af_f(R, i, tw)) % AF_a
    t = AF_BASE + hi * AF_ERA + L * AF_b + R
    if len(tag) == 6:
        k = AF_ZL.find(tag[5])
        if k < 0 or k >= len(AF_OFF):
            return None
        return t, AF_OFF[k]
    return t, _af_home(t)


def read_id(pid):
    """A post id, in either spelling -> (the wall clock it was written at, as a
    naive datetime, and the instant as a Unix time). None if it is not one."""
    m = re.match(r'^[YZ]-(\d{4})(\d{2})(\d{2})-(\d{2})(\d{2})(\d{2})$', pid)
    if m:
        w = datetime.datetime(*map(int, m.groups()))
        return w, w.astimezone().timestamp()
    m = re.match(r'^[YZ]-([0-9A-Za-z]{5,6})$', pid)
    r = af_read(m.group(1)) if m else None
    if not r:
        return None
    t, off = r
    return (datetime.datetime.fromtimestamp(t + off, datetime.timezone.utc)
            .replace(tzinfo=None), t)

POST_RE = re.compile(r'<blockquote id="(%s)">(.*?)</blockquote>' % ID_RE, re.S)
POST_OPEN = re.compile(r'<blockquote id="(%s)">' % ID_RE)
# A post deleted ONLINE signs itself with this invisible marker (gadget
# v1.36): without it, a deletion there is indistinguishable from a local
# post not yet pushed, and the merge below would faithfully resurrect it.
TOMBSTONE = re.compile(r'<!--deleted:(%s)-->' % ID_RE)
BQ_OPEN = re.compile(r'<blockquote\b', re.I)
BQ_CLOSE = re.compile(r'</blockquote\s*>', re.I)

# ---- scheduled posts: the composer's <!-- --> drafts (V2.2) ---------------
# The shape draftify() writes and y-publish looks for: the two markers on
# lines of their own, the post — heading and all, when its day was new —
# between them, verbatim. Only a comment holding a post's <blockquote> is a
# draft; any other comment on the page is left exactly where it is.
DRAFT_RE = re.compile(r'(?ms)^<!--\n(.*?)\n-->$')
HEAD_RE = re.compile(
    r'^==\s*\{\{[Tt]hisday\|(\d+)\|([A-Za-z]+)\|(\d+)\}\}\s*==\s*$')


def lift_drafts(text):
    """-> (the page without its draft blocks, {pid: block}) — the block being
    the whole thing, markers included, since the one operation ever made on
    it is y-publish taking those two lines off and nothing inside may move.
    """
    drafts = {}

    def cut(m):
        ids = POST_OPEN.findall(m.group(1))
        if len(ids) != 1:
            return m.group(0)             # not ours (or crowded: see below)
        drafts[ids[0]] = m.group(0)
        return ''
    return DRAFT_RE.sub(cut, text), drafts


def strip_drafts(text):
    return lift_drafts(text)[0]


def crowded_drafts(text):
    """Comment blocks holding MORE than one post: the shape the v1.65 bug
    produced (the whole rest of the page wrapped into the new draft), which
    lift_drafts() leaves alone so that the page is refused, not rebuilt."""
    return [POST_OPEN.findall(m.group(1)) for m in DRAFT_RE.finditer(text)
            if len(POST_OPEN.findall(m.group(1))) > 1]


def heading_day(line):
    """The day a '== {{thisday|d|Month|y}} ==' line stands for, as YYYYMMDD,
    or None if the line is not one."""
    m = HEAD_RE.match(line.strip())
    if not m or m.group(2) not in MONTHS:
        return None
    return '%04d%02d%02d' % (int(m.group(3)), MONTHS.index(m.group(2)) + 1,
                             int(m.group(1)))


def keep_drafts(local, online, live, dead, pushed):
    """Which draft blocks the rebuilt page gets back, and what to say.

    `local` and `online` are the two pages' drafts, `live` the ids of the
    posts the merge is keeping, `dead` the tombstoned ids, `pushed` the
    drafts online held at the end of the last clean sync. Returns (kept,
    notes). Local is the base of record: its copy wins over online's, a
    draft whose post is live or dead goes, and a draft online alone is a
    cancellation if the base carried it (local took it off by hand and the
    push will take it off online) — otherwise it was written online and
    comes in, said out loud because nothing will publish it by itself.
    """
    kept, notes = {}, []
    for pid, block in local.items():
        if pid in live:
            notes.append('scheduled post %s is live: its draft goes' % pid)
        elif pid in dead:
            notes.append('scheduled post %s was deleted: its draft goes' % pid)
        else:
            kept[pid] = block
            if pid in online and online[pid] != block:
                notes.append("scheduled post %s differs online; local's "
                             "copy kept" % pid)
    for pid, block in online.items():
        if pid in local or pid in live or pid in dead:
            continue
        if pid in pushed:
            notes.append('scheduled post %s was cancelled here; the next '
                         'push takes it off the online page' % pid)
        else:
            kept[pid] = block
            notes.append('scheduled post %s was written ONLINE — pulled in, '
                         'but it has no line in [[Private:Scheduled posts]] '
                         'and will not publish by itself (y-publish --now %s '
                         '--push, or add its line)' % (pid, pid))
    return kept, notes


def place_drafts(text, drafts):
    """Put the draft blocks back into a rebuilt page, each where its id's
    moment puts it — the composer's insertIntoPage() rule, applied to an
    opaque block: under its day's heading, above the first post older than
    it (or closing the day); for a day the page does not have, above the
    first day older than its own; older than every day, after the last post.
    Every block is placed against the same draft-free page, newest first at
    a shared spot, so two drafts never decide each other's place.

    No blank lines are added around a block: draftify() wrapped whatever
    blank line the splice itself added, so the block carries its own.
    """
    if not drafts:
        return text
    lines = text.rstrip('\n').split('\n')
    heads = [(i, heading_day(l)) for i, l in enumerate(lines)]
    heads = [(i, d) for i, d in heads if d]
    ends = [text[:e].count('\n') for _p, _i, _s, e in find_posts(text)]
    slots = []
    for pid, block in drafts.items():
        r = read_id(pid)
        day, instant = (r[0].strftime('%Y%m%d'), r[1]) if r else ('99999999', 0)
        mine = [i for i, d in heads if d == day]
        if mine:
            h = mine[0]
            end = next((i for i, _d in heads if i > h), len(lines))
            at, seen = None, False
            for k in range(h + 1, end):
                m = POST_OPEN.match(lines[k].strip())
                if not m:
                    continue
                seen = True
                q = read_id(m.group(1))
                if q and q[1] < instant:
                    at = k                        # above the first older one
                    break
            if at is None and seen:
                at = end                          # older than all: closes it
                while at > h + 1 and lines[at - 1].strip() == '':
                    at -= 1
            elif at is None:
                at = h + 1                        # a heading with nothing yet
                while at < len(lines) and lines[at].strip() == '':
                    at += 1
        else:
            older = [i for i, d in heads if d < day]
            at = older[0] if older else (max(ends) + 1 if ends else len(lines))
        slots.append((at, -instant, pid, block))
    slots.sort()
    out, pos = [], 0
    for at, _neg, _pid, block in slots:
        out.extend(lines[pos:at])
        out.extend(block.split('\n'))
        pos = at
    out.extend(lines[pos:])
    return '\n'.join(out).lstrip('\n') + '\n'

# The 𝕐 month pages read NEWEST FIRST (since 2026-07-27): the newest day at the
# top of the page, the newest post at the top of its day. Set this to False to
# write them back the old way round, oldest first.
NEWEST_FIRST = True
TEXT_RE = re.compile(r'<text\b[^>]*>(.*?)</text>', re.S)
EDIT_RE = re.compile(r'<!--lastedited:(\d{8}-\d{6})-->')
FILE_RE = re.compile(r'\[\[\s*(?:File|Image)\s*:\s*([^|\]]+?)\s*(?:\|[^\]]*)?\]\]', re.I)


# ---- read (Special:Export, no auth needed for public reads) ---------------
def fetch(base, title):
    url = base + 'Special:Export/' + urllib.parse.quote(title.replace(' ', '_'))
    headers = {'User-Agent': USER_AGENT}
    headers.update(EXTRA_HEADERS)
    req = urllib.request.Request(url, headers=headers)
    with urllib.request.urlopen(req, timeout=30) as r:
        xml = r.read().decode('utf-8', 'replace')
    m = TEXT_RE.findall(xml)
    return html.unescape(m[-1]) if m else ''   # '' if the page doesn't exist


# ---- merge ----------------------------------------------------------------
def last_changed(pid, inner):
    """When this copy of the post was last touched, as 'YYYYMMDD-HHMMSS' — the
    latest edit stamp in it, or, for a post never edited, the moment it was
    written. That moment used to be pid[2:], the id being the timestamp; an
    Anno Fabri id has to be decoded to give it, and must be, since the two
    copies of a post are compared on this string."""
    m = EDIT_RE.findall(inner)
    if m:
        return max(m)
    r = read_id(pid)
    return r[0].strftime('%Y%m%d-%H%M%S') if r else pid[2:]


def find_posts(text):
    """Yield (pid, inner, start, end) for every post.

    A post may quote something in a <blockquote> of its own, so the first
    </blockquote> after the opening tag is not necessarily the post's: count
    the nesting. (POST_RE did take the first one, which silently truncated
    such a post — and everything the truncated tail held, its {{thistime}}
    included — every time a page was synced.)
    """
    for m in POST_OPEN.finditer(text):
        depth, pos = 1, m.end()
        while depth:
            nxt_c = BQ_CLOSE.search(text, pos)
            if not nxt_c:
                break                      # unbalanced: leave this one alone
            nxt_o = BQ_OPEN.search(text, pos)
            if nxt_o and nxt_o.start() < nxt_c.start():
                depth, pos = depth + 1, nxt_o.end()
            else:
                depth, pos = depth - 1, nxt_c.end()
                if depth == 0:
                    yield m.group(1), text[m.end():nxt_c.start()], m.start(), pos


def collect(texts):
    kept, clashes = {}, {}
    for t in texts:
        for pid, inner, _start, _end in find_posts(t):
            lc = last_changed(pid, inner)
            if pid not in kept:
                kept[pid] = (lc, inner)
            else:
                old_lc, old_inner = kept[pid]
                if lc > old_lc:
                    kept[pid] = (lc, inner)
                elif lc == old_lc and old_inner.strip() != inner.strip():
                    clashes.setdefault(pid, [old_inner]).append(inner)
    return kept, clashes


def build(kept, clashes):
    # No id is ever invented here (V1.7). A clash used to be resolved by
    # keeping both copies and giving the second a made-up id — Z-…-155034b —
    # and such a post is broken the moment it is written: a post's id IS its
    # timestamp, and {{thistime}} builds its own stamp link from what the
    # composer gave it: it lands on …-155034, never on …-155034b. The stamp
    # therefore no longer matches its own blockquote, the gadget cannot find
    # it to lift it to the top of the post (it stays at the foot, which is the
    # visible symptom), and the {{continue}} hand-off written by the other
    # branch points at an anchor that is not on the page — the link from 𝕐 to
    # its ℤ half simply dies. The losing text is written out beside the backup
    # instead, and the clash reported, for Fabrice to merge by hand.
    posts = {pid: inner for pid, (lc, inner) in kept.items()}
    # The day a post belongs to, and the instant that orders it inside that
    # day, both come out of its id — read, not sliced: an Anno Fabri id is
    # its date just as much as the old spelling was, but it neither shows it
    # at pid[2:10] nor sorts by it (the low three characters are scrambled on
    # purpose). An id that will not decode keeps its string order, which is
    # the best that can be done with it and puts it nowhere surprising.
    days, when = {}, {}
    for pid in posts:
        r = read_id(pid)
        days.setdefault(r[0].strftime('%Y%m%d') if r else pid[2:10],
                        []).append(pid)
        when[pid] = r[1] if r else 0
    out = []
    for day in sorted(days, reverse=NEWEST_FIRST):
        y, mo, d = int(day[0:4]), int(day[4:6]), int(day[6:8])
        out.append('== {{thisday|%d|%s|%d}} ==' % (d, MONTHS[mo - 1], y))
        first = True
        for pid in sorted(days[day], key=lambda q: (when[q], q),
                          reverse=NEWEST_FIRST):
            if not first:
                out.append('-----')
            first = False
            out.append('<blockquote id="%s">%s</blockquote>' % (pid, posts[pid]))
        out.append('')
    return '\n'.join(out).rstrip() + '\n'


def normalize(kept):
    """Post-level view for change detection (ignores whitespace/formatting)."""
    return {pid: inner.strip() for pid, (lc, inner) in kept.items()}


# ---- preserve non-post chrome (e.g. the {{MonthNav}} header) --------------
def split_chrome(text):
    """Return (preamble, postamble): the page content before the first post
    element and after the last one. This keeps non-post material — above all
    the {{MonthNav}} header with its trip annotations — from being lost when
    the posts section is regenerated.

    Until V2.1 the caller read this from the LOCAL text and nothing else:
    "local is authoritative for chrome, online's header is ignored". That is
    the posts' V2.0 blind spot in its last hiding place — chrome edited at
    laussy.org was not merged, not backed up, not reported, just quietly
    replaced by local's copy on the next sync. See settle_chrome()."""
    heads = [m.start() for m in re.finditer(r'(?m)^==\s*\{\{[Tt]hisday\b', text)]
    posts = list(find_posts(text))
    starts = heads + [p[2] for p in posts]
    if not starts:
        return text.strip('\n'), ''          # no posts: keep it all as preamble
    first = min(starts)
    last = max([p[3] for p in posts]) if posts else first
    return text[:first].strip('\n'), text[last:].strip('\n')


def orphan_text(text):
    """Wikitext BETWEEN the posts that the rebuild would not put back.

    build() regenerates the posts region from the posts it could parse, and
    with_chrome() keeps the preamble and the postamble — so anything else
    living between the first post and the last simply ceases to exist. A
    stray paragraph, or a post whose <blockquote> wrapper got damaged by a
    hand edit, is destroyed by the next sync without a word.

    That is not a theory: the 3 June 2026 Würzburg post was lost exactly so.
    A section edit on 7 June left its text on the page but outside any
    blockquote; the sync of 9 June ("syncy: import online posts") swept it
    away, and it was two months before anyone noticed.

    Returns whatever would be dropped, stripped of the furniture the rebuild
    does reproduce — day headings, separators, tombstones, blank space.
    """
    posts = sorted(find_posts(text), key=lambda p: p[2])
    if not posts:
        return ''
    heads = [m.start() for m in re.finditer(r'(?m)^==\s*\{\{[Tt]hisday\b', text)]
    first = min([posts[0][2]] + heads)
    last = posts[-1][3]
    gaps, pos = [], first
    for _pid, _inner, st, en in posts:
        if st > pos:
            gaps.append(text[pos:st])
        pos = max(pos, en)
    if last > pos:
        gaps.append(text[pos:last])
    rest = '\n'.join(gaps)
    rest = re.sub(r'(?m)^==\s*\{\{[Tt]hisday\|[^}]*\}\}\s*==\s*$', '', rest)
    rest = re.sub(r'(?m)^-{3,}\s*$', '', rest)
    rest = TOMBSTONE.sub('', rest)
    return rest.strip()


def with_chrome(preamble, posts, postamble):
    chunks = [c for c in (preamble.strip('\n'), posts.strip('\n'),
                          postamble.strip('\n')) if c]
    return '\n\n'.join(chunks) + '\n'


# ---- LOCAL API session (login once; reuse for query, upload, edit) --------
def load_creds():
    u, p = os.environ.get('SYNCY_USER'), os.environ.get('SYNCY_PASS')
    if u and p:
        return u, p
    path = os.path.expanduser('~/.syncy')
    if os.path.exists(path):
        d = {}
        for line in open(path, encoding='utf-8'):
            line = line.strip()
            if line and not line.startswith('#') and '=' in line:
                k, v = line.split('=', 1)
                d[k.strip().lower()] = v.strip()
        if d.get('user') and d.get('pass'):
            return d['user'], d['pass']
    sys.exit('No credentials: set SYNCY_USER/SYNCY_PASS or create ~/.syncy '
             '(user=..., pass=...).')


def make_api():
    cj = http.cookiejar.CookieJar()
    opener = urllib.request.build_opener(urllib.request.HTTPCookieProcessor(cj))

    def call(params, post=False):
        params = dict(params); params['format'] = 'json'
        if post:
            req = urllib.request.Request(
                LOCAL_API, data=urllib.parse.urlencode(params).encode('utf-8'),
                headers={'User-Agent': USER_AGENT})
        else:
            req = urllib.request.Request(
                LOCAL_API + '?' + urllib.parse.urlencode(params),
                headers={'User-Agent': USER_AGENT})
        with opener.open(req, timeout=30) as r:
            return json.loads(r.read().decode('utf-8'))
    return call, opener


def login(call):
    user, pwd = load_creds()
    lt = call({'action': 'query', 'meta': 'tokens', 'type': 'login'})
    lt = lt['query']['tokens']['logintoken']
    r = call({'action': 'login', 'lgname': user, 'lgpassword': pwd, 'lgtoken': lt}, post=True)
    if r.get('login', {}).get('result') != 'Success':
        sys.exit('LOGIN FAILED: ' + json.dumps(r.get('login', r)))


def csrf_token(call):
    return call({'action': 'query', 'meta': 'tokens'})['query']['tokens']['csrftoken']


def edit_page(call, csrf, title, text, summary):
    r = call({'action': 'edit', 'title': title, 'text': text,
              'summary': summary, 'token': csrf, 'assert': 'user'}, post=True)
    if 'error' in r:
        sys.exit('EDIT FAILED: ' + json.dumps(r['error']))
    return r['edit']


# ---- files: find references, see what's missing locally, transfer ---------
def canon(name):
    """A file name as MediaWiki's title normalization sees it: underscores,
    and the first letter uppercased (titles are ucfirst, so the gadget's
    lw_*.jpg answers locally to Lw_*.jpg)."""
    n = name.strip().replace(' ', '_')
    return (n[0].upper() + n[1:]) if n else n


def referenced_files(text):
    names, seen = [], set()
    for m in FILE_RE.finditer(text):
        n = m.group(1).strip().replace(' ', '_')
        if n and canon(n) not in seen:
            names.append(n)
            seen.add(canon(n))
    return names


def local_missing(call, names):
    """Read-only (no login): which of these File: names lack a file locally."""
    missing = []
    for i in range(0, len(names), 50):
        chunk = names[i:i + 50]
        r = call({'action': 'query', 'prop': 'imageinfo', 'iiprop': 'timestamp',
                  'titles': '|'.join('File:' + n for n in chunk), 'formatversion': '2'})
        present = set()
        for pg in r.get('query', {}).get('pages', []):
            if 'imageinfo' in pg:        # a file actually exists for this title
                present.add(canon(pg['title'].split(':', 1)[1]))
        for n in chunk:
            if canon(n) not in present:
                missing.append(n)
    return missing


def download_online_file(name):
    """Exact bytes of a file from the online wiki via Special:FilePath (which
    302-redirects to the real image; urllib carries our headers across it)."""
    url = ONLINE_BASE + 'Special:FilePath/' + urllib.parse.quote(name)
    headers = {'User-Agent': USER_AGENT}
    headers.update(EXTRA_HEADERS)
    req = urllib.request.Request(url, headers=headers)
    with urllib.request.urlopen(req, timeout=120) as r:
        return r.read()


def upload_local(opener, csrf, name, data, comment):
    boundary = '----syncy' + os.urandom(8).hex()

    def field(n, v):
        return ('--%s\r\nContent-Disposition: form-data; name="%s"\r\n\r\n%s\r\n'
                % (boundary, n, v)).encode('utf-8')

    head = b''.join([field('action', 'upload'), field('filename', name),
                     field('token', csrf), field('ignorewarnings', '1'),
                     field('comment', comment), field('format', 'json')])
    filehdr = ('--%s\r\nContent-Disposition: form-data; name="file"; filename="%s"\r\n'
               'Content-Type: application/octet-stream\r\n\r\n' % (boundary, name)).encode('utf-8')
    body = head + filehdr + data + b'\r\n' + ('--%s--\r\n' % boundary).encode('utf-8')
    req = urllib.request.Request(LOCAL_API, data=body, headers={
        'User-Agent': USER_AGENT,
        'Content-Type': 'multipart/form-data; boundary=%s' % boundary})
    with opener.open(req, timeout=300) as r:
        return json.loads(r.read().decode('utf-8'))


def transfer_files(opener, csrf, names):
    done, already, failed = [], [], []
    for n in names:
        try:
            data = download_online_file(n)
            r = upload_local(opener, csrf, n, data, 'syncy: import file from online')
            if r.get('upload', {}).get('result') in ('Success', 'Warning'):
                done.append(n)
            elif r.get('error', {}).get('code') == 'fileexists-no-change':
                already.append(n)      # byte-identical copy is here: a no-op, not a failure
            else:
                failed.append((n, json.dumps(r)[:200]))
        except Exception as e:
            failed.append((n, str(e)))
    return done, already, failed


def backup(title, text):
    os.makedirs(BACKUP_DIR, exist_ok=True)
    from datetime import datetime
    name = '%s.%s.txt' % (title.replace(' ', '_').replace('/', '_'),
                          datetime.now().strftime('%Y%m%d-%H%M%S'))
    path = os.path.join(BACKUP_DIR, name)
    with open(path, 'w', encoding='utf-8') as f:
        f.write(text)
    return path


def base_path(title):
    return os.path.join(BASE_DIR,
                        title.replace(' ', '_').replace('/', '_') + '.json.gz')


def read_base(title, raw=False):
    """The local and online texts as they stood when this page last synced
    cleanly. {} when there is none yet — a first run, or a run left with a
    question outstanding. The base holds the pages as they were, drafts and
    all; readers get them draft-free (V2.2) unless they ask for raw."""
    try:
        with gzip.open(base_path(title), 'rt', encoding='utf-8') as f:
            base = json.load(f)
    except Exception:
        return {}
    if not raw:
        for side in ('local', 'online'):
            base[side] = strip_drafts(base.get(side, ''))
    return base


def write_base(title, local_text, online_text):
    os.makedirs(BASE_DIR, exist_ok=True)
    tmp = base_path(title) + '.tmp'
    with gzip.open(tmp, 'wt', encoding='utf-8') as f:
        json.dump({'local': local_text, 'online': online_text,
                   'when': datetime.datetime.now().isoformat(timespec='seconds')},
                  f, ensure_ascii=False)
    os.replace(tmp, base_path(title))


def bodies(text):
    return {pid: inner for pid, inner, _s, _e in find_posts(text)}


def settle(title, clashes, both, local_text, online_text):
    """Three-way merge for the posts the last-edit marker cannot separate.

    The marker is written by the gadget and by nothing else, so a post edited
    at laussy.org in the ordinary editor carries the SAME marker as the local
    copy it no longer matches. Comparing the two copies to each other can
    therefore only ever say "same age, different words" — which is not an
    answer. Comparing each of them to what it was at the last sync is: the
    side that has moved is the side somebody wrote on.

    Returns (resolved, unresolved); `both` is updated in place for the ones
    that could be decided.
    """
    base = read_base(title)
    bloc, bonl = bodies(base.get('local', '')), bodies(base.get('online', ''))
    loc, onl = bodies(local_text), bodies(online_text)
    resolved, unresolved = {}, {}
    for pid, variants in clashes.items():
        l, o = loc.get(pid), onl.get(pid)
        if l is None or o is None or pid not in bloc or pid not in bonl:
            unresolved[pid] = variants          # no base to judge against
            continue
        moved_local = l.strip() != bloc[pid].strip()
        moved_online = o.strip() != bonl[pid].strip()
        if moved_online and not moved_local:
            both[pid] = (both[pid][0], o)
            resolved[pid] = 'online moved since the last sync, local did not'
        elif moved_local and not moved_online:
            both[pid] = (both[pid][0], l)
            resolved[pid] = 'local moved since the last sync, online did not'
        elif not moved_local and not moved_online:
            both[pid] = (both[pid][0], l)
            resolved[pid] = 'neither side moved; kept local'
        else:
            unresolved[pid] = variants
    return resolved, unresolved


def chrome_norm(s):
    """Chrome as it counts for comparison. Tombstones are rebuilt by the
    caller out of both texts, so their coming and going is bookkeeping, not
    something anybody wrote."""
    return TOMBSTONE.sub('', s or '').strip()


def settle_chrome(title, local_text, online_text):
    """Three-way merge for the chrome: everything on the page that is not a
    post — the {{MonthNav}} header above the first one, and whatever sits
    below the last.

    settle() asks of a post "which side moved since the last clean sync?".
    Chrome was never asked anything: local's copy was taken and online's
    dropped, so an edit made to the header at laussy.org died on the next
    sync — not merged, not backed up, not printed, gone. That is exactly the
    silence V2.0 was written to end, still in force one function away.

    It happened on 1 September 2026. Two chrome edits made online at 17:55
    and 17:58 — a link retargeted to a new {{anchor}}, then its ☕ label
    doubled to ☕☕ — were reverted by the 18:04 sync. The post he had also
    written online, in the same editor, minutes apart, came through intact:
    it had an id.

    The base answers here as it does for posts. Returns (preamble, postamble,
    resolved, unresolved): `resolved` maps a part name to why, `unresolved`
    maps it to (local, online) for the caller to write out and refuse over.
    """
    base = read_base(title)
    lpre, lpost = split_chrome(local_text)
    opre, opost = split_chrome(online_text)
    blpre, blpost = split_chrome(base.get('local', ''))
    bopre, bopost = split_chrome(base.get('online', ''))
    chosen, resolved, unresolved = {}, {}, {}
    for name, l, o, bl, bo in (('the header above the posts', lpre, opre,
                                blpre, bopre),
                               ('the text below the posts', lpost, opost,
                                blpost, bopost)):
        if chrome_norm(l) == chrome_norm(o):
            chosen[name] = l
            continue
        if not base:
            # No base is no answer, and the old default (take local) is the
            # very thing that loses the words. Ask instead.
            chosen[name] = l
            unresolved[name] = (l, o)
            continue
        moved_local = chrome_norm(l) != chrome_norm(bl)
        moved_online = chrome_norm(o) != chrome_norm(bo)
        if moved_online and not moved_local:
            chosen[name] = o
            resolved[name] = 'online moved since the last sync, local did not'
        elif moved_local and not moved_online:
            chosen[name] = l
            resolved[name] = 'local moved since the last sync, online did not'
        elif not moved_local and not moved_online:
            # They differed at the last sync too and neither has been touched
            # since: local is the base of record, and the next push makes it
            # true online anyway.
            chosen[name] = l
            resolved[name] = 'neither side moved; kept local'
        else:
            chosen[name] = l
            unresolved[name] = (l, o)
    return (chosen['the header above the posts'],
            chosen['the text below the posts'], resolved, unresolved)


def save_clash(title, pid, variants):
    """Write the copies of one clashing post side by side, so that a text
    syncy declined to publish is never a text syncy lost."""
    os.makedirs(BACKUP_DIR, exist_ok=True)
    from datetime import datetime
    name = '%s.%s.CLASH-%s.txt' % (title.replace(' ', '_').replace('/', '_'),
                                   datetime.now().strftime('%Y%m%d-%H%M%S'), pid)
    path = os.path.join(BACKUP_DIR, name)
    with open(path, 'w', encoding='utf-8') as f:
        f.write('%s%d copies with the same last-edit time.\n'
                'The FIRST is the one kept (local is the base of record); the\n'
                'rest were declined. Merge by hand and edit the post normally.\n'
                % (pid, len(variants)))
        for i, v in enumerate(variants):
            f.write('\n==== copy %d %s ====\n%s\n'
                    % (i + 1, '(kept)' if i == 0 else '(declined)', v.strip()))
    return path


# ---- page selection -------------------------------------------------------
def page_for(year, month, prefix=''):
    return '%s%s (%d)' % (prefix, MONTHS[month - 1], year)


def current_pages():
    """The current month of every branch (V1.5: 𝕐 and ℤ)."""
    t = datetime.date.today()
    return [page_for(t.year, t.month, pre) for pre, _y, _m in BRANCHES]


def all_pages():
    """Every month of every branch, from that branch's start to the current
    month, inclusive."""
    t = datetime.date.today()
    pages = []
    for pre, y, m in BRANCHES:
        while (y, m) <= (t.year, t.month):
            pages.append(page_for(y, m, pre))
            m += 1
            if m > 12:
                m, y = 1, y + 1
    return pages


# ---- one page -------------------------------------------------------------
class Api:
    """Holds the cookie session; logs in lazily and caches the CSRF token,
    so a --all run authenticates once and reuses it for every month."""
    def __init__(self):
        self.call, self.opener = make_api()
        self.csrf = None

    def auth(self):
        if self.csrf is None:
            login(self.call)
            self.csrf = csrf_token(self.call)
        return self.csrf


def sync_page(api, title):
    try:
        local_text = fetch(LOCAL_BASE, title)
    except Exception as e:
        raise RuntimeError('reading local failed: %s' % e)
    try:
        online_text = fetch(ONLINE_BASE, title)
    except Exception as e:
        raise RuntimeError('reading online failed: %s '
                           '(Cloudflare? check EXTRA_HEADERS)' % e)

    # Scheduled posts (V2.2) are lifted out whole before anything reads the
    # page and put back once it is rebuilt. The raw texts are what the
    # backup and the base record: the pages as they stand.
    local_raw, online_raw = local_text, online_text
    for side, raw in (('local', local_raw), ('online', online_raw)):
        crowded = crowded_drafts(raw)
        if crowded:
            print('\u2716 %s: REFUSING to write — a <!-- --> block on the %s '
                  'page holds %d posts (%s): a scheduled post is wrapped '
                  'alone. Fix the page by hand, then sync again.'
                  % (title, side, len(crowded[0]), ', '.join(crowded[0])))
            return 'refused'
    local_text, ldrafts = lift_drafts(local_raw)
    online_text, odrafts = lift_drafts(online_raw)

    both, clashes = collect([local_text, online_text])
    local_only, _ = collect([local_text])
    # Deletion beats every copy, and it travels BOTH ways (V1.7). This merge
    # is a UNION of the two texts, so a post deleted on one side is simply
    # handed back by the other unless something says it was deleted on
    # purpose. Online deletions have said so since V1.6; local ones said
    # nothing, on the theory that the push propagates them — true only if the
    # push happens first. Sync first instead and the post walked straight back
    # in, was pushed back online, and came back for ever after. So a tombstone
    # from EITHER text buries the post.
    dead = (set(TOMBSTONE.findall(online_text))
            | set(TOMBSTONE.findall(local_text)))
    killed = sorted(pid for pid in dead if pid in both)
    for pid in dead:
        both.pop(pid, None)
        clashes.pop(pid, None)
    # Decide the clashes BEFORE the page is rebuilt, so a post the base can
    # settle goes into the merge like any other (V2.0).
    resolved, unresolved = settle(title, clashes, both, local_text, online_text)
    for pid in resolved:
        clashes.pop(pid, None)

    lpre0, lpost0 = split_chrome(local_text)
    preamble, postamble, chrome_ok, chrome_stuck = settle_chrome(
        title, local_text, online_text)
    # Chrome taken from online is a change to the local page even though no
    # post moved — without this the merge is built and never written (V2.1).
    chromechanged = (chrome_norm(preamble) != chrome_norm(lpre0)
                     or chrome_norm(postamble) != chrome_norm(lpost0))
    # A tombstone guards against exactly one thing: the other side still
    # holding the post. Once online no longer has it there is nothing left to
    # guard, and the marker goes — so a month page does not silt up with the
    # record of every post ever deleted from it. Until then it stays, which is
    # what stops the next sync (or the next push) from digging the post up.
    online_ids = {pid for pid, _inner, _s, _e in find_posts(online_text)}
    stones = sorted(dead & online_ids)
    preamble = TOMBSTONE.sub('', preamble)
    postamble = TOMBSTONE.sub('', postamble).strip('\n')
    if stones:
        postamble = (postamble + '\n' if postamble else '') + '\n'.join(
            '<!--deleted:%s-->' % pid for pid in stones)
    drafts, notes = keep_drafts(
        ldrafts, odrafts, both, dead,
        lift_drafts(read_base(title, raw=True).get('online', ''))[1])
    for note in notes:
        print('  \u00b7 %s: %s' % (title, note))
    merged = place_drafts(with_chrome(preamble, build(both, clashes),
                                      postamble), drafts)

    # A tombstone that has expired is a real change to the local page even
    # when no post moved, or it would never actually be cleared away. So is
    # a draft pulled in or dropped (V2.2).
    textchanged = (normalize(both) != normalize(local_only)
                   or set(TOMBSTONE.findall(local_text)) != set(stones)
                   or chromechanged
                   or set(drafts) != set(ldrafts))
    new_ids = sorted(set(both) - set(local_only))
    upd_ids = sorted(pid for pid in both if pid in local_only
                     and both[pid][1].strip() != local_only[pid][1].strip())

    refs = referenced_files(merged)
    missing = local_missing(api.call, refs)       # read-only, no login required

    # A clash is a question, not a change: two copies of one post carrying the
    # same last-edit time and different words. The kept copy is local's, and
    # every declined one is written out where it can be read, named and loudly
    # reported — on stdout, since a warning on stderr inside an llw2lw run is
    # a warning nobody sees.
    # Nothing between the posts may be thrown away (V1.8). The rebuild can
    # only put back what it could parse, so if the local page carries text
    # there that is not a post, writing this page would destroy it — refuse,
    # and say exactly what is at stake. Read from the LOCAL text, which is the
    # one about to be overwritten. Fix the page (usually by giving the orphan
    # text its <blockquote id="…"> back) and the sync goes through.
    orphan = orphan_text(local_text)
    if orphan:
        preview = orphan if len(orphan) <= 300 else orphan[:300] + '…'
        print('\u2716 %s: REFUSING to write — %d character(s) sit between the '
              'posts without being in one, and the rebuild would drop them:'
              % (title, len(orphan)))
        for line in preview.split('\n'):
            print('      | %s' % line)
        print('    Put that text back inside a <blockquote id="…"> (or move it '
              'above the first post / below the last), then sync again.')
        return 'refused'

    for pid in sorted(resolved):
        print('\u2699 %s: %s settled — %s' % (title, pid, resolved[pid]))
    for name in sorted(chrome_ok):
        if chrome_ok[name].startswith('neither'):
            continue                   # nothing moved; do not narrate it
        print('\u2699 %s: %s settled — %s' % (title, name, chrome_ok[name]))

    # What the base could not settle is a question for Fabrice, and a question
    # must stop the machinery rather than be printed past. The marker line is
    # what llw2lw 3.9.0 reads to refuse the push of this page.
    clash_files = [save_clash(title, pid, variants)
                   for pid, variants in sorted(unresolved.items())]
    if unresolved:
        print('\n\u2716 %s: %d POST(S) CHANGED ON BOTH WIKIS SINCE THE LAST '
              'SYNC — nobody but you can say which words to keep.'
              % (title, len(unresolved)))
    for pid, path in zip(sorted(unresolved), clash_files):
        print('    %s: kept the local copy for now, the other is in %s'
              % (pid, path))
        print('SYNCY-UNRESOLVED: %s %s %s' % (title, pid, path))
    if unresolved:
        print('    Merge each by hand, edit the post normally, and sync again. '
              'Until then this page is not pushed and the base is not moved.\n')

    # Chrome has no id to be named by, so it is named by where it sits; the
    # rest is the posts' machinery exactly (V2.1).
    if chrome_stuck:
        print('\n\u2716 %s: CHROME CHANGED ON BOTH WIKIS SINCE THE LAST SYNC '
              '— nobody but you can say which words to keep.' % title)
    for name in sorted(chrome_stuck):
        path = save_clash(title, name, list(chrome_stuck[name]))
        print('    %s: kept the local copy for now, the other is in %s'
              % (name, path))
        print('SYNCY-UNRESOLVED: %s %s %s'
              % (title, name.replace(' ', '_'), path))
    if chrome_stuck:
        print('    Merge by hand, edit the page normally, and sync again. '
              'Until then this page is not pushed and the base is not moved.\n')

    if not textchanged and not missing:
        print('\u00b7 %s: up to date' % title)
        if not unresolved and not chrome_stuck:
            write_base(title, local_raw, online_raw)
        return 'unresolved' if (unresolved or chrome_stuck) else 'ok'

    csrf = api.auth()                             # uploads/edits need auth
    done, already, failed = transfer_files(api.opener, csrf, missing) \
        if missing else ([], [], [])

    rev = bpath = None
    if textchanged:
        bpath = backup(title, local_raw)
        rev = edit_page(api.call, csrf, title, merged,
                        'syncy: import online posts').get('newrevid')

    bits = ['%d new, %d updated' % (len(new_ids), len(upd_ids))] if textchanged \
        else ['no text changes']
    if killed:
        bits.insert(1 if textchanged else 0,
                    '%d deleted online (%s)' % (len(killed), ', '.join(killed)))
    bits.append('%d files transferred, %d present, %d failed'
                % (len(done), len(refs) - len(missing) + len(already), len(failed)))
    print('\u2713 %s: %s' % (title, '; '.join(bits)))
    if textchanged:
        print('    local rev %s; backup %s' % (rev, bpath))
    for n, why in failed:
        print('    FAILED %s: %s' % (n, why))

    # The base moves only when the page is settled. Leaving it where it is
    # while a question is open is what keeps the question being asked: move it
    # now and the next run would find neither side changed, hand the post to
    # local, and the declined words would be gone for good.
    if not unresolved and not chrome_stuck:
        write_base(title, merged if textchanged else local_raw, online_raw)
    return 'unresolved' if (unresolved or chrome_stuck) else 'ok'


# ---- a page that only grows (V2.3) ----------------------------------------
GALLERY_RE = re.compile(r'<gallery\b[^>]*>(.*?)</gallery\s*>', re.S | re.I)


def gallery_files(text):
    """The files a page's <gallery> blocks name: one a line, the name up to
    the first '|', with or without its File: prefix."""
    names, seen = [], set()
    for block in GALLERY_RE.findall(text):
        for line in block.splitlines():
            n = line.split('|', 1)[0].strip()
            n = re.sub(r'^(?:File|Image)\s*:\s*', '', n, flags=re.I).replace(' ', '_')
            if n and canon(n) not in seen:
                names.append(n)
                seen.add(canon(n))
    return names


SECTION_RE = re.compile(r'^==[^=].*==\s*$')
AF_HEAD_RE = re.compile(r'\{\{\s*af\s*\|\s*([0-9A-Za-z]{5,6})\s*\|')


def split_sections(text):
    """(intro, [(heading line, body)]) — the page cut at its == headings =="""
    intro, sections, cur = [], [], None
    for line in text.rstrip().split('\n') if text.strip() else []:
        if SECTION_RE.match(line):
            cur = [line.rstrip(), []]
            sections.append(cur)
        elif cur is None:
            intro.append(line)
        else:
            cur[1].append(line)
    return ('\n'.join(intro).strip('\n'),
            [(h, '\n'.join(b).strip('\n')) for h, b in sections])


def section_moment(heading):
    """When a section's {{af|TAG|h}} heading says it is: (day, instant)."""
    m = AF_HEAD_RE.search(heading)
    r = read_id('Y-' + m.group(1)) if m else None
    return (r[0].date(), r[1]) if r else None


def blocks_of(body):
    return [b.strip('\n') for b in re.split(r'\n\s*\n', body) if b.strip()]


def three_way(local, base, online):
    """One piece of text on three sides -> (what to keep, both moved?)."""
    if local == online or online == base:
        return local, False
    if local == base:
        return online, False
    return None, True


def merge_union(local_text, base_text, online_text):
    """Three-way merge of a page that only grows, by what it is made of —
    day sections, and in them blocks set apart by a blank line — so that a
    <gallery> is never cut in two, which a merge by lines does as soon as
    both wikis file on the same day (the closing tags look alike, and diff3
    pairs them off). Nothing is a conflict: a piece one side changed since
    the base is taken from that side, a section only one side has is kept
    unless the base had it (then the other side deleted it), and where both
    sides wrote in the same section local's blocks stand and online's new
    ones follow them. Newest day first, one section a day: an online section
    for a day local already has is poured into local's."""
    # Whole pages first: a page only one side touched is that side's, to the
    # letter — nothing is re-laid-out that did not have to be.
    whole, both = three_way(local_text.rstrip(), base_text.rstrip(),
                            online_text.rstrip())
    if not both:
        return whole + '\n'
    l_intro, l_secs = split_sections(local_text)
    b_intro, b_secs = split_sections(base_text)
    o_intro, o_secs = split_sections(online_text)
    intro, both = three_way(l_intro, b_intro, o_intro)
    if both:
        intro = l_intro
    lmap, bmap, omap = dict(l_secs), dict(b_secs), dict(o_secs)

    out = []                                  # [heading, body], local's order
    for h, body in l_secs:
        if h in omap:
            kept, both = three_way(body, bmap.get(h), omap[h])
            if both:
                mine = blocks_of(body)
                was = set(blocks_of(bmap.get(h, '')))
                kept = '\n\n'.join(mine + [b for b in blocks_of(omap[h])
                                            if b not in mine and b not in was])
            out.append([h, kept])
        elif h in bmap and bmap[h] == body:
            continue                          # deleted online, untouched here
        else:
            out.append([h, body])
    for h, body in o_secs:
        if h in lmap:
            continue
        if h in bmap:
            if bmap[h] == body:
                continue                      # deleted here, untouched online
        when = section_moment(h)
        # The same day under another tag: one section a day.
        twin = next((s for s in out if when and section_moment(s[0])
                     and section_moment(s[0])[0] == when[0]), None)
        if twin:
            mine = blocks_of(twin[1])
            twin[1] = '\n\n'.join(mine + [b for b in blocks_of(body)
                                           if b not in mine])
            continue
        # Newest first: ahead of the first section that is older than it.
        at = len(out)
        if when:
            for i, s in enumerate(out):
                sm = section_moment(s[0])
                if sm and sm[1] < when[1]:
                    at = i
                    break
        out.insert(at, [h, body])
    parts = ([intro] if intro else []) + ['%s\n%s' % (h, b) if b else h
                                          for h, b in out]
    return '\n\n'.join(parts) + '\n'


def same_text(a, b):
    return a.rstrip() == b.rstrip()


def sync_union(api, title):
    try:
        local_text = fetch(LOCAL_BASE, title)
    except Exception as e:
        raise RuntimeError('reading local failed: %s' % e)
    try:
        online_text = fetch(ONLINE_BASE, title)
    except Exception as e:
        raise RuntimeError('reading online failed: %s '
                           '(Cloudflare? check EXTRA_HEADERS)' % e)
    if not online_text.strip():
        print(%s: nothing online to pull' % title)
        return 'ok'
    base_text = read_base(title, raw=True).get('online', '')
    merged = merge_union(local_text, base_text, online_text)
    textchanged = not same_text(merged, local_text)

    refs = list(dict.fromkeys(referenced_files(merged) + gallery_files(merged)))
    missing = local_missing(api.call, refs) if refs else []

    if not textchanged and not missing:
        print(%s: up to date' % title)
        if same_text(local_text, online_text):
            write_base(title, local_text, online_text)
        return 'ok'

    csrf = api.auth()
    done, already, failed = transfer_files(api.opener, csrf, missing) \
        if missing else ([], [], [])
    rev = bpath = None
    if textchanged:
        bpath = backup(title, local_text)
        rev = edit_page(api.call, csrf, title, merged,
                        'syncy: import what was filed online').get('newrevid')
    added = len([l for l in merged.splitlines() if l.strip()]) \
        - len([l for l in local_text.splitlines() if l.strip()])
    print('✓ %s: %s; %d files transferred, %d present, %d failed'
          % (title, ('%+d line(s) from online' % added) if textchanged
             else 'no text changes', len(done),
             len(refs) - len(missing) + len(already), len(failed)))
    if textchanged:
        print('    local rev %s; backup %s' % (rev, bpath))
    for n, why in failed:
        print('    FAILED %s: %s' % (n, why))
    # The base is what BOTH wikis hold. After a pull that is true only if
    # online had nothing to learn from local; otherwise llw2lw moves it, once
    # the push that makes it true is verified.
    if same_text(merged, online_text):
        write_base(title, merged, online_text)
    return 'ok'


# ---- main -----------------------------------------------------------------
def seed_base(title):
    """Declare the two wikis, as they stand right now, to be the reference a
    later clash is judged against.

    Needed once, when the base store is introduced or after a page has been
    repaired by hand: without a base every clash is unresolvable, and the run
    would stop on a difference that is in fact already settled.
    """
    write_base(title, fetch(LOCAL_BASE, title), fetch(ONLINE_BASE, title))
    print('\u2261 %s: base set to the two wikis as they stand now' % title)


def main():
    do_all, seed, titles = False, False, []
    for a in sys.argv[1:]:
        if a in ('-a', '--all'):
            do_all = True
        elif a == '--seed-base':
            seed = True
        elif a in ('-h', '--help'):
            sys.stdout.write(__doc__)
            return
        elif a.startswith('-'):
            sys.exit('unknown option: %s' % a)
        else:
            titles.append(a)

    if do_all:
        pages = all_pages() + UNION_PAGES
    elif titles:
        pages = titles
    else:
        pages = current_pages() + UNION_PAGES

    if seed:
        for title in pages:
            try:
                seed_base(title)
            except Exception as e:
                sys.stderr.write('  ! %s: %s\n' % (title, e))
        return

    api = Api()
    multi = do_all or len(pages) > 1
    stuck = []
    for title in pages:
        try:
            sync = sync_union if title in UNION_PAGES else sync_page
            if sync(api, title) in ('unresolved', 'refused'):
                stuck.append(title)
        except Exception as e:
            if multi:
                sys.stderr.write('  ! %s: %s (skipped)\n' % (title, e))
            else:
                sys.exit('ERROR with "%s": %s' % (title, e))
    # Exit 2 says "I could not settle everything". llw2lw reads it, and so
    # does anyone running syncy from a script: a clash left open must not look
    # like a clean run (V2.0).
    if stuck:
        print('Unresolved on: %s. Nothing on those pages should be pushed '
              'until they are merged by hand.' % ', '.join(stuck))
        sys.exit(2)


if __name__ == '__main__':
    main()