feat: discussion artifacts as wiki pages, in two new layers

A discussion leaves behind a directory of markdown somewhere outside this
repo, and the only durable home for it is the Gitea wiki. Getting it
there by hand means re-deriving the same three things every time: what
each file should be called, where it goes, and whether the page already
exists. Two skills wrap that, along the split the repo already uses.

`skills/page` is domain, offline, stdlib-only, and knows nothing about
Gitea. It imports a directory into a space under `tmp/wiki/`, titles
every file, records the result in `.pages.json`, and writes the index.
`skills/wiki` is the bridge — `wikimap.py` translates, and the transport
is `_gitea.py`, the same one the issue side uses. There is no second
transport, and `tea` has no wiki subcommand to offer one.

The Gitea wiki is flat, and that fact shapes everything

There are no directories. A title of `a/b` is stored as one file named
`a%2Fb.md`, and Gitea escapes it by rules of its own: space becomes `-`,
`/` becomes `%2F`, and a literal `-` forces a trailing `.-` marker so the
two stay distinct. `Chain decisions — DC` under two levels of prefix
comes back as `Simple-Chains%2FParked%2FChain-decisions-%E2%80%94-DC`.

So `sub_url` is the identity, it is read back from whatever the API
returned, and it is never constructed. One built by hand that is almost
right does not fail — it creates a second page and abandons the first.

And a real subdirectory committed into a wiki's git repository is a ghost:
the file exists, the API and the web UI do not see it. `folder/page.md`
in this repo's own wiki is one. Nothing here clones a wiki repo.

A title is a decision, not a derivation

Titles come from the first heading, because there is no mechanical route
from `03-q-01-do-we-know-the-chain-participant-by-name.md` to
`Q-01. Do We Know the Chain Participant by Name`. But they are derived
exactly once. A re-import replaces bodies and keeps titles, so editing a
heading cannot rename a published page — which would not rename it, it
would publish a second one.

`--retitle` opts in. It finds the prior entry by `source` rather than by
path, because the path is derived from the title and a retitle moves it;
looked up by path the page would read as new and the next push would
duplicate it. The old file goes, `sub_url` comes along, and `pushed` is
cleared — a rename can leave the body byte-identical, and push decides by
body hash alone, so a stale hash would skip the rename forever.

Ordering is a `NN-` file-name prefix and never reaches the title. `00-`
means "this is the directory's own page", and that page is named for the
directory, not for its own heading: a child's title has to extend its
parent's exactly, and `ideas/00-intro.md` opens with "Ideas for chain
business requirements".

Path collisions are reported and never resolved. Picking a winner is how
a discussion loses a document.

The index is navigation, not decoration

Nothing draws a tree from flat titles. `page_index.py` writes one as an
ordinary page, nested by title depth rather than by manifest path order —
those disagree, since on disk `Top/System.md` sorts before
`Top/Ideas/Scale.md` while in the hierarchy System is a child and Scale a
grandchild. A parent with no page of its own still gets a node, so its
children are not hidden.

Links use `sub_url` when there is one and Gitea's `[[Title|label]]`
syntax when there is not, so the order is push, rebuild, push.

The same stances as the issue store, for the same reasons

Pull overwrites, push is additive and never deletes, change detection is
one hash and there is no drift model. A page with no `sub_url` has never
been published, and that is a durable state.

Issues gain a `wiki:` field holding page titles — titles, not URLs, so
the reference stays in the domain. It already round-trips as a foreign
key; this documents it.

Verified against a live Gitea 1.26.1: create, update with a message,
unchanged-skip, prefix-filtered pull, byte-identical round trip, and the
per-page revision history carrying the operator's own words. The probe
pages were deleted afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
naudachu
2026-08-10 16:29:55 +05:00
parent 257c547e22
commit 1815d91cdf
16 changed files with 2319 additions and 26 deletions
+77
View File
@@ -0,0 +1,77 @@
#!/usr/bin/env python3
"""
wiki_ls.py — list what is actually in a wiki. One call, no bodies.
Cheap enough to run before a pull: it tells you what titles exist, which is the
only thing a prefix filter can be built from, and it shows the `sub_url` Gitea
settled on for each — worth a look the first time a title contains a dash or a
slash, because the escaping is not what anyone guesses.
wiki_ls.py
wiki_ls.py --prefix "Simple Chains"
wiki_ls.py --repo other/repo --titles
Usage:
wiki_ls.py [--repo owner/repo] [--prefix TITLE] [--titles] [--urls]
"""
import argparse
import os
import sys
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
sys.path.insert(0, os.path.join(_HERE, "..", "..", "sync", "scripts"))
import _gitea # noqa: E402
import wikimap # noqa: E402
def cell(v):
return (str(v or "").strip().replace("|", "\\|")) or ""
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--repo", help="owner/repo (default: the repo in CWD)")
ap.add_argument("--prefix", default="", help="only titles at or under this")
ap.add_argument("--titles", action="store_true",
help="print one title per line and nothing else")
ap.add_argument("--urls", action="store_true", help="add the browser URL")
a = ap.parse_args()
login = _gitea.require_login()
base = _gitea.repo_base(a.repo)
slug = _gitea.repo_slug(login, a.repo)
listing = _gitea.paginate(login, "%s/wiki/pages" % base)
if not isinstance(listing, list):
_gitea.die("unexpected listing from %s/wiki/pages" % base)
rows = sorted((p for p in listing
if wikimap.matches_prefix(p.get("title") or "", a.prefix)),
key=lambda p: p.get("title") or "")
if not rows:
print("no page at or under %r in %s" % (a.prefix, slug) if a.prefix
else "%s has no wiki pages" % slug)
return 0
if a.titles:
for p in rows:
print(p.get("title") or "")
return 0
head = ["title", "sub_url", "updated", "by"] + (["url"] if a.urls else [])
print("| %s |" % " | ".join(head))
print("|%s|" % "|".join("---" for _ in head))
for p in rows:
c = (p.get("last_commit") or {}).get("author") or {}
row = [cell(p.get("title")), "`%s`" % cell(p.get("sub_url")),
cell((c.get("date") or "")[:10]), cell(c.get("name"))]
if a.urls:
row.append(cell(p.get("html_url")))
print("| %s |" % " | ".join(row))
print("\n%d page(s) in %s" % (len(rows), slug))
return 0
if __name__ == "__main__":
sys.exit(main())
+127
View File
@@ -0,0 +1,127 @@
#!/usr/bin/env python3
"""
wiki_pull.py — fetch wiki pages into a local space.
wiki_pull.py the whole wiki
wiki_pull.py --prefix "Simple Chains" a page and everything under it
wiki_pull.py --repo other/repo --space docs from elsewhere, into a named space
The wiki is flat, so "a page and its children" is a prefix test on the title,
run against one listing call. One GET per page follows. There is no tree
endpoint to ask for a subtree, and no way to fetch bodies in bulk.
Pulling OVERWRITES the local body — a fetch, not a merge, the same stance the
issue store takes. `sha` and `synced` tell you how old your copy is; re-pull
when it matters. Nothing tracks drift.
Usage:
wiki_pull.py [--repo owner/repo] [--prefix TITLE] [--space SPACE]
[--dry-run] [--out DIR]
"""
import argparse
import os
import sys
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
sys.path.insert(0, os.path.join(_HERE, "..", "..", "sync", "scripts"))
sys.path.insert(0, os.path.join(_HERE, "..", "..", "page", "scripts"))
import _gitea # noqa: E402
import page # noqa: E402
import wikimap # noqa: E402
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--repo", help="owner/repo (default: the repo in CWD)")
ap.add_argument("--prefix", default="", help="only titles at or under this")
ap.add_argument("--space", help="local space (default: the owner/repo slug)")
ap.add_argument("--dry-run", action="store_true")
ap.add_argument("--out", help="wiki cache root (default: <repo>/tmp/wiki)")
a = ap.parse_args()
login = _gitea.require_login()
base = _gitea.repo_base(a.repo)
slug = _gitea.repo_slug(login, a.repo)
space = a.space or slug
root = a.out or page.WIKI_ROOT
space_dir = page.space_root(space, root)
listing = _gitea.paginate(login, "%s/wiki/pages" % base)
if not isinstance(listing, list):
_gitea.die("unexpected listing from %s/wiki/pages" % base)
wanted = [p for p in listing
if wikimap.matches_prefix(p.get("title") or "", a.prefix)]
if not wanted:
if a.prefix:
_gitea.die("no page at or under %r in %s (%d page(s) in the wiki)"
% (a.prefix, slug, len(listing)))
_gitea.die("%s has no wiki pages" % slug)
manifest = page.load_manifest(space, root)
# Two remote titles can land on one local path — Gitea keeps them apart with
# its `.-` marker, a filesystem does not. Caught before anything is written,
# because the failure mode otherwise is one page silently overwriting
# another and the manifest pointing both entries at the survivor.
seen = {}
for p in wanted:
seen.setdefault(page.path_for_title(p["title"]), []).append(p["title"])
clashes = {k: v for k, v in seen.items() if len(v) > 1}
for path, titles in sorted(clashes.items()):
sys.stderr.write("collision: %s <- %s\n" % (path, " | ".join(titles)))
created = not os.path.isdir(space_dir)
n = 0
for p in sorted(wanted, key=lambda x: x.get("title") or ""):
title = p["title"]
relpath = page.path_for_title(title)
if relpath in clashes:
continue
if a.dry_run:
print("%-44s %s" % (relpath, title))
n += 1
continue
full = _gitea.api(login, wikimap.page_endpoint(base, p["sub_url"]))
if not isinstance(full, dict):
_gitea.warn("could not read %r; skipped" % title)
continue
text = wikimap.decode(full)
dest = os.path.join(space_dir, relpath)
os.makedirs(os.path.dirname(dest), exist_ok=True)
with open(dest, "w", encoding="utf-8") as f:
f.write(text)
# The prior entry is the base so `order` — a local decision the wiki
# cannot hold — survives a pull.
e = dict(manifest["pages"].get(relpath, {}))
e.update(wikimap.from_payload(full))
e["synced"] = _gitea.now_iso()
# What is on disk is now exactly what is published, so push has nothing
# to do until the file is edited.
e["pushed"] = page.body_hash(text)
manifest["pages"][relpath] = e
print("%-44s %s" % (relpath, title))
n += 1
if a.dry_run:
print("\ndry run — %d page(s) would be written to %s" % (n, space_dir))
return 1 if clashes else 0
page.save_manifest(manifest, root)
if created:
sys.stderr.write("created space %s\n" % space_dir)
print("\n%d page(s) from %s -> %s" % (n, slug, space_dir))
if clashes:
sys.stderr.write("%d collision(s) skipped — rename them in the wiki\n"
% len(clashes))
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env python3
"""
wiki_push.py — publish a local space to a wiki.
wiki_push.py -m "Import the simple-chains discussion"
wiki_push.py --prefix "Simple Chains/Ideas" -m "Rework B-04"
wiki_push.py -m "Fix the send-timing table" Simple-Chains/Ideas/Send-timing.md
Every page in the selection is compared against `pushed` — the hash of what was
last published — and only the ones that differ are sent. That is the whole of
change detection: no timestamps, no drift model.
A page with no `sub_url` is created; a page with one is edited in place. The
title comes from the manifest, never re-derived from the file, because sending
a different title to the edit endpoint is a RENAME and leaves nothing behind at
the old address.
Pushing is additive. A page deleted locally is NOT deleted in the wiki — the
manifest simply stops mentioning it. Removing a published page is an explicit
act; do it in the web UI or with a DELETE through /tea:use.
Usage:
wiki_push.py -m MESSAGE [--space SPACE] [--repo owner/repo]
[--prefix TITLE] [--dry-run] [--out DIR] [PATH-or-TITLE ...]
"""
import argparse
import os
import sys
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
sys.path.insert(0, os.path.join(_HERE, "..", "..", "sync", "scripts"))
sys.path.insert(0, os.path.join(_HERE, "..", "..", "page", "scripts"))
import _gitea # noqa: E402
import page # noqa: E402
import wikimap # noqa: E402
def select(manifest, prefix, targets):
"""The pages to consider, in tree order.
A positional argument matches a manifest path or a title, exactly. Exact
because a near-miss that silently selects nothing is indistinguishable from
a clean no-op run, and the operator finds out only when the page never
appears."""
pages = (page.children_of(manifest, prefix) if prefix
else page.sorted_pages(manifest))
if not targets:
return pages, []
want, chosen, hit = set(targets), [], set()
for p, e in pages:
if p in want or e.get("title") in want:
chosen.append((p, e))
hit.add(p if p in want else e.get("title"))
return chosen, sorted(want - hit)
def main():
ap = argparse.ArgumentParser()
ap.add_argument("targets", nargs="*", metavar="PATH-or-TITLE")
ap.add_argument("-m", "--message", required=True,
help="wiki commit message for this push")
ap.add_argument("--repo", help="owner/repo (default: the repo in CWD)")
ap.add_argument("--space", help="local space (default: the owner/repo slug)")
ap.add_argument("--prefix", default="", help="only titles at or under this")
ap.add_argument("--dry-run", action="store_true")
ap.add_argument("--out", help="wiki cache root (default: <repo>/tmp/wiki)")
a = ap.parse_args()
login = _gitea.require_login()
base = _gitea.repo_base(a.repo)
slug = _gitea.repo_slug(login, a.repo)
space = a.space or slug
root = a.out or page.WIKI_ROOT
space_dir = page.space_root(space, root)
if not os.path.isdir(space_dir):
_gitea.die("no such space: %s (looked in %s). Import or pull first."
% (space, space_dir))
manifest = page.load_manifest(space, root)
if not manifest["pages"]:
_gitea.die("space %s has no pages in its manifest" % space)
chosen, missing = select(manifest, a.prefix, a.targets)
for t in missing:
_gitea.warn("not in the manifest: %s" % t)
if not chosen:
_gitea.die("nothing selected")
created = updated = skipped = 0
for relpath, e in chosen:
title = e.get("title")
full = os.path.join(space_dir, relpath)
if not title:
_gitea.warn("%s has no title in the manifest; skipped" % relpath)
continue
if not os.path.isfile(full):
_gitea.warn("%s is in the manifest but not on disk; skipped" % relpath)
continue
with open(full, encoding="utf-8") as f:
text = f.read()
h = page.body_hash(text)
if e.get("sub_url") and h == e.get("pushed"):
skipped += 1
continue
verb = "create" if not e.get("sub_url") else "update"
print("%-7s %-44s %s" % (verb, relpath, title))
if a.dry_run:
created += verb == "create"
updated += verb == "update"
continue
if verb == "create":
payload = wikimap.new_payload(title, text, a.message)
got = _gitea.api(login, "%s/wiki/new" % base, method="POST",
payload=payload, payload_name="wiki-new",
out_root=space_dir)
else:
payload = wikimap.edit_payload(title, text, a.message)
got = _gitea.api(login, wikimap.page_endpoint(base, e["sub_url"]),
method="PATCH", payload=payload,
payload_name="wiki-edit", out_root=space_dir)
if not isinstance(got, dict) or not got.get("sub_url"):
_gitea.warn("%s: no page returned; the manifest is unchanged for it"
% title)
continue
# sub_url comes back from Gitea and is stored as given. It is the only
# address this page has, and it is not something we could have computed.
e.update(wikimap.from_payload(got))
e["synced"] = _gitea.now_iso()
e["pushed"] = h
manifest["pages"][relpath] = e
created += verb == "create"
updated += verb == "update"
if a.dry_run:
print("\ndry run — %d to create, %d to update, %d unchanged"
% (created, updated, skipped))
return 0
page.save_manifest(manifest, root)
print("\n%d created, %d updated, %d unchanged -> %s wiki"
% (created, updated, skipped, slug))
return 0
if __name__ == "__main__":
sys.exit(main())
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env python3
"""
wikimap.py — md <-> Gitea wiki JSON. The whole translation, and only the
translation.
Pure functions: no network, no filesystem, no argparse. Give it a payload and
it hands back a page; give it a page and it hands back a request body. That
purity is the point — it can be reasoned about and tested without a Gitea
anywhere, and it is the single file to open when the two representations
disagree.
Direction of knowledge: this module imports the domain (page.py) and is
imported by the transport's callers. The domain never imports this.
What crosses the boundary, and what does not:
domain Gitea note
----------------------------------------------------------------------
title title verbatim, both ways; `/` is the
only hierarchy either side has
path — local only; derived from the title
order — local only; the wiki cannot sort
body content_base64 base64, utf-8, verbatim
— sub_url lands in the manifest as sub_url
— last_commit.sha lands as sha
— html_url lands as url
sub_url is the identity, and it is NOT derivable
------------------------------------------------
Gitea stores a wiki page as one flat file whose name it escapes from the title,
and the escaping is not a mapping worth reimplementing:
"Abstract Issue" -> Abstract-Issue.md space -> dash
"zz-probe/child" -> zz-probe%2Fchild.-.md / -> %2F, and a
LITERAL dash forces a
`.-` marker so the two
cases stay distinct
Every rule there is Gitea's to change. So `sub_url` is read back from whatever
the API returned and stored; it is never constructed here, and a caller that
needs to address a page fetches the listing rather than guessing. Building one
by hand is how you get a second page instead of an edit.
The wiki is flat, and only titles are structured
------------------------------------------------
There are no directories. A real subdirectory committed into the wiki's git
repository is invisible to the API and to the web UI — a ghost file. All nesting
lives in the title, which is why `page.py` treats `/` as its only separator.
"""
import base64
# A page's whole shape on the wire, for reference and for tests. Gitea also
# returns `commit_count`, `sidebar` and `footer` on a single-page GET; none of
# them describe the page itself, so none of them cross.
WIRE_KEYS = ("title", "sub_url", "html_url", "content_base64", "last_commit")
def decode(payload):
"""content_base64 -> text. Missing content is "" and not None: a page that
exists with an empty body is a real state, and the caller writing a file
should not have to tell the two apart."""
b = payload.get("content_base64") or ""
return base64.b64decode(b).decode("utf-8", "replace") if b else ""
def encode(text):
return base64.b64encode(text.encode("utf-8")).decode("ascii")
def from_payload(payload):
"""Gitea JSON -> the manifest fields the wiki layer owns, plus the title
the domain owns. The caller merges this into the existing entry so that
domain keys it does not mention (`order`) survive."""
commit = payload.get("last_commit") or {}
author = commit.get("author") or {}
return {
"title": payload.get("title") or "",
"sub_url": payload.get("sub_url") or "",
"url": payload.get("html_url") or "",
"sha": commit.get("sha") or "",
"remote-updated": author.get("date") or "",
}
def new_payload(title, text, message):
"""POST /repos/{owner}/{repo}/wiki/new.
`title` carries the hierarchy; Gitea derives the filename from it and
returns the sub_url it settled on. `message` is the wiki commit message —
the operator's words, not a generated one, because this is the only record
of why a page changed."""
return {"title": title, "content_base64": encode(text), "message": message}
def edit_payload(title, text, message):
"""PATCH /repos/{owner}/{repo}/wiki/page/{sub_url}.
The same shape as a create. Sending the unchanged title is a no-op; sending
a different one is a RENAME, which moves the file and leaves nothing at the
old sub_url — so callers pass the title from the manifest unless the
operator asked for a rename."""
return {"title": title, "content_base64": encode(text), "message": message}
def page_endpoint(base, sub_url):
"""The address of one page. `sub_url` goes in verbatim — Gitea hands it
back already escaped (`%2F` and all), and re-encoding it here would produce
a path that resolves to nothing."""
return "%s/wiki/page/%s" % (base, sub_url)
def revisions_endpoint(base, sub_url):
return "%s/wiki/revisions/%s" % (base, sub_url)
def matches_prefix(title, prefix):
"""Is this page at, or under, a title prefix?
The wiki being flat, "children" is exactly this test and nothing more:
there is no tree to walk, only a naming convention to trust. An empty
prefix matches everything."""
if not prefix:
return True
return title == prefix or title.startswith(prefix + "/")