fix: apply --limit to the write, not to the selection
pull.py's docstring said "the limit is on the write, not on the selection", and the code did the opposite: `list_issues` truncated the payload list to `limit`, and pull.py dropped the closed ones after that. A milestone whose first issues are closed therefore answered `--limit 20` with twelve files, and the only statement about the behavior anywhere was the false one. The limit now counts what the run leaves in the store. `list_issues` takes a `keep` predicate, pages keep arriving until `limit` payloads have satisfied it, and the ones that did not are still returned — they were enumerated, and pull.py still reports them as "N closed, not stored". What `keep` means stays the caller's business; the transport only counts. pull.py hands it `lands_in_store`, which is the same test the walk itself applies: a closed issue counts only when the store already has it, since that one is refreshed rather than dropped. Pagination is the other half, and it cuts both ways. `paginate` is now a thin wrapper over a new `pages` generator, so the page after the one that fills the budget is never requested. In the other direction "fetch until N are kept" is "fetch the whole tracker" on a filter that matches mostly closed issues, so a keep-bounded read scans at most PAGE_SLACK times the pages the limit would need if nothing were dropped, then warns on stderr and returns short. Raising --limit raises that ceiling with it. --deps is outside the count: a dependency is followed because an issue named it. remote.py keeps the old meaning and now says so in as many words — it writes nothing, so there is no write for a limit to bound, and its --limit caps the listing, closed issues included. Same flag, two jobs, documented in both scripts and in the skill's command table. Also refuses `--limit 0` instead of dividing by the page size and raising ZeroDivisionError. tests/test_pull_limit.py stubs the transport with a fake that serves `page=`/`limit=` itself, so the request pattern is observed rather than assumed: exactly N files out of a half-closed selection, the second page fetched and the third not, the scan stopping at the budget with a warning, and remote.py's listing unchanged. 251 tests, no network. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -28,11 +28,29 @@ not exposed (404 on Gitea 1.26) — use milestones or labels, or the web UI.
|
||||
|
||||
A closed issue is not a unit of work, so filter mode enumerates it but leaves
|
||||
it out of the store: `--state all` still shows the whole picture, and only
|
||||
`--state closed` writes one. The limit is on the write, not on the selection —
|
||||
an issue already on disk is refreshed either way, so the local copy learns it
|
||||
was closed instead of staying open forever, and the count of the ones left out
|
||||
goes to stderr. Key mode is exempt: an address is not a bulk read, and
|
||||
`pull.py 1` fetches a closed issue as it always did.
|
||||
`--state closed` writes one. An issue already on disk is refreshed either way,
|
||||
so the local copy learns it was closed instead of staying open forever, and the
|
||||
count of the ones left out goes to stderr. Key mode is exempt: an address is not
|
||||
a bulk read, and `pull.py 1` fetches a closed issue as it always did.
|
||||
|
||||
**`--limit` is on the write, not on the selection.** It counts the issues this
|
||||
run puts in the store — written, or left in place by `--cached` — and never the
|
||||
closed ones it enumerated and threw away. `--limit 20` over a milestone whose
|
||||
first 30 issues are closed still writes 20, if 20 open ones are there to write:
|
||||
pages keep coming until the budget is full. Two boundaries keep that honest:
|
||||
|
||||
- Pages stop the moment the budget is full. Never one page more.
|
||||
- A filtered read may scan at most `_gitea.PAGE_SLACK` times the pages the limit
|
||||
would need if nothing were dropped. A filter that matches almost only closed
|
||||
issues therefore ends in a warning and a short answer, not in a walk of the
|
||||
whole tracker. Narrow the filter, or raise `--limit`, which raises the budget
|
||||
with it.
|
||||
- `--deps` is outside the count: a dependency is followed because an issue
|
||||
named it, not because the filter selected it.
|
||||
|
||||
`remote.py` is the deliberate exception, and it is not the same flag twice: it
|
||||
writes nothing at all, so there is no write to bound and its `--limit` means
|
||||
what it says — how many lines to print.
|
||||
|
||||
Comments ride along by default, in both modes and for every issue written:
|
||||
the thread lands in tmp/issues/<id>.comments.md, beside the issue. It costs
|
||||
@@ -98,6 +116,23 @@ def id_for(payload, store_ids, remote_map, repo, root):
|
||||
taken=store_ids)
|
||||
|
||||
|
||||
def lands_in_store(payload, drop_closed, store_ids, remote_map, repo, root):
|
||||
"""Would this payload leave a file in the store? The `--limit` predicate.
|
||||
|
||||
It has to be the same test the walk below applies, or the budget is spent on
|
||||
issues that never land — which is the bug this exists to prevent. So: a
|
||||
closed issue counts only when the store already has it (it is refreshed, and
|
||||
that is a write); anything else counts, including one `--cached` will skip,
|
||||
because a skipped issue is still an issue the store holds when the run ends.
|
||||
|
||||
Cheap in the common case: only a closed payload costs an `id_for`, and that
|
||||
is a lookup plus, at worst, a stat."""
|
||||
if not (drop_closed and payload.get("state") == "closed"):
|
||||
return True
|
||||
id = id_for(payload, store_ids, remote_map, repo, root)
|
||||
return os.path.isfile(issue.path_of(root, id))
|
||||
|
||||
|
||||
def comments_path(root, id):
|
||||
"""Where an issue's comment thread lives — beside it, under the same slug.
|
||||
Named in `_gitea` because push.py has to delete the same file."""
|
||||
@@ -131,7 +166,9 @@ def main():
|
||||
ap.add_argument("-q", "--query", help="search text in title/body")
|
||||
ap.add_argument("--state", default="open", choices=["open", "closed", "all"],
|
||||
help="filter mode only (default: open)")
|
||||
ap.add_argument("--limit", type=int, default=100, help="filter mode cap (default: 100)")
|
||||
ap.add_argument("--limit", type=int, default=100,
|
||||
help="filter mode: how many issues to STORE, not to enumerate"
|
||||
" (default: 100)")
|
||||
ap.add_argument("--deps", action="store_true", help="follow dependencies and pull them")
|
||||
ap.add_argument("--depth", type=int, default=3, help="max dependency depth (default: 3)")
|
||||
ap.add_argument("--cached", action="store_true",
|
||||
@@ -180,9 +217,14 @@ def main():
|
||||
|
||||
# ---- seeds -----------------------------------------------------------
|
||||
if filtered:
|
||||
# The limit bounds the write, so the transport is told what a write is
|
||||
# and counts those; the closed ones it enumerated on the way come back
|
||||
# in the list anyway, to be reported and dropped below.
|
||||
payloads, ms_title = _gitea.list_issues(
|
||||
login, base, state=args.state, labels=args.label, query=args.query,
|
||||
milestone=args.milestone, limit=args.limit)
|
||||
milestone=args.milestone, limit=args.limit,
|
||||
keep=lambda p: lands_in_store(p, drop_closed, store_ids, remote_map,
|
||||
repo, root))
|
||||
if not payloads:
|
||||
_gitea.die("no issues match that filter")
|
||||
what = []
|
||||
|
||||
Reference in New Issue
Block a user