Skip to content

fix(cli): partial failures set the exit code and the headline (CLI-6) - #747

Merged
soustruh merged 6 commits into
mainfrom
claude/issue-745-partial-failure-exit-code
Oct 9, 2026
Merged

soustruh merged 6 commits into
mainfrom
claude/issue-745-partial-failure-exit-code

Conversation

@padak

@padak padak commented Sep 8, 2026 •

Copy link
Copy Markdown
Member

What was wrong

  • Several commands keep going when one item fails, and collect the failure in the result. They then printed a green Success: line and exited 0.
  • A sync clone where every config failed printed Success ... 0 created and exited 0. A script could not tell a clean run from a total failure.
  • sync push --all-projects counted a project whose push returned per-config errors[] as a success, and printed OK for it.
  • storage describe-batch --json exited 0 when items failed. Only human mode exited 1.
  • flow schedule-remove dropped a failed schedule delete when it deleted another schedule, and printed Success:.
  • semantic-layer build did not show a table that it left out of the model because it could not read the table schema.

What changed

  • A command exits 1 when at least one item failed. Exit 1 means that something the command did failed, reads included. These commands changed: sync push, sync push/pull/diff --all-projects, sync clone (also for bucket_errors[]), org setup, project refresh, project invite --from-csv, workspace gc, semantic-layer import, promote, build and edit metric --new-name, storage describe-batch --json and flow schedule-remove.
  • A --dry-run of these commands exits 1 when it reports a failed item, like the real run. storage delete-bucket --dry-run and describe-migrate --dry-run already did this. A clean dry-run exits 0.
  • Commands that printed a green Success: headline (sync push, sync clone, storage describe-batch, storage describe-migrate, flow schedule-remove) now print Failed: with the failed count. The other commands already list failures in a table or summary line. sync push --all-projects marks a failed project with x instead of OK.
  • The --json payload is printed before the exit. Its keys stay the same, with three additions: sync push --all-projects counts a project with per-config errors[] in summary.failed, flow schedule-remove returns an errors[] list, and org setup --refresh lists failed token refreshes under projects_refresh_failed.
  • The new helper item_failure_exit_code() in commands/_helpers.py returns the exit code, and the caller raises typer.Exit, as with map_error_to_exit_code().
  • Other read-only commands across several projects (config list, job list, billing credits, schedule list and others) still exit 0 when one project fails.
  • gotchas.md, commands-reference.md, CLAUDE.md, AGENT_CONTEXT, keboola-expert.md and the sync and workspace workflow references document the new exit codes with (since vNEXT).

Tests

  • tests/test_partial_failure_exit_code.py covers each changed command: a failed item exits 1, a clean run exits 0, a --dry-run with a failure exits 1, a clean --dry-run exits 0, and --json prints the full payload.
  • It also covers the push_all count in the service, the x mark and the failed count in the --all-projects push line, the clone bucket errors, and the interactive preview of org setup and project refresh.
  • tests/test_flow_service.py covers the new errors[] of remove_flow_schedule.
  • If you remove a fix, its test fails.
  • make check passes.

Fixes #745

Commands that survive a per-item failure collected it into the result and
then printed a fixed green "Success:" line and exited 0 anyway. A `sync
clone` where every config failed reported `Success ... 0 created` and exit
0; a `sync push` where half the configs failed reported success too. The
failures showed only as warnings underneath, so nothing in the exit code
distinguished a clean run from a total failure and no script could branch
on it.

Add `exit_on_item_failures()` to the command helpers and apply it to the
commands the audit found still missing it. The exit code follows the
convention the bulk storage commands already used (any per-item failure is
a general error, exit 1):

- `sync push` (single project and `--all-projects`)
- `sync clone`
- `sync pull` / `sync diff` with `--all-projects`
- `org setup` (`projects_failed`)
- `project invite --from-csv` (`failed`)

The human headline now states the failure instead of claiming success --
`Failed: Pushed: 3 created, 0 updated, 0 deleted, 1 failed` -- so the
failed count is in the main line rather than only in the warnings below
it. `sync push`'s inline human rendering moved into `_render_push_result`
because its early `return`s for `no_changes` / `dry_run` would otherwise
bypass the exit call.

`--json` output is unchanged and is emitted BEFORE the non-zero exit, so a
JSON caller still receives the complete payload including the per-item
error detail.

Read-only multi-project fan-outs (`billing credits`, `job list`, `schedule
list`, `notification list`) are deliberately excluded: they document
per-project degradation as intended behavior, and failing the whole read
because one project is unreachable would break existing callers.
@linear-code

linear-code Bot commented Sep 8, 2026

Copy link
Copy Markdown

CLI-6

@soustruh

soustruh commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Dear Claude, without reviewing this PR any further, I'd just like to highlight this one thing – exit_on_item_failures raises typer.Exit itself, while the existing map_error_to_exit_code is a pure mapper that returns the code and lets the caller raise. Two exit-code helpers side by side with opposite shapes is inconsistent. Consider returning the code (int | None) and keeping the raise at the call site, matching the existing convention.

@soustruh
soustruh marked this pull request as ready for review October 8, 2026 13:10

@keboola-pr-reviewer-bot keboola-pr-reviewer-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: needs_human (risk 3/5) · profile keboola-mcp-server

Well-tested CLI exit-code contract change across ~15 commands; behavior change with CI blast radius warrants a human sign-off.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 4 potential issues.

Devin Review

Comment thread src/keboola_agent_cli/commands/org.py Outdated
Comment on lines +289 to +290
if code := item_failure_exit_code(len(result.get("projects_failed", []))):
raise typer.Exit(code=code)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Failed token refreshes report success

When org setup --refresh fails to refresh a skipped project, result omits that project's failure. refresh_result contributes only successful refreshes, so the command exits 0 despite the failed refresh.

Learn more

Organization setup first registers new projects, then --refresh processes already-registered projects with refresh_tokens. The refresh result has its own projects_failed list, but org_setup copies only its successful projects_refreshed list. The final exit check reads only the setup result's projects_failed, so it never sees those refresh failures.

Example: A registered project prod has an invalid token. org setup --refresh --yes skips it during onboarding; its token refresh fails. The final output contains no failed project and exits 0, leaving its invalid token unchanged.

Recommended fix: Merge refresh_result['projects_failed'] into result['projects_failed'] before formatting and computing the exit code. Ensure the interactive preview also reflects refresh failures where applicable.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +653 to +656
# A workspace that could not be deleted, or a project that could not be
# listed, is collected in errors[]; the run is then not a success (#745).
if code := item_failure_exit_code(len(result.get("errors", []))):
raise typer.Exit(code=code)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Dry-run cleanup conceals project failures

When workspace gc --dry-run cannot list a project, result triggers exit 1 without printing its errors. The human output lists only workspaces, leaving users unable to identify the failed project.

Learn more

The gc_workspaces dry-run result includes listing failures in errors. The human dry-run renderer workspace_gc prints only would_delete, while the new exit logic fails on errors. Thus the exit code changes but the user cannot see which project was not scanned.

Example: Cleanup checks prod and dev, but listing dev fails. Human dry-run output shows prod workspaces and exits 1 without identifying dev or showing its API error.

Recommended fix: Render errors for both dry-run and real runs, using the listing error's project_alias and message keys and the delete error's workspace_id and error keys.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +5865 to +5866
- the human headline starts with `Failed:` and states the failed count, for
example `Failed: Pushed: 3 created, 0 updated, 0 deleted, 1 failed`;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Partial-failure headline promise is inconsistent

The documentation promises Failed: headlines, but org setup, project refresh, bulk invite, semantic-layer import and promote retain their existing headings. Align the renderers or narrow the promise.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +5892 to +5894
- `flow schedule-remove` (`errors[]`, a new key: a schedule whose delete
failed while other schedules were deleted. Before, that failure was dropped
and the command printed `Success: Removed N schedule(s)`).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Total schedule failures use a different response

When every schedule delete fails, remove_flow_schedule still raises instead of returning errors[]. The full JSON payload promise applies only when at least one delete succeeds.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

…re-exit-codes

# Conflicts:
#	plugins/kbagent/skills/kbagent/references/commands-reference.md
@soustruh
soustruh requested a review from zajca October 9, 2026 12:43

@zajca zajca left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable findings from the automated review.

Comment thread src/keboola_agent_cli/commands/org.py Outdated

@soustruh soustruh left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review of #747 — fix(cli): partial failures set the exit code and the headline

Generated by kbagent-pr-reviewer subagent. Verdict and findings below are advisory; the human author retains every veto. CI-coverable issues (lint, format, tests) are confirmed via make check, not duplicated here.

Summary

The PR makes the commands that collect per-item failures exit 1 and print a Failed: headline (sync push, sync clone, storage describe-batch, flow schedule-remove), and it adds the same exit code to the multi-item commands of the org, project, workspace, semantic-layer and sync --all-projects groups. I checked the result shape that each service returns against the key that each command reads, and they match. make check passes. Verdict: REQUEST CHANGES, because three older statements in gotchas.md and member-workflow.md still say that these commands exit 0 on a partial failure, which contradicts the new behavior and the new section in the same file.

Verdict

  • Verdict: REQUEST CHANGES
  • Blocking findings: 1
  • Non-blocking findings: 3
  • Nits: 3

Blocking findings

[B-1] plugins/kbagent/skills/kbagent/references/gotchas.md:2600 — older entries still say a partial failure exits 0

Three statements that were written before this PR now say the opposite of the new behavior. (1) gotchas.md:2600 says project invite --from-csv "exits 0 with failed > 0 reflected in the JSON summary" and "mirrors the org setup partial-success exit semantics". (2) member-workflow.md:90 says "Partial-success exits 0 with failed > 0 reflected in the JSON; this mirrors org setup". (3) gotchas.md:5813 says that a 403 on every new item of semantic-layer import / promote is "listed under failed, exit 0"; this text comes from main (it was added with the scope options of import / promote) and is merged into this branch. This PR changes all four commands to exit 1 and says so in the new section at gotchas.md:5878, so the same file now gives two opposite answers, and member-workflow.md (the workflow file of project invite --from-csv) is not in the diff at all. An agent that reads the older entry expects exit 0 and treats exit 1 as a total failure. Fix: change the three statements to exit 1, tag the change (since vNEXT), and keep the note that 0.97.0 and older exit 0.

Non-blocking findings

[NB-1] src/keboola_agent_cli/commands/org.py:284 — org setup --refresh --json changes the projects_failed payload, and the PR does not say so

The new code appends the failures of refresh_tokens to result["projects_failed"]. Before this PR those failures did not reach the output. They have another shape than the setup failures: {alias, project_name, error} with no project_id (services/org_service.py:394-399; pinned by tests/test_partial_failure_exit_code.py:628). A consumer that reads projects_failed[*].project_id now fails on a refresh failure. The PR description says the JSON keys stay the same "with two additions", and neither gotchas.md nor commands-reference.md names this third change. Fix: document the mixed shape in gotchas.md and in the PR description, or add project_id to the refresh entries.

[NB-2] src/keboola_agent_cli/commands/semantic_layer.py:637 — promote failure counting is not tested for 3 of 5 item types

I removed each type in turn from the hard-coded plurals tuple and re-ran the PR tests. Without datasets, relationships or constraints the promote tests stay green. Only metrics and glossary are pinned (tests/test_partial_failure_exit_code.py:913-967), so "if you remove a fix, its test fails" does not hold for these three. The tuple also repeats PUSH_ORDER (services/_semantic_layer_internals.py:79), so a new item type would not count towards the exit code. Fix: run the promote failure test for all five types (parametrize) and keep one shared tuple.

[NB-3] src/keboola_agent_cli/commands/sync.py:1138 — four command files over the soft size ceiling grow, and sync.py has 11 code lines left before the hard ceiling

make loc-check is green (warnings only). On main these files already exceed the 800-line soft ceiling, and this PR adds code to each: sync.py 1161 to 1189 (hard ceiling 1200), project.py 1012 to 1022, semantic_layer.py 878 to 895, flow.py 802 to 813. CONTRIBUTING.md says the next PR that adds material to such a file should split it first. The next change to sync.py will fail loc-check. Fix: move the push renderers (_format_push_result, _push_one_liner, _render_push_result) into a private module, as this PR's base already did for the clone renderer in _sync_clone_render.py.

Nits

  • [NIT-1] PR description — it says "The human headline starts with Failed:" for all changed commands. The code does this for sync push, sync clone, storage describe-batch, storage describe-migrate and flow schedule-remove, and semantic-layer build prints a Failed: section. The other commands list the failures in a table or summary line. CLAUDE.md:1016 and the new gotchas.md section say this correctly; fix the description so a squash commit does not overstate it.
  • [NIT-2] CLAUDE.md:1015, src/keboola_agent_cli/commands/context.py:2463, plugins/kbagent/skills/kbagent/references/gotchas.md:5925 — they say flow schedule-remove returns errors[] "only for a partial failure". The service returns the key on every successful call (services/flow_service.py:1009), empty when nothing failed. Say "errors[] is empty unless a delete failed".
  • [NIT-3] CLAUDE.md:1010 — the note on partial-failure exit codes sits under sync push but covers 14 commands in seven groups. A reader of the org setup or workspace gc lines does not see it. Consider the header comment at the top of ## All CLI Commands.

Verification log

  • gh pr view 747 → state OPEN, base main, 24 files, +1567/-84, title prefix fix(cli): matches a behavior fix ✓. No version bump and no changelog.py change in the diff ✓.
  • Local HEAD is 708a6360, the same SHA as headRefOid. The local branch name (fix/745-partial-failure-exit-codes) differs from the PR head ref (claude/issue-745-partial-failure-exit-code); I used the SHA match as the check.
  • Read the full diff against origin/main (already merged into the PR head): all source files, tests, CLAUDE.md, context.py and every plugin file.
  • 3-layer greps over the diff → no typer/click/formatter in services/, no httpx in commands/, no raw error_code literal in src/, no bare except:, no print(, no token-like string ✓.
  • make check on a copy of the PR head (git archive into a scratch directory, uv sync --frozen, a throwaway git repo because skill-check needs one) → ruff check ✓, ruff format ✓, ty 0 errors (1 warning: hatchling not importable in scripts/hatch_build.py, an environment gap) ✓, skill-check ✓, version-check ✓, version-gate-check ✓, command-sync-check ✓ (all CLI commands registered and documented), endpoints-check ✓, check-error-codes ✓, check-sentinel-guards ✓, loc-check ✓ (warnings, see NB-3), unit suite 0 failed ✓. changelog-check needs the gh repo context, which the copy lacks; I ran scripts/generate_changelog.py --check from the PR checkout instead → all releases have entries ✓.
  • Mutation check: I applied 14 single-line mutations to the new exit and count logic in the copy and re-ran the PR tests. 13 were caught. The promote tuple mutation survived; per-type results are in NB-2.
  • Behavior check without a Keboola project (no credentials were used, no real project was called): real FlowService with a mocked client behind the real CLI, flow schedule-remove with one of two deletes failing. Human mode → Failed: Removed 1 schedule(s) from flow 5, 1 failed, then Warning: Error: schedule 78: boom [x] with the brackets escaped, exit 1 ✓. --json → full payload with errors[], exit 1 ✓.
  • The other changed commands were not run against a project. I checked them through the PR's CliRunner tests and by reading each service result against the key the command reads: push_all summary, sync push / sync clone errors[] and bucket_errors[], org setup and refresh_tokens projects_failed, workspace gc errors[], promote / import failed[], build fetch_errors[], edit metric cascaded_constraints[] ✓.
  • Plugin synchronization map: the PR adds no command, so OPERATION_REGISTRY, CommandHint and the keboola-expert.md matrix need no row (check_command_sync.py agrees). AGENT_CONTEXT, CLAUDE.md, commands-reference.md, gotchas.md (vNEXT tags) and keboola-expert.md §3 are updated; keboola-expert.md is 68930 B of the 70000 B budget. grep for older exit-0 statements found the three in B-1.
  • E2E: no new command, so no new E2E test is required. Not run (no credentials).

Open questions for the author

  • sync diff --all-projects exits 1 when one project fails, while billing credits, job list, schedule list and config list exit 0 in the same case. The docs name the second group as a deliberate exception, so this is not a defect. Is the stricter rule for the sync family intended, and should agents get one rule?
  • With a partial failure the --json envelope still has "status": "ok" and the process exits 1 (seen on flow schedule-remove: envelope status: ok, data.status: removed, exit 1). The docs say to parse the payload and then check the exit code. Is a consumer that keys on the envelope status expected to treat exit 1 as the failure signal?

@soustruh
soustruh requested a review from zajca October 9, 2026 13:14

@zajca zajca left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No actionable findings were found by the automated review.

@soustruh
soustruh merged commit c0f5596 into main Oct 9, 2026
4 checks passed
@soustruh
soustruh deleted the claude/issue-745-partial-failure-exit-code branch October 9, 2026 15:31
@padak padak mentioned this pull request Oct 9, 2026
padak added a commit that referenced this pull request Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Commands report success and exit 0 even when per-item operations fail

4 participants