ci(triage): name the real fault when the Claude credential is rejected

The AI pipeline has been down since ~2026-08-25. Every Claude-backed step fails
with:

  "result": "Failed to authenticate. API Error: 401 OAuth access token has been
             revoked."
  "error": "authentication_failed", "api_error_status": 401

The CLAUDE_CODE_OAUTH_TOKEN secret (last updated 2026-07-24) has been revoked.
That is not fixable in code — it needs regenerating — but the six days it went
unnoticed are, because nothing on the way out said so.

What a maintainer actually saw was the action reporting:

  "--json-schema was provided but Claude did not return structured_output.
   Result subtype: success"

which points at the schema, and then Apply verdict's generic "usually a
workflow-level fault ... left untouched for a retry". Neither mentions
credentials, and the failure presents per-issue while the real scope is every
tier at once: tier 1 cannot label, so tier 2's batch is empty and the sweep
reports success daily having done nothing.

GATE 1 now checks the execution file for authentication_failed / HTTP 401
before the max-turns branch and says what is wrong and what to do — regenerate
with `claude setup-token`, update the secret in the CR_PAT environment, and set
AI_DISABLED=true to silence the runs meanwhile. Same array guard and
fail-closed posture as hit_max_turns: an unrecognised shape is simply not an
auth failure and falls through to the generic branch.

Behaviour is otherwise unchanged — this branch already exited 1 without
touching labels, which was correct for a systemic fault.

Verified against the exact execution-file shape captured from the live 08-30
failure (both the issues and catch-up paths report the new error), and
regression-checked that max_turns still escalates on the automated retry, that
generic failures keep the generic message with and without an execution file,
and that the verdict paths are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
claude-ai-fix[bot]
2026-08-31 10:20:17 +02:00
parent c8ca11604e
commit 3ac46b0cd5

View File

@@ -362,6 +362,25 @@ jobs:
"$EXECUTION_FILE" >/dev/null 2>&1
}
# Did the run die because the Claude credential is bad? The action
# reports this uselessly — a revoked token surfaces as "--json-schema
# was provided but Claude did not return structured_output", which
# points at the schema and not at auth. The execution file carries the
# truth: api_retry / result objects with error "authentication_failed"
# and a 401. Same array guard and fail-closed posture as above; an
# unrecognised shape simply is not an auth failure and falls through
# to the generic branch.
hit_auth_failure() {
[ -n "${EXECUTION_FILE:-}" ] && [ -s "${EXECUTION_FILE:-}" ] || return 1
jq -e '(type == "array") and
any(.[]?;
(type == "object") and
(((.error? // "") == "authentication_failed") or
((.error_status? // 0) == 401) or
((.api_error_status? // 0) == 401)))' \
"$EXECUTION_FILE" >/dev/null 2>&1
}
# GATE 1 — did the action itself run? This is checked BEFORE looking
# at the payload, because the action can fail *after* having written
# a valid structured output: the object would sail through the shape
@@ -380,6 +399,16 @@ jobs:
if [ "${CLASSIFY_OUTCOME:-}" = "failure" ]; then
restore_needs_info
# Checked before anything else, because it is the one failure with
# a specific remedy and it takes down every tier at once — tier 1
# cannot label, so tier 2's batch is empty and the whole pipeline
# goes quiet while each run still fails in a way that reads like a
# per-issue problem. Say plainly what is wrong and what to do.
if hit_auth_failure; then
echo "::error::CLAUDE_CODE_OAUTH_TOKEN is rejected (HTTP 401 / authentication_failed). This is NOT a problem with issue #$ISSUE — every AI workflow is down until the credential is replaced. Regenerate it with 'claude setup-token' and update the CLAUDE_CODE_OAUTH_TOKEN secret in the CR_PAT environment. Set the AI_DISABLED repo variable to 'true' to silence these runs meanwhile."
exit 1
fi
# ...with one exception. Exhausting the turn budget is NOT a
# workflow fault: the action ran fine and this particular issue was
# just too tangled to finish inside the turn budget. Treating it as systemic