Compare commits

...

17 Commits

Author SHA1 Message Date
Alexandre
ca0f78f10e fix(portainer): request identity encoding on the ingress listener (#2943)
* fix(portainer): request identity encoding on the ingress listener

Portainer compresses its own responses, so every response reached Home
Assistant's ingress relay gzipped and chunked, with no Content-Length.
Both relay hops (Supervisor api/ingress.py and Core hassio/ingress.py)
only take their buffered path for responses carrying a Content-Length
under 4 MB; everything else goes through the streaming path, where an
aiohttp error surfaces to the browser as 502 Bad Gateway even though the
add-on's own nginx logged a 200.

proxy_params.conf already stripped Accept-Encoding, but the location
block declares its own proxy_set_header directives, and nginx discards
every server-level proxy_set_header once a location sets any of its own
(the comment above those lines warns about exactly this). The strip was
therefore dead config.

The same server block serves both the ingress listener and the direct
web UI port, so the strip is scoped through a map on $server_port:
ingress gets identity, direct access keeps compression. The map keys on
the direct-access port rather than the templated ingress port, so the
default stays correct if the ingress port ever changes.

Verified with a local nginx against the live Portainer backend:
- ingress listener, client sending "Accept-Encoding: gzip, deflate"
  -> identity, Content-Length: 14203
- direct listener, same request -> Content-Encoding: gzip, chunked
- direct listener, no Accept-Encoding -> identity, Content-Length
- websocket upgrade through ingress still reaches Portainer (401 auth)
- nginx -t passes for both the ssl and non-ssl rendered variants

Partial mitigation only: vendor.js (5.7 MB) and main.js (7.0 MB) exceed
the 4 MB buffering threshold uncompressed and still stream. This
supersedes 2.43.0.1, which disabled nginx's own gzip module rather than
the compressor that was actually running.

Refs #2766

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(portainer): request identity explicitly on the ingress listener

Address review feedback: use "identity" rather than an empty value as
the map default. Both were verified to produce identity responses with a
Content-Length from Portainer, and both leave direct access on 9099
compressed, but "identity" states the intent explicitly instead of
relying on the server's choice when no Accept-Encoding is present.

Also reword the changelog entry for readability.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 22:48:26 +02:00
Alexandre
3e561a3f9f refactor(skill): shorten hassio-addon-workflow via progressive disclosure (#2942)
* refactor(skill): shorten hassio-addon-workflow via progressive disclosure

SKILL.md was 365 lines, loaded in full on every add-on task. Split it per
Anthropic's Agent Skills best practices: keep steps, completion criteria,
and the mechanism ladder inline; push rationale, war-story examples, and
Codex CLI invocation details into reference files loaded only when that
branch is taken.

- SKILL.md: 365 -> 171 lines. The "ship the simplest solution" rule was
  stated three times; now once. Steps 3/6/9 keep their load-bearing
  checklist but point to detail files instead of inlining it.
- references/evidence.md (new): measurement methodology, the
  host-generalization failure examples, the merged-and-inert case studies
- references/codex-review.md (new): CLI invocation, prompt guidance,
  plan-attack checklist
- references/simplify.md (new): mechanism-ladder case studies

Also replaces the old "Token efficiency" section (which named this
maintainer's personal MCP tools - rtk/headroom/tokensave - not guaranteed
present for CI agents or other collaborators using the checked-in copy)
with a subagent-delegation instruction: Codex's plan/code review and
PR-comment triage on >5 threads should run in a subagent that returns a
condensed summary, not raw output, into the calling session.

* fix(skill): address PR review feedback

- SKILL.md: define $SKILL once and state that all scripts/ and references/
  shorthand paths are relative to it — they read as repo-root-relative
  otherwise and don't resolve from an add-on directory
- codex-review.md: drop the reference to a "global CLAUDE.md note" that
  isn't in this repo's CLAUDE.md; describe --sandbox read-only accurately
  (reads allowed, writes/exec blocked, approvals disabled) instead of
  "cannot run anything"; add the missing & so the example is actually
  backgrounded as the prose claims
- evidence.md: use smaps_rollup for process RSS — plain smaps prints one
  Rss: line per mapping (47 for a trivial process here), not a total
- simplify.md: drop the absolute "removals cannot regress" claim, which
  contradicted the /dev/shm case study in evidence.md

* refactor(skill): apply independent quality-review findings

From an independent model review of the skill against the Agent Skills
best-practice guides:

- traps.md: CHANGELOG was described as "the only hard gate", contradicting
  SKILL.md — the HA add-on linter and the image build also block PRs
  (verified against onpr_check-pr.yaml). Gates list corrected.
- codex-review.md: the "attack your own plan" checklist is run by the main
  agent, but lived in the file whose stated consumer is the delegated
  subagent — an agent that delegates correctly would never see it. Moved
  inline into SKILL.md step 3.
- SKILL.md: delegation instruction now says exactly what to do (spawn a
  subagent whose prompt includes references/codex-review.md and follows its
  invocation) instead of "delegated per the rule above".
- SKILL.md: traps.md is no longer mandatory reading on the light path —
  the light-path facts it holds (versioning, CHANGELOG format) are inline
  in step 7; it stays required when touching scripts/Dockerfiles/env.
- SKILL.md: dropped the bundled-files table (every row already cited at
  point of use) and the "read the script when you use it" anti-instruction
  — scripts are run, not read. 177 -> 167 lines.
- description: 982 -> 550 chars; removed workflow narrative that does
  nothing for skill selection and the stale-prone model name, kept all
  trigger terms.

* fix(skill): note that CI hard gates skip on non-addon PRs

The three hard gates in onpr_check-pr.yaml (changelog check, addon-linter,
check-build) are each matrixed over check-addon-changes.outputs.changedAddons
and if:-skipped when it is '[]'. A PR touching only docs, .github/ or .claude/
therefore shows them as "skipping" rather than passing — which should not be
read as a green build. Verified against onpr_check-pr.yaml lines 64, 83, 101
and against this PR's own check output.
2026-08-05 10:11:11 +02:00
Alexandre
5abcb923c4 fix(webtrees): chown persistent dirs to the uid the app actually runs as (#2941)
/docker-entrypoint.py renumbers www-data to PUID:PGID after cont-init has
run, then re-chowns only its DATA_DIR. Chowning /config by name beforehand
left /config/modules_v4 owned by the image's original www-data uid (33),
so the running Apache/PHP process (uid 1000) could not create directories
in it -- breaking custom module installation from the webtrees UI.

Fixes #2940

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 08:35:20 +02:00
Alexandre
74456a03b6 chore: add hassio-addon-workflow skill for Claude Code (#2939)
* chore: add hassio-addon-workflow skill for Claude Code

Checks in the repo-specific skill so future Claude Code sessions get the
tiered scope->measure->plan->implement->verify workflow, repo traps, and
helper scripts (preflight/measure/env_trace/validate/pr_review) without
depending on a local machine's ~/.claude config.

* fix(skill): address PR review feedback from Copilot and Codex

- SKILL.md: repo-relative script invocation (skill is now checked in);
  correct the CI-gates list — the PR add-on linter is blocking, only the
  weekly Super-Linter is non-blocking
- preflight.sh: git-aware repo detection (worktrees have a .git file)
- pr_review.sh: header now documents resolve's actual --all behavior
- measure.sh: CPU% uses getconf CLK_TCK; sample all processes, not the
  top-24 by RSS
- env_trace.sh: validate VAR as a strict env-var name before regex use
- validate.sh: shellcheck also covers extensionless run/finish; --vs-master
  skips visibly when a linter is missing instead of reporting a false clean

* fix(skill): address CodeRabbit review feedback

- pr_review.sh: resolve exits nonzero unless every thread actually
  resolved; watch exits nonzero and says so when checks settle with
  failures instead of reporting bare "settled"
- validate.sh: pass the config.yaml path to Python as argv instead of
  interpolating $ADDON into the source (CWE-94)
2026-08-04 22:29:20 +02:00
Alexandre
7dfeb78c37 fix(claude_desktop): stop the GPU flags from disabling the GPU; make max_resolution work; drop the dead driver install (#2938)
The GPU acceleration shipped in 2026.08.03 was not inert — it was what turned
the GPU off. `--use-gl=angle --use-angle=gl-egl` forces Mesa's EGL X11 platform,
which offers no window-capable EGLConfig under this Xvfb. The GPU process logged
`gl_surface_egl.cc:262 No suitable EGL configs found`, abandoned GL, and was
relaunched with `--use-gl=disabled` while every renderer got
`--disable-gpu-compositing`.

Chromium already renders on the GPU here with no flags at all: LSIO's Xvfb runs
`-vfbdevice /dev/dri/renderD128`, so its GLX is backed by the real render node.
The premise that Xvfb offers only an indirect/software path was wrong for this
base image. Measured on a separate display, including at the production
15360x8640 screen — with the flags the GPU process loads libEGL_mesa and holds
1 fd on the render node; without them it loads the Mesa gallium megadriver over
DRI3, holds 8, and no renderer carries --disable-gpu-compositing.

claude-gpu-probe was not wrong about the hardware, it answered the wrong
question: it exercised ANGLE's default GLX path, which works, so it passed while
the flags it gated disabled the GPU. Removed with the gpu_acceleration option.

Also fixes two other changes from the same release that never did anything:

- max_resolution wrote MAX_RES into the s6 container_environment, but svc-xorg
  starts `#!/usr/bin/env bashio`, not with-contenv, and never reads it. Renaming
  the option to MAX_RES makes the add-on env layer inject it into every service
  run script, which is how DRINODE already reaches Xvfb. It ships with no
  default, so nothing changes until it is set: carrying over the old 1920x1080
  default would have silently shrunk every existing desktop, since that value
  had never taken effect. Schema bounds each axis to 100-9999.

- The amd64 driver install has never run in any release: guarded by `if [[ ]]`
  with no SHELL directive, so it runs under dash, which has no `[[` — condition
  false, RUN still exit 0. It is deleted rather than repaired, because it had
  nothing to add. Hardware GL already works without it, Mesa already comes from
  the LSIO base at a newer backports version, and intel-media-va-driver-non-free
  was installed live on the running add-on and changed nothing: VA-API failed
  identically to the free driver on both DRM nodes. That failure is below the
  add-on, in the host i915 stack.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 19:38:09 +02:00
Alexandre
f7c00ca818 Use upstream app_config-capable add-on linter (#2936)
* Add one-shot app_config linter migration workflow

* Add app_config linter compatibility action

* Add add-on map compatibility normalizer

* Test add-on map compatibility normalizer

* Run app_config compatibility preparation on pull request

* Validate app_config short and long map forms separately

* Export validated workflow updates

* Use app_config-compatible linter for PR checks

* Use app_config-compatible linter for builds

* Remove temporary linter preparation workflow

* Prepare minimal upstream linter update

* Use upstream app_config-capable linter

* Use upstream app_config-capable linter

* Remove local add-on linter wrapper

* Remove local linter compatibility script

* Remove local linter compatibility tests

* Remove temporary linter preparation workflow
2026-08-04 09:25:52 +02:00
dependabot[bot]
1ab1799276 Bump actions/stale from 10 to 11 (#2935)
Bumps [actions/stale](https://github.com/actions/stale) from 10 to 11.
- [Release notes](https://github.com/actions/stale/releases)
- [Changelog](https://github.com/actions/stale/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/stale/compare/v10...v11)

---
updated-dependencies:
- dependency-name: actions/stale
  dependency-version: '11'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-03 23:07:03 +02:00
GitHub Actions
49dd9f0e4e Revert "Migrate legacy addon_config maps to app_config (#2933)"
This reverts commit d3d3476986.
2026-08-03 14:22:18 +00:00
GitHub Actions
b453c60c33 Revert "GitHub bot: sanitize (spaces + LF endings) & chmod [nobuild]"
This reverts commit aceeaf36b4.
2026-08-03 14:22:18 +00:00
github-actions
aceeaf36b4 GitHub bot: sanitize (spaces + LF endings) & chmod [nobuild] 2026-08-03 13:42:52 +00:00
Alexandre
9faae5cc90 perf(claude_desktop): render on the GPU and stop duplicating MCP servers (#2934)
* perf(claude_desktop): render on the GPU and stop duplicating MCP servers

Measured live inside a running add-on (amd64, 4 cores): 3388 MB RSS across
73 processes, with the Electron renderer burning ~44% of a core even with no
browser client connected.

GPU: Chromium was rendering everything on the CPU. Under Xvfb it probes GLX,
finds only Xvfb's indirect/software path, and falls back to `--use-gl=disabled`
plus `--disable-gpu-compositing` — while a perfectly good iGPU sits idle behind
/dev/dri. Claude Desktop is now launched through ANGLE's OpenGL backend over
EGL, but only when the new claude-gpu-probe confirms that Desktop's own bundled
ANGLE can create a hardware GL context on this host; the probe rejects
llvmpipe/SwiftShader, is bounded by a timeout, and any failure leaves the
command line exactly as it was. New `gpu_acceleration` option (auto|on|off).

MCP: every stdio MCP server is a separate process per client, and Desktop
starts another full set for each Claude Code session it hosts. Claude Code now
reaches the Home Assistant MCP server over its native HTTP transport instead of
the mcp-proxy stdio bridge, removing the most expensive duplicate (~45 MB of
private RSS per copy). Desktop keeps the bridge: its remote-entry config schema
could not be confirmed, and guessing would silently break it. New
`mcp_servers_desktop` / `mcp_servers_code` options let each client register only
what it actually uses; defaults are unchanged.

Display: new `max_resolution` option (default 1920x1080) caps the virtual screen
via the base image's MAX_RES. Xvfb ran at 15360x8640, so it and the Selkies
capture loop tracked damage over a 133-megapixel area continuously. This is a
CPU saving, not a memory one — the framebuffer is a lazily populated shared
segment whose unused portion was never resident.

Dockerfile: the Intel graphics block was dead code. It was gated on TARGETARCH,
which this repo's builder does not pass, so it never ran — the shipped amd64
image has no vainfo and no intel-media-va-driver-non-free, and its apt history
contains no matching install. It now uses BUILD_ARCH, and its Vulkan ICD check
no longer names intel_icd.x86_64.json, a file Debian does not ship.

Also removes `--disable-dev-shm-usage` (a workaround for a 64 MB /dev/shm; this
image has 7.7 GB) and fixes stale Home Assistant MCP registrations, including
the bearer token inside them, being left behind when enable_ha_mcp is disabled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(claude_desktop): satisfy static analysis in claude-gpu-probe

Codacy flagged two new issues, both in the probe: a broad exception catch and
too many locals in main(). Split the EGL bring-up into load_angle(),
open_angle_display(), make_current_context() and describe_renderer(), each
raising a dedicated ProbeFailure, so the failure paths read as intent rather
than as a chain of early returns. The catch-all remains — a probe must never
stop the desktop from starting — but is now explicit and narrowly scoped.

No behaviour change: exits 0 with a hardware renderer under DISPLAY, 1 without.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(claude_desktop): address PR review on shm and MCP entry ownership

Two review findings, both correct.

/dev/shm: dropping --disable-dev-shm-usage outright was generalised from one
host. The flag was added to fix a real Electron renderer crash loop on Docker's
64 MB default, and Home Assistant ignores the add-on's shm_size, so the size
genuinely varies per install and cannot be asserted from this repo. The size is
now read at startup: the workaround is kept below 256 MB, dropped above it, and
kept when the size cannot be determined.

MCP ownership: claiming an HTTP 'homeassistant' entry by URL and shape would
have deleted a user's own manually configured server on the first boot after
upgrade, since ha_mcp_url defaults to the same public endpoint that a hand-
written entry would use, and enable_ha_mcp defaults to false. An HTTP entry is
now only ever modified or removed when the add-on recorded writing it, in
~/.config/claude_desktop_addon/managed-mcp.json. Anything not written by the
add-on is untouchable regardless of how it looks.

Also drops the invalid '?' optional marker from the list *item* type in the
mcp_servers_* schema; both keys always carry defaults, so the marker was
meaningless as well as wrong.

Tests cover the regression directly: a user-owned HTTP entry on the default URL
now survives both a disabled and an enabled boot, while the add-on's own entry
is still removed with its token when enable_ha_mcp is turned off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(claude_desktop): tighten max_resolution validation and reporting

grep anchors ^...$ per line, so a multi-line max_resolution such as
"1920x1080\n640x480" passed validation on its first line and was then written
to MAX_RES verbatim, leaving svc-xorg with a corrupt screen size. Bash's =~
anchors the whole string and rejects the embedded newline.

Also stop reporting success when no s6 environment directory existed and
nothing was written — the cap silently did not apply, and the log said it did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 15:27:30 +02:00
Alexandre
d3d3476986 Migrate legacy addon_config maps to app_config (#2933)
* Run app config map migration

* Migrate app config map names

* Add one-shot workflow to address PR 2933 review comments

* Trigger PR 2933 review-fix workflow

* Temporarily apply PR 2933 review fixes

* Restore lint workflow

* Remove temporary review-fix workflow

* Fix Cleanuparr app_config persistence path

* Document Cleanuparr app_config mount

* Fix qBittorrent migration paths and messages

* Correct qbit_manage configuration filename

* Align custom script assignment indentation

* Update Mealie app_config documentation path

* Apply remaining PR 2933 review fixes

* Correct Mealie historical path count

* Commit review fixes without workflow changes

* Address remaining PR review comments

* Restore lint workflow after review fixes

* Document Webtrees app-config host folder

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-03 15:24:16 +02:00
github-actions
c0aa5af0ce Github bot : image compressed 2026-08-02 23:22:36 +00:00
github-actions
87c9b60be0 GitHub bot: changelog [nobuild] 2026-08-02 15:59:48 +00:00
Alexandre
82ac7957c8 Update version in config.yaml 2026-08-02 17:57:45 +02:00
Alexandre
f0d20d77fa Update config.yaml 2026-08-02 17:57:20 +02:00
Alexandre
988cecb122 fix(birdnet-pi): restore the ingress Caddy site (502 Bad Gateway) (#2931)
* fix(birdnet-pi): restore the ingress Caddy site, ingress returned 502

nginx forwards ingress traffic to 127.0.0.1:8082 (rootfs/etc/nginx/servers/
ingress.conf:11), a site appended to the Caddyfile by helpers/caddy_ingress.sh.
The upstream script $HOME/BirdNET-Pi/scripts/update_caddyfile.sh regenerates
/etc/caddy/Caddyfile from scratch and 02-caddy.sh runs it immediately before
`exec caddy run`, so 91-nginx_ingress.sh injected a call to caddy_ingress.sh
into that script to re-add the site. The injection was anchored on
`sudo caddy fmt --overwrite`.

Since 2026.07.10-1 the Dockerfile strips `sudo ` from every BirdNET-Pi script at
build time (Dockerfile:110), so the shipped line is `caddy fmt --overwrite` and
the anchor stopped matching. sed reports success when a pattern matches nothing,
so this failed silently: Caddy came up listening only on :8081 and every ingress
request got connection-refused on 8082. Confirmed by extracting the script from
the published ghcr.io/alexbelgium/birdnet-pi-amd64 image - it contains no `sudo`.

- 91-nginx_ingress.sh: make the anchor accept the line with or without `sudo`,
  skip the injection when it is already there (cont-init re-runs on restart),
  and verify afterwards, falling back to appending the call if the anchor is
  ever gone again.
- caddy_ingress.sh: return early when a `:8082` site already exists. The script
  now runs from more than one place, and a duplicate site address makes Caddy
  refuse to start.
- 02-caddy.sh: re-add the ingress site just before starting Caddy if it is
  missing, so a future upstream change to update_caddyfile.sh cannot silently
  bring back the 502.
- 91-nginx_ingress.sh: drop /ingress_url when ingress is off, so that marker is
  a truthful signal for the check above even across an in-container restart.

Verified with a harness that replays the boot sequence (build-time sudo strip,
81-modifications.sh, 91-nginx_ingress.sh, 02-caddy.sh) against the real upstream
update_caddyfile.sh: master ends with no :8082 site, this branch ends with
exactly one, on fresh boot, on restart, and when the `caddy fmt` anchor is
removed entirely; standalone mode still gets no ingress site.

Closes #2928

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(birdnet-pi): harden the ingress-site recovery check in 02-caddy.sh

Review feedback on the pre-start check:

- Require /etc/caddy/Caddyfile to exist and silence grep's stderr. `! grep` is
  also true for grep's exit code 2, so an unreadable or missing Caddyfile read
  as "ingress site missing" and would have produced an ingress-only Caddyfile
  out of an error state. Letting caddy fail on the missing config is easier to
  diagnose.
- Report a failure of caddy_ingress.sh instead of swallowing it. Do not exit:
  /custom-services.d scripts are LSIO longruns that s6 restarts when they
  return, so exiting would flap the service and take the WebUI down on 8081 as
  well, which is strictly worse than ingress alone being broken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 17:56:19 +02:00
158 changed files with 1501 additions and 42 deletions

View File

@@ -0,0 +1,167 @@
---
name: hassio-addon-workflow
description: >-
Workflow for alexbelgium/hassio-addons Home Assistant add-on work: diagnose with real
measurements, independent Codex review, implement, open a PR, resolve CodeRabbit / Copilot /
Codex bot review comments, verify in production. Use for any task touching an add-on in this
repo — bugs, RAM/CPU/performance tuning, Dockerfile, config.yaml, cont-init.d or s6 changes,
version bumps, opening or iterating PRs — and when asked to "check with codex", "verify with
chatgpt", or resolve bot comments. Cheap for small asks: a light path skips the heavy steps.
---
# Home Assistant add-on workflow
Triage first, then one of two paths:
- **Light** — typo/doc fixes, CHANGELOG edits, version bumps, one-file edits at ladder levels
1-3 (below), simple questions: scope → implement → validate (`$SKILL/scripts/validate.sh
<addon> --vs-master`; `$SKILL` defined below) → PR (version bump + CHANGELOG still required) →
resolve bot comments.
- **Full loop** — performance/RAM/CPU work, diagnosis, anything changing a shipped default,
ladder levels 4-6, or an explicit Codex-check request: scope → measure → plan → Codex reviews
the plan → implement → simplify → Codex reviews the code → PR → resolve comments → verify in
production → report.
Escalate mid-flight if a light task grows — touches a default, needs a new script or service, or
reveals a deeper problem.
**Standing rule:** ship the simplest solution that works. Complexity is bought only by a
**measurement** showing a concrete, user-visible cost on a real host — never by reasoning about
hypothetical performance.
**Repo layout.** `alexbelgium/hassio-addons`; each add-on is a top-level directory. This skill is
checked in at `.claude/skills/hassio-addon-workflow/` (canonical copy). Set the skill root once,
then every `scripts/…` and `references/…` path below is relative to it:
```bash
SKILL="$(git rev-parse --show-toplevel)/.claude/skills/hassio-addon-workflow"
bash "$SKILL/scripts/preflight.sh" # and likewise for the other scripts
```
**Non-negotiables:**
- Docker build cannot be tested locally (no dockerd) — CI is the only gate.
- Never `git stash` under `/data/claude` — `refs/stash` is shared across worktrees.
- Work in a worktree under `/data`, not `/tmp` (`/tmp` is noexec).
**Delegate heavy output to a subagent.** Codex reviews and multi-thread PR triage produce output
you don't need verbatim in your own context. For Codex's plan review (step 3), Codex's code
review (step 6), and PR-comment listing when there are more than ~5 threads (step 8): launch a
subagent to run the command and report back only the objections/findings and your assessment of
each, not the raw transcript.
---
## 1. Scope
State goal, non-goals, constraints, and definition of done — two sentences, explicit. A diagnosis
ask ("why is it slow?") is not automatically a fix ask. Changing a shipped default is the user's
call, not yours — ask before implementing.
## 2. Evidence before reasoning (full loop)
Measure the running add-on rather than reasoning from source (`$BUILD_VERSION` set,
`HOME=/data/data`) — reviewers hold you to the numbers. Tool per question:
- RAM/CPU → `scripts/measure.sh` (≥20 s sample)
- "I set an option and nothing happened" → `scripts/env_trace.sh <VAR> <process>`
- Is this flag/driver/package actually present? → inspect the artifact: `/proc/<pid>/cmdline`,
`command -v`, `/var/log/apt/history.log`
Verify you're reading the right revision first — `scripts/preflight.sh` catches a stale branch
before it costs a full analysis pass. Measurement methodology, gotchas, and real failure examples:
`references/evidence.md`.
## 3. Plan — choose the mechanism level, then Codex reviews it (full loop)
Rank mechanisms, pick the lowest (simplest) one that solves it, and state the choice in the plan:
1. A config value — an option, a schema constraint, an existing env var.
2. An existing knob the base image already reads (`MAX_RES`, `DRINODE`, `SELKIES_*`).
3. A few lines in an existing script, at the point that already runs.
4. A new init script.
5. A new service, wrapper, or long-running process.
6. Custom protocol code, or patching someone else's internals.
Levels 4-6 need a reason that survives being said out loud ("upstream has no knob for this, and I
checked" is one; "it felt cleaner" is not) and mean full loop.
Attack your own plan before implementing:
- What does this do on a host **unlike this one** — no GPU, small `/dev/shm`, aarch64, a VM?
- What happens on **upgrade** to someone who configured this by hand?
- What is the **blast radius** if the assumption underneath it is wrong?
- What am I **inferring** that I could instead **detect at runtime** or **record explicitly**?
Highest-yield question here — see `references/evidence.md`'s failure-mode section.
Full loop only, before writing code: get Codex's independent read on the plan. Spawn a subagent
whose prompt includes the path `references/codex-review.md` and tells it to follow that file's
invocation, then report back only Codex's objections and an assessment of each — not the raw
transcript.
## 4. Implement
Touching a shell script, Dockerfile, or env option? Read `references/traps.md` first — skim the
headings, read the sections you're about to touch; the bashio, s6-env, arch-guard and versioning
traps are all live. (The light-path facts it holds — versioning format, CHANGELOG heading — are
already inline in step 7.) Validate with `scripts/validate.sh <addon> --vs-master`. Write
behavioural tests for anything with branches, targeting **the regression a reviewer described**,
not just the happy path.
## 5. Simplify
Before requesting review, check: did the diff stay at the ladder level chosen in step 3? Can this
be solved by deleting instead of adding? Is the fix bigger than what it fixes? How does it fail in
three years? Case studies of what happens when this check is skipped: `references/simplify.md`.
## 6. Codex attacks the code (full loop only)
Same delegated invocation, pointed at `git diff origin/master...HEAD` plus your reasoning per
hunk. Details in `references/codex-review.md`.
## 7. Open the PR
CI gates on a PR: **`CHANGELOG.md` updated** (hard fail), the **HA add-on linter**
(`frenck/action-addon-linter`, blocking — not the weekly Super-Linter, which is non-blocking), and
the **add-on image build**. Bump `version` anyway (`X.Y.Z.N`, never `X.Y.Z-N`, see
`references/traps.md#versioning`) — Supervisor won't offer a rebuild without it. Update
`README.md` if you added options; match the CHANGELOG heading format `## X.Y (DD-MM-YYYY)`.
Write the body to a file, `gh pr create --body-file`: state what was measured, what changed,
**what is not verified**, and how to roll back the riskiest hunk alone.
## 8. Resolve review comments
`scripts/pr_review.sh list|reply|resolve|status|watch <PR>`. For every comment, **reproduce the
claim before agreeing or disagreeing** — reviewers are frequently right and occasionally
confidently wrong; a reproduction takes a minute and decides it either way. Reply with the
evidence, then resolve. **Push back when you're right**, on the thread — a resolved-but-wrong
thread is worse than an open one.
## 9. Verify before declaring done
Never blur these three: **Verified** (you ran it and observed the result), **Checked** (parses,
lints, type-checks), **Assumed** (reasoning only — name the assumption). Do not write "this should
work" — either it was exercised, or say plainly it wasn't.
Light path: verification is `validate.sh` plus CI; anything beyond that is Assumed. Full loop: CI
passing proves the build works, not that the change does anything — re-run the measurement that
motivated the work once the rebuilt add-on is running. Real "merged and inert" examples, and what
to do when a fix can't be self-verified: `references/evidence.md`.
## 10. Calibrate and report
Close against the scope from step 1, not against what you ended up doing:
```
What was asked / what shipped — mapped to the original scope
Evidence — the numbers, before and after
Verified — observed, with how
Not verified — and why (e.g. no dockerd locally; CI is the gate)
Known broken / left out — explicitly, including anything descoped
Risk + rollback — the riskiest hunk and how to revert it alone
```
Lead with anything that did not work — a merged PR that achieved nothing is the single most
important sentence in the report. Give confidence per claim, not one blanket number.
Scripts are meant to be **run, not read** — each is cited at its point of use above; read one
only if its output surprises you.

View File

@@ -0,0 +1,41 @@
# Codex review — invocation and prompt guidance
Used for step 3 (plan review) and step 6 (code review), full loop only. Codex is a genuinely
different model reading the files itself; on this workload it has repeatedly been worth the
minutes. Delegate the invocation to a subagent (see SKILL.md's subagent-delegation note) so its
output doesn't land verbatim in your context — have the subagent return only Codex's objections
and your assessment of each.
## Invocation
**Use the CLI, not the MCP tool, for prompts of this size.** The `codex` MCP tool timed out twice
on ~4 KB prompts (2026-08-03); the CLI with the same content succeeded. The MCP tool is still fine
for short questions.
`--sandbox read-only` lets Codex read files but blocks writes and command execution, and
`approval_policy=never` means it will not be prompted for permission to run anything either — so
paste every number into the prompt rather than expecting Codex to gather it. `- <` feeds the
prompt file on stdin. Run it in the background so you are not blocked for the several minutes it
takes (`&` here, or your harness's background-task mechanism):
```bash
codex exec --model gpt-5.6-sol --sandbox read-only --skip-git-repo-check \
-c approval_policy='"never"' - < prompt.md > codex_out.txt 2>&1 &
```
## Writing the prompt
- **Plan review (step 3):** include the files to read, your measurements **with numbers**, the
proposed changes, and explicit instructions to challenge you. Ask direct questions ("is this
really add-on-fixable?", "give the precise flag set") rather than "review this".
- **Code review (step 6):** point it at `git diff origin/master...HEAD` plus the reasoning behind
each hunk. Ask specifically what breaks: upgrade paths, hosts unlike this one, users who
configured things by hand. Ask directly whether a simpler mechanism would achieve the same
thing — an outside reader spots one-level-too-deep framing far more easily than the person who
just built it.
- Codex's sandbox often cannot run local commands and falls back to reading GitHub, so paste the
evidence in rather than assuming it will find it.
**Codex agrees with confident premises.** It has confirmed a wrong conclusion stated too
confidently, and separately caught a genuine methodology error in the same review. Treat its
confirmations with the same scepticism as its objections — especially about the build.

View File

@@ -0,0 +1,54 @@
# Evidence — measurement methodology and case studies
## Why summed RSS and reserved-vs-resident both matter
- **Summed RSS double-counts shared pages.** Removing a duplicate process frees its *private*
memory, not its RSS. `scripts/measure.sh` reports PSS and private alongside RSS — quote
**private** when arguing "removing this saves N MB".
- **A big mapping is not necessarily resident.** Large SysV/tmpfs segments are lazily populated;
reserved size is reported separately from resident for this reason.
- `/proc/meminfo` and `free` show **host** figures (no memory cgroup namespace here) — never
attribute those to the add-on.
- Sample duration matters: a 3 s CPU sample measured 2.3% where a 20 s sample measured 21.6% for
the same process. Use ≥20 s for anything you report.
## Before asserting anything, ask what would show it false
- "This process is duplicated" → is it? `ps -ef --forest`, compare parents and start times.
- "This costs 500 MB" → is it resident? `grep Rss /proc/<pid>/smaps_rollup` — plain `smaps` prints
one `Rss:` line per mapping (dozens of them), not a process total.
- "This block never runs" → is its payload in the image? `command -v`, `apt` history.
- "The flag isn't set" → `tr '\0' '\n' < /proc/<pid>/cmdline`.
When you correct yourself mid-analysis, keep the correction visible in your notes and in what you
report — a retracted claim that stays retracted is worth more than one quietly dropped.
## The failure mode this loop keeps producing
Every bug shipped from the source session came from one move: **measuring this host correctly,
then generalising it to all hosts.**
- `/dev/shm` was 7.7 GB here, so a flag looked useless — but Home Assistant ignores `shm_size`, so
elsewhere it is Docker's 64 MB default and removing the flag reintroduces a crash loop.
- An MCP entry was identified by its URL — but that URL is the documented default, so the rule
would have deleted a user's hand-written configuration.
- A GPU probe created a hardware context — but that proved the driver worked, not that Chromium's
GPU path did.
The pattern is always *inference standing in for detection*. Before changing a default, ask what
this is like on a host unlike yours. Prefer detecting the condition at runtime over asserting it.
When ownership matters, **record it rather than infer it**.
## "Merged and inert" — CI passing proves the build works, not that the change does anything
Both changes in the session that produced this skill passed CI, merged, and were **inert**:
- The Xvfb resolution cap wrote its env file correctly and Xvfb still started at the base-image
default — wrong env mechanism for that service.
- The GPU flags reached Chromium's command line exactly as intended, and the GPU process still
reported `--use-gl=disabled`, having overridden them after its own init failed.
Once the rebuilt add-on is running, re-run the measurement that motivated the work. Some fixes
cannot be self-verified — a service that reads its environment only at start makes an env-var fix
unproven until the add-on restarts, which needs the user or `ha-cli` with their agreement. If you
cannot restart, the change is **Assumed**, not Verified, and must be reported that way.

View File

@@ -0,0 +1,30 @@
# Simplify — case studies
Evidence for why the mechanism ladder in SKILL.md step 3 exists, and why levels 4-6 need a reason
that survives being said out loud. In all three cases the simpler option existed and was skipped.
Being able to build the complicated thing is not a reason to.
- A rejected PR spent a **388-line TCP proxy plus a 142-line monkeypatch of a private upstream
method** to reclaim 159 MB — placing custom transport code in the path of every API request.
Both independent reviewers said close it rather than iterate on it.
- A ~180-line `ctypes` probe was written to decide whether to enable GPU flags. It worked
perfectly, proved the driver was fine, and the change **still did nothing**, because the
question it answered was not the question that mattered.
- A resolution cap shipped as a **new init script writing an s6 envdir** — the wrong mechanism
entirely (ladder level 4). Renaming the option to the env var the service already reads
(level 1) would have worked, and the new script did not.
## Checks worth running against your own diff
- **Did the diff stay at the ladder level chosen in step 3?** If it crept up a level, either
justify that out loud or redo it at the level you chose.
- **Can this be solved by deleting instead of adding?** A flag that shouldn't be passed, a
process that shouldn't start, a registration that shouldn't be duplicated. Deleting usually
shrinks the regression surface — but not always: the `/dev/shm` case in
`references/evidence.md` is a removal that reintroduced a crash loop on hosts unlike this one.
A removal that depends on a host default still needs the same verification as an addition.
- **Is the fix bigger than the thing it fixes?** That is a smell, not a rule — but it usually
means the problem was framed one level too deep.
- **How does this fail in three years**, when the base image, Electron, or upstream has moved?
Code that reads a documented knob keeps working. Code that reaches into private internals
does not.

View File

@@ -0,0 +1,218 @@
# Repo-specific traps
Things that look correct and are not. Each cost real time or shipped broken. Read this before
implementing; skim the headings, read the ones you're about to touch.
The repo's own `CLAUDE.md` documents structure, Dockerfile conventions, `updater.json`, CI
workflows and lint rules — that is not repeated here.
## Contents
- [Environment and workspace](#environment-and-workspace)
- [Measurement](#measurement)
- [Passing values into base-image services](#passing-values-into-base-image-services)
- [Shell and bashio](#shell-and-bashio)
- [Dockerfile and architecture](#dockerfile-and-architecture)
- [Versioning](#versioning)
- [Chromium / Electron under Xvfb](#chromium--electron-under-xvfb)
- [CI and review bots](#ci-and-review-bots)
---
## Environment and workspace
**The checkout is probably on the wrong branch.** Checkouts under `/data/claude` are shared and
persistent; another session leaves them wherever it finished. A stale branch looks entirely
normal. Compare the add-on's `config.yaml` `version` against the running `$BUILD_VERSION` before
trusting anything you read. `scripts/preflight.sh` does this.
**Never run `git stash` under `/data/claude`.** `refs/stash` is shared across every worktree and
concurrent session, so it is *not* isolated even in your own worktree. A bare `stash` / `stash
pop` pair in a clean worktree once restored another session's stash, producing conflict markers
in six untouched files. To compare a file against another revision use
`git show <rev>:<path> > /tmp/x`. If a pop does go wrong: a conflicted pop **keeps** the stash
entry, so nothing is lost — confirm `git rev-parse HEAD` matches what you pushed, then
`git reset --hard HEAD`.
**Work in a worktree under `/data`, not `/tmp`** — `/tmp` is `noexec`, so scripts there won't run.
```bash
git worktree add --detach /data/claude/.work/<task> origin/master
```
**You cannot test the Docker build.** dockerd does not start in this environment. CI is the only
gate. One observed run took ~3 hours, with 20+ runs queued against 2 executing — that was account
runner contention, not the diff. Check `gh run list` before concluding your PR is stuck. Poll in
a background task, and never claim the build is verified when it hasn't run.
## Measurement
**Summed RSS overstates savings.** Shared library pages are counted once per process, so removing
a duplicate frees its *private* memory, not its RSS. Measured example: four MCP shims summed to
882 MB RSS but 643 MB PSS / 564 MB private, and per-process private ranged 54 MB down to 2 MB —
which completely changes which duplicate is worth removing. Quote private when arguing "removing
this saves N MB".
**A large mapping is often not resident.** SysV/tmpfs segments are lazily populated. Xvfb's
506 MB framebuffer shows `Rss: 0` in `/proc/<pid>/smaps`. Check before calling anything a leak.
**`/proc/meminfo` and `free` show host figures** — there is no memory cgroup namespace here.
Never attribute those totals to the add-on.
**`rtk` filters some command output.** For a complete listing, redirect to a file and read that
(`ps ... > $SP/ps.txt`), or use `rtk proxy <cmd>`.
## Passing values into base-image services
The plumbing has four stages. `scripts/env_trace.sh <VAR> <process>` walks all four and tells
you which one drops the value — use it rather than reasoning about this from memory.
1. `/data/options.json` — the user's saved options.
2. **Injected export block** — `.templates/00-global_var.sh` writes a literal
`export <option>='<value>'` block into *every* service `run` script, using the option name
**verbatim**. So `max_resolution` *is* injected; it just isn't a name any service reads.
`MAX_RES` would be both injected and read.
3. `container_environment` — s6's envdir, read **only** by services whose shebang is
`#!/usr/bin/with-contenv`.
4. The running process — the only stage that decides behaviour.
**Name the option exactly as the env var the service reads** (uppercase), the way `DRINODE`,
`KEYBOARD` and `TZ` already do. Verified live: `DRINODE` appears as `export DRINODE=…` in all 16
service run scripts including `svc-xorg`, and Xvfb runs with `-vfbdevice /dev/dri/renderD128`.
**Two consequences that are easy to get wrong:**
- `00-global_var.sh` is cont-init **00**. Any cont-init script numbered higher runs *after* the
injection, so it cannot change what a service will see through stage 2.
- LSIO's `svc-xorg` starts `#!/usr/bin/env bashio`, **not** `with-contenv`, so it never reads
stage 3 at all. Writing `container_environment` for it is a silent no-op — that shipped: the
file was written 6 seconds before Xvfb started, and Xvfb still came up at the base-image
default.
**Renaming an option to match a base-image env var moves validation out of your script and into
the schema.** `00-global_var.sh` exports empty strings (only objects/arrays/nulls are dropped),
and base-image scripts typically test `${VAR+x}` — *set*-ness, not emptiness. So an empty
`MAX_RES` becomes `Xvfb -screen 0 "x24"` and the X server does not start. If you make this move,
constrain the value in `config.yaml` (`match(^[0-9]{1,5}x[0-9]{1,5}$)?`) in the same commit, or
keep a guard script.
**Open question, unresolved:** what Supervisor does with a stored `options.json` key that no
longer exists in the new schema — error, warn, or silently drop. This decides whether renaming an
option is safe on upgrade. The `monica` add-on shipped exactly such a rename
(`MEILISEARCH_KEY` → `meilisearch_key`) with no migration, which is weak evidence it is
tolerated. The base image has an `init-migrations` oneshot reading `/migrations` if a migration
is needed. Confirm before renaming a shipped option.
Whichever mechanism you use, verify the service actually received it:
```bash
tr '\0' '\n' < /proc/<pid>/environ | grep <VAR>
```
**`cont-init.d` runs as root before s6 services start** — that part is true and is the right
place for filesystem and permission setup.
**Anything needing an X display must not run in `cont-init.d`** — Xvfb isn't up yet. Put it in
the openbox autostart. (ANGLE's OpenGL backend, for instance, fails with "Could not open the
default X display".)
## Shell and bashio
**`bashio::config` for lists**: `while read ... < <(bashio::config ...)` silently yields an empty
list under errexit — bashio's internals return non-zero and process substitution inherits the
failure. Capture with `$(...)` first, then feed a here-string.
**Scripts shared by symlink**: `80-configuration.sh` and friends are shared with the webtop
add-ons. Put add-on-specific logic in a new numbered script instead of editing them.
**`grep -E '^…$'` anchors per line.** A multi-line config value passes validation on its first
line and is then used verbatim. Use bash's `[[ =~ ]]`, which anchors the whole string.
## Dockerfile and architecture
**Prefer `BUILD_ARCH` over `TARGETARCH`** — the repo's builder passes `BUILD_ARCH` explicitly,
while `TARGETARCH` is BuildKit-provided and may or may not be populated.
Either way the variable must be declared with `ARG <NAME>` **in the build stage that uses it**;
without that it expands empty, the guard never matches, and the block silently does nothing —
which is the same dead-`if` failure described just below, and the usual cause of it.
**Verify a guarded block actually ran** rather than assuming. Check whether its payload exists in
the running image (`command -v <tool>`), and cross-check `/var/log/apt/history.log` for the
matching `apt-get install` line. An `if` block whose condition never matched leaves no trace and
no error — one such block sat dead for weeks while appearing to guarantee driver verification.
**Don't test for distro-specific filenames.** A guard on
`/usr/share/vulkan/icd.d/intel_icd.x86_64.json` named a file Debian does not ship (it installs
`intel_icd.json`), so fixing the arch variable alone would have turned dead code into a failing
build.
## Versioning
**`X.Y.Z.N`, never `X.Y.Z-N`.** A hyphen parses as a semver pre-release, which Supervisor treats
as *older* than `X.Y.Z` — the update is never offered.
Date-based versions (`2026.08.03`) are common here. Check whether master has already moved to the
version you were about to use.
## Chromium / Electron under Xvfb
**Xvfb offers only indirect/software GLX**, so Chromium probes it, fails, and falls back to CPU
rendering — the GPU process runs `--use-gl=disabled` and the renderer `--disable-gpu-compositing`.
**Passing ANGLE flags is not sufficient.** `--ozone-platform=x11 --use-gl=angle
--use-angle=gl-egl` reached Chromium's command line exactly as intended and the GPU process
*still* reported `--use-gl=disabled`, having overridden the flag after its own init failed.
**A standalone ANGLE probe proves less than it appears to.** Loading Claude Desktop's bundled
`libEGL.so`, initializing the OpenGL backend and reading back
`ANGLE (Intel, Mesa Intel(R) Graphics (ADL-N), OpenGL 4.6)` proves the driver and device work —
not that Chromium's GPU process, sandbox, dmabuf import and X11 presentation path work. Codex
flagged this distinction during review and was right.
`--use-angle=gles-egl` is rejected outright by Mesa ("Intel or NVIDIA OpenGL ES drivers are not
supported").
**`--disable-dev-shm-usage`** is a workaround for Docker's 64 MB default `/dev/shm`. Home
Assistant **ignores** the add-on's `shm_size`, so the real size varies per install — it was 7.7 GB
on one host. Detect at runtime rather than assuming either way; keep the flag when the size
cannot be determined, because the crash it prevents is worse than its overhead.
## CI and review bots
**What CI actually gates** — checked against the workflows, because assuming costs a cycle.
All three hard gates below are matrixed over `check-addon-changes.outputs.changedAddons` and
`if:`-skipped when it is `[]`, so a PR that touches no add-on directory (docs, `.github/`,
`.claude/`) shows them as *skipping*, not passing — do not read that as a green build.
- **CHANGELOG updated** — hard gate (`onpr_check-pr.yaml` exits 1 without it).
- **HA add-on linter** — hard gate: `frenck/action-addon-linter` in `onpr_check-pr.yaml` has no
`continue-on-error`, so a config.yaml schema error fails the PR.
- **Add-on image build** — hard gate, and slow; one run took ~3 h.
- **Weekly super-linter** — `lint.yml` runs with `continue-on-error: true` at both call sites, so
it *cannot* fail a PR. Fix real findings anyway, but do not treat it as a blocker.
- **Version bump** — no workflow checks it. It is repo convention, and required for Supervisor to
offer the rebuild, but it will not fail CI.
**CI rewrites your shell scripts.** `lint.yml` runs `shfmt -w -i 4 -ci -bn -sr` over every `*.sh`
and `run`, plus a `chmod +x` pass, on schedule. Repo-wide reformatting commits land on master
without your involvement — another reason a shared checkout goes stale mid-task.
**Reviewers**: CodeRabbit (deepest — often runs scripts to prove a claim; reviews ~9 minutes
after the PR opens, or on `@coderabbitai review`), chatgpt-codex-connector, Copilot, Codacy.
**Codacy `action_required` is this repo's normal state.** Other open PRs show the same. It
exposes no annotations via the API, so its findings are only visible in the maintainer's Codacy
account. Note it and move on rather than guessing.
**Resolving a review thread requires GraphQL** (`resolveReviewThread`); the REST API cannot do it.
`scripts/pr_review.sh` wraps fetch / reply / resolve.
**The repo's `.markdownlint.yaml` does not disable MD022/MD032**, so a CHANGELOG will show
dozens of pre-existing heading/list findings. They are noise because lint is `continue-on-error`,
not because the config exempts them — don't cite the config as a reason to ignore a finding.
**Separate new lint findings from pre-existing ones** by linting the same file at `origin/master`
and diffing the result sets — otherwise you chase warnings that were already there.
`scripts/validate.sh --vs-master` does this.

View File

@@ -0,0 +1,136 @@
#!/usr/bin/env bash
# Trace one env var through the whole add-on plumbing, to answer "I set the option and nothing
# happened".
#
# This is the single highest-value diagnostic for this repo, because the plumbing has four
# separate stages and a value can be present at stage 3 and absent at stage 4 while every script
# involved reports success. That exact case shipped: MAX_RES was correctly written to
# /run/s6/container_environment/MAX_RES six seconds before Xvfb started, and Xvfb still came up
# at the base-image default.
#
# The four stages:
# 1. /data/options.json the user's saved add-on options
# 2. injected export block .templates/00-global_var.sh writes `export <option>='<v>'`
# into every service run script — as cont-init 00, i.e. BEFORE
# any higher-numbered cont-init script can influence it
# 3. container_environment s6's envdir, read only by services using `with-contenv`
# 4. the running process the only stage that actually matters
#
# Two consequences worth internalising:
# * A cont-init.d script numbered >00 cannot change what stage 2 injected.
# * A service starting `#!/usr/bin/env bashio` (LSIO's svc-xorg does) never reads stage 3, so
# writing container_environment for it is a silent no-op.
#
# Usage: env_trace.sh <VAR> [process-name-or-pid]
# env_trace.sh MAX_RES Xvfb
# env_trace.sh DRINODE Xvfb # a working example, for comparison
set -uo pipefail
VAR="${1:?usage: env_trace.sh <VAR> [process-name-or-pid]}"
TARGET="${2:-}"
# VAR is interpolated into grep/sed patterns below — restrict it to a valid env var name
case "$VAR" in
[A-Za-z_]*) [ -z "${VAR//[A-Za-z0-9_]/}" ] || { echo "invalid env var name: $VAR" >&2; exit 1; } ;;
*) echo "invalid env var name: $VAR" >&2; exit 1 ;;
esac
echo "== tracing ${VAR} =="
echo
echo "1. /data/options.json (the user's saved options)"
if [ -f /data/options.json ]; then
python3 - "$VAR" <<'PY'
import json, sys
var = sys.argv[1]
try:
opts = json.load(open('/data/options.json'))
except Exception as err:
print(f" could not parse: {err}"); raise SystemExit
hit = {k: v for k, v in opts.items() if k.lower() == var.lower()}
if hit:
for k, v in hit.items():
shown = '<empty string>' if v == '' else repr(v)
print(f" {k} = {shown}")
if k != var:
print(f" NOTE: option is named '{k}', not '{var}' — the injected export uses the")
print(f" option name verbatim, so a service reading ${var} will not see it.")
else:
print(f" absent (so the add-on default from config.yaml applies, if any)")
PY
else
echo " /data/options.json not present (not running as an add-on?)"
fi
echo
echo "2. injected 'ADDON ENV' export block in service run scripts"
found=0
for d in /etc/s6-overlay/s6-rc.d /etc/services.d; do
[ -d "$d" ] || continue
while IFS= read -r rs; do
if grep -qE "^export ${VAR}=" "$rs" 2> /dev/null; then
echo " $(grep -E "^export ${VAR}=" "$rs" | head -1) <- $rs"
found=1
fi
done < <(find "$d" -name run -type f 2> /dev/null)
done
[ "$found" -eq 0 ] && echo " ${VAR} not injected into any service run script"
echo
echo "3. s6 container_environment (only read by services using #!/usr/bin/with-contenv)"
seen3=0
for d in /var/run/s6/container_environment /run/s6/container_environment; do
if [ -f "$d/$VAR" ]; then
echo " $d/$VAR = [$(cat "$d/$VAR")]"; seen3=1
fi
done
[ "$seen3" -eq 0 ] && echo " not present in either envdir"
echo
echo "4. the running process (the only stage that decides behaviour)"
if [ -z "$TARGET" ]; then
echo " no target given; pass a process name or pid as \$2"
else
# Prefer an exact process-name match. A -f substring match picks up this script's own shell
# (its command line contains the name you searched for), which produces confusing noise.
if [ -d "/proc/$TARGET" ]; then
pids="$TARGET"
else
pids=$(pgrep -x "$TARGET" 2> /dev/null | head -3)
[ -z "$pids" ] && pids=$(pgrep -f "$TARGET" 2> /dev/null | grep -vE "^($$|$PPID)$" | head -3)
fi
if [ -z "$pids" ]; then
echo " no process matching '$TARGET'"
else
for pid in $pids; do
comm=$(tr -d '\0' < "/proc/$pid/comm" 2> /dev/null)
val=$(tr '\0' '\n' < "/proc/$pid/environ" 2> /dev/null | sed -n "s/^${VAR}=//p")
if [ -n "$val" ]; then
echo " pid=$pid ($comm): ${VAR}=[$val]"
else
echo " pid=$pid ($comm): ${VAR} NOT SET"
# Naming the shebang is usually the whole answer.
for d in /etc/s6-overlay/s6-rc.d /etc/services.d; do
while IFS= read -r rs; do
if grep -qiE "exec .*${comm}|${comm}" "$rs" 2> /dev/null; then
echo " its service $rs starts: $(head -1 "$rs")"
head -1 "$rs" | grep -q with-contenv \
&& echo " -> uses with-contenv, so stage 3 WOULD reach it" \
|| echo " -> NOT with-contenv, so stage 3 can never reach it"
break 2
fi
done < <(find "$d" -name run -type f 2> /dev/null)
done
fi
# What it was actually launched with beats any theory about its environment.
tr '\0' '\n' < "/proc/$pid/cmdline" 2> /dev/null | tail -n +2 |
grep -iE "res|screen|${VAR}" | head -3 | sed 's/^/ argv: /'
done
fi
fi
echo
echo "== reading the result =="
echo " present at 4 -> the value reached the process; the bug is elsewhere"
echo " at 1+2 but not 4 -> service started before injection, or reads a different name"
echo " at 1+3 but not 2 or 4 -> classic silent no-op: wrong mechanism for this service"
echo " at 1 only -> option name does not match any env var a service reads"

View File

@@ -0,0 +1,99 @@
#!/usr/bin/env bash
# RAM/CPU snapshot of the running add-on, built to avoid the two mistakes that make such
# snapshots wrong:
#
# 1. Summed RSS double-counts shared pages. Removing a duplicate process frees its *private*
# memory, not its RSS. So PSS and private are reported alongside, and private is the number
# to quote when arguing "removing this saves N MB".
# 2. A big mapping is not necessarily resident. Large SysV/tmpfs segments are lazily populated,
# so reserved size is reported separately from resident.
#
# Note /proc/meminfo and free show HOST figures (no memory cgroup namespace) — never attribute
# those to the add-on.
#
# Usage: measure.sh [cpu-sample-seconds] (default 20)
set -uo pipefail
SAMPLE="${1:-20}"
if [ "$SAMPLE" -lt 20 ]; then
echo "WARNING: a ${SAMPLE}s sample understates CPU badly (a 3s sample measured 2.3% where" >&2
echo " 20s measured 21.6% for the same process). Use >=20s for anything you report." >&2
fi
OUT="${SCRATCH:-${TMPDIR:-/tmp}}/addon-measure.$$"
mkdir -p "$OUT"
# rtk filters some output; redirect to a file to get the complete list.
ps -eo pid,ppid,user,rss,pcpu,etimes,args --sort=-rss > "$OUT/ps.txt" 2>&1
echo "== totals =="
awk 'NR>1{s+=$4; n++} END{printf " processes=%d summed RSS=%.0f MB (overstates: shared pages counted per-process)\n", n, s/1024}' "$OUT/ps.txt"
awk '{t+=$2} END{printf " threads=%d\n", t}' <(ps -eo pid,nlwp --no-headers 2> /dev/null)
echo
echo "== per-process memory (top 20 by PSS) =="
printf ' %-28s %8s %8s %8s\n' COMMAND RSS PSS PRIVATE
python3 - "$OUT" <<'PY'
import os, sys
rows = []
for pid in filter(str.isdigit, os.listdir('/proc')):
try:
cmd = open(f'/proc/{pid}/cmdline', 'rb').read().replace(b'\x00', b' ').decode(errors='replace').strip()
if not cmd:
continue
rss = pss = priv = 0
for line in open(f'/proc/{pid}/smaps_rollup'):
k, _, v = line.partition(':')
v = v.split()[0] if v.split() else '0'
if k == 'Rss': rss = int(v)
elif k == 'Pss': pss = int(v)
elif k in ('Private_Dirty', 'Private_Clean'): priv += int(v)
except Exception:
continue
rows.append((pss, rss, priv, pid, cmd))
rows.sort(reverse=True)
tr = tp = tv = 0
for pss, rss, priv, pid, cmd in rows:
tr += rss; tp += pss; tv += priv
for pss, rss, priv, pid, cmd in rows[:20]:
name = (cmd[:26] + '..') if len(cmd) > 28 else cmd
print(f" {name:<28} {rss/1024:7.0f}M {pss/1024:7.0f}M {priv/1024:7.0f}M")
print(f"\n {'TOTAL':<28} {tr/1024:7.0f}M {tp/1024:7.0f}M {tv/1024:7.0f}M")
print(" ^ quote PRIVATE when claiming what removing a process would free.")
PY
echo
echo "== reserved-but-not-resident (lazy allocations, NOT leaks) =="
ipcs -m 2>/dev/null | awk 'NR>3 && $5 ~ /^[0-9]+$/ && $5 > 50000000 {printf " SysV shm %.0f MB (owner %s) — check Rss in /proc/<pid>/smaps before calling it used\n", $5/1048576, $3}'
echo
echo "== CPU over ${SAMPLE}s (idle unless you are driving the UI) =="
# utime+stime. Parsed after the LAST ')' because field 2 is (comm) and may contain spaces —
# a plain $14+$15 is wrong for anything like 'npm exec @foo' and silently reports a fabricated
# number rather than failing.
jiffies() { awk -F') ' '{n=split($NF,a," "); print a[12]+a[13]}' "/proc/$1/stat" 2>/dev/null; }
# jiffies are USER_HZ units — almost always 100, but read it rather than assume it
HZ=$(getconf CLK_TCK 2>/dev/null) && [ "$HZ" -gt 0 ] 2>/dev/null || HZ=100
# Sample EVERY readable process, not the top-N of ps.txt: that list is sorted by RSS,
# and the busiest process is not necessarily a big one.
declare -A t0
for d in /proc/[0-9]*; do
pid=${d#/proc/}
[ -r "$d/stat" ] && t0[$pid]=$(jiffies "$pid")
done
sleep "$SAMPLE"
for pid in "${!t0[@]}"; do
[ -r "/proc/$pid/stat" ] || continue
t1=$(jiffies "$pid") || continue
[ -n "$t1" ] && [ -n "${t0[$pid]}" ] || continue
delta=$(( t1 - ${t0[$pid]} ))
[ "$delta" -gt 0 ] || continue
pct=$(awk -v d="$delta" -v s="$SAMPLE" -v hz="$HZ" 'BEGIN{printf "%.2f", d*100/(hz*s)}')
comm=$(tr -d '\0' < "/proc/$pid/comm" 2>/dev/null)
echo "$pct $pid $comm"
done | sort -rn | head -12 | awk '{printf " %6s%% %-8s %s\n", $1, $2, $3}'
echo
echo " established conns on :8082/:3000/:3001 = $(ss -tn 2>/dev/null | grep -cE 'ESTAB.*:(8082|3000|3001)')"
echo " (those are claude_desktop/webtop viewer ports; 0 here means CPU above is idle burn)"
echo " raw ps: $OUT/ps.txt"

View File

@@ -0,0 +1,112 @@
#!/usr/bin/env bash
# Work through bot review comments on a PR. Resolving a thread needs the GraphQL API (the REST
# API cannot do it), which is the only reason this script exists.
#
# pr_review.sh list <PR> every inline comment, grouped
# pr_review.sh status <PR> checks + unresolved thread count
# pr_review.sh reply <PR> <COMMENT_ID> <text|@file>
# pr_review.sh resolve <PR> <THREAD_ID...|--all> --all = every unresolved, asks first
# pr_review.sh watch <PR> [minutes] poll checks (run this backgrounded)
#
# Reviewers seen here: coderabbitai (deepest; reviews ~9 min after open, or on
# "@coderabbitai review"), chatgpt-codex-connector, Copilot, Codacy.
#
# Verify every claim before agreeing. Bots are frequently right and occasionally confidently
# wrong; a reproduction takes a minute and decides it either way. Push back with evidence when
# you are right — a resolved-but-wrong thread is worse than an open one.
set -uo pipefail
REPO="${HASSIO_REPO:-}"
[ -z "$REPO" ] && REPO=$(gh repo view --json nameWithOwner --jq .nameWithOwner 2> /dev/null)
[ -z "$REPO" ] && { echo "cannot determine repo; set HASSIO_REPO=owner/name" >&2; exit 1; }
echo "repo: $REPO" >&2
CMD="${1:-}"; PR="${2:-}"
[ -z "$CMD" ] || [ -z "$PR" ] && { sed -n '2,16p' "$0" | sed 's/^# \?//'; exit 1; }
case "$CMD" in
list)
echo "== inline comments on #$PR =="
gh api "repos/$REPO/pulls/$PR/comments" --paginate \
--jq 'sort_by(.created_at)[] | "=== [\(.id)] \(.user.login) | \(.path):\(.line // .original_line) ===\n\(.body)\n"'
echo "== review bodies =="
gh api "repos/$REPO/pulls/$PR/reviews" \
--jq '.[] | select(.body != "") | "--- \(.user.login) (\(.state)) ---\n\(.body[0:4000])\n"'
;;
status)
gh pr checks "$PR" 2>&1 | head -15
echo
gh api graphql -f query="{repository(owner:\"${REPO%%/*}\",name:\"${REPO##*/}\"){pullRequest(number:$PR){reviewThreads(first:50){nodes{id isResolved path comments(first:1){nodes{author{login}}}}}}}}" \
--jq '.data.repository.pullRequest.reviewThreads.nodes[] | "\(if .isResolved then "resolved" else "OPEN " end) \(.id) \(.comments.nodes[0].author.login) \(.path)"'
;;
reply)
ID="${3:?comment id}"; BODY="${4:?text or @file}"
if [ "${BODY#@}" != "$BODY" ]; then
out=$(gh api "repos/$REPO/pulls/$PR/comments/$ID/replies" -F body=@"${BODY#@}" --jq '.id' 2>&1)
rc=$?
else
out=$(gh api "repos/$REPO/pulls/$PR/comments/$ID/replies" -f body="$BODY" --jq '.id' 2>&1)
rc=$?
fi
# Silently "succeeding" here is worse than failing: a later session reads the transcript and
# believes a reviewer was answered when they were not.
if [ "$rc" -eq 0 ] && [ -n "$out" ]; then
echo "replied to $ID (comment $out)"
else
echo "FAILED to reply to $ID: $out" >&2
echo " (top-level review bodies have different ids and cannot take replies here)" >&2
exit 1
fi
;;
resolve)
shift 2
ids="$*"
if [ "${ids:-}" = "--all" ]; then
echo "About to resolve EVERY unresolved thread. Only do this if you have read and"
echo "answered each one — a resolved-but-wrong thread is worse than an open one."
gh api graphql -f query="{repository(owner:\"${REPO%%/*}\",name:\"${REPO##*/}\"){pullRequest(number:$PR){reviewThreads(first:100){nodes{isResolved path comments(first:1){nodes{author{login} body}}}}}}}" \
--jq '.data.repository.pullRequest.reviewThreads.nodes[] | select(.isResolved==false) | " - \(.comments.nodes[0].author.login) \(.path): \(.comments.nodes[0].body[0:90])"'
printf 'Type "yes" to resolve all: '; read -r ok
[ "$ok" = "yes" ] || { echo "aborted"; exit 1; }
ids=""
elif [ -z "$ids" ]; then
echo "usage: pr_review.sh resolve <PR> <THREAD_ID...> (or --all, with confirmation)" >&2
echo "resolve each thread as you answer it; get ids from: pr_review.sh status $PR" >&2
exit 1
fi
if [ -z "$ids" ]; then
ids=$(gh api graphql -f query="{repository(owner:\"${REPO%%/*}\",name:\"${REPO##*/}\"){pullRequest(number:$PR){reviewThreads(first:50){nodes{id isResolved}}}}}" \
--jq '.data.repository.pullRequest.reviewThreads.nodes[] | select(.isResolved==false) | .id')
fi
[ -z "$ids" ] && { echo "nothing unresolved"; exit 0; }
rfail=0
for id in $ids; do
r=$(gh api graphql -f query="mutation{resolveReviewThread(input:{threadId:\"$id\"}){thread{isResolved}}}" \
--jq '.data.resolveReviewThread.thread.isResolved' 2>&1)
echo " $id -> $r"
[ "$r" = "true" ] || rfail=1
done
# exiting 0 on a failed mutation would let a session believe threads were resolved
exit "$rfail"
;;
watch)
MINS="${3:-180}" # the addon build alone has taken ~3h; 20 was far too short
for i in $(seq 1 "$MINS"); do
c=$(gh pr checks "$PR" 2> /dev/null | awk '{print $1"="$2}' | tr '\n' ' ')
if [ -z "$c" ]; then
# Normal in the first minutes after `gh pr create`, and also whenever gh errors.
# Calling that "settled" would report success for checks that never ran.
echo "[$i] no checks reported yet (gh returned nothing) — still waiting"
sleep 60; continue
fi
echo "[$i] $c"
case "$c" in
*pending*) sleep 60 ;;
*fail* | *error* | *cancel*) echo "settled — with FAILURES (see above)"; wfail=1; break ;;
*) echo "settled — all passing"; break ;;
esac
done
echo "note: long queues here are usually account runner contention, not your diff."
exit "${wfail:-0}"
;;
*) echo "unknown: $CMD"; exit 1 ;;
esac

View File

@@ -0,0 +1,87 @@
#!/usr/bin/env bash
# Orient before starting add-on work: what tools exist, are we inside the running add-on, and
# — the one that actually bites — does the checkout match what is running?
#
# Usage: preflight.sh [repo-path] [addon-slug]
set -uo pipefail
REPO="${1:-/data/claude/hassio-addons}"
SLUG="${2:-}"
echo "== tools =="
for c in gh git codex rtk headroom tokensave shellcheck hadolint yamllint python3 jq; do
printf ' %-11s %s\n' "$c" "$(command -v "$c" > /dev/null 2>&1 && echo yes || echo MISSING)"
done
[ -x /data/codex/bin/codex-real ] && echo " codex-real yes (prefer: codex exec --model gpt-5.6-sol)"
echo
echo "== running add-on =="
if [ -n "${BUILD_VERSION:-}" ]; then
echo " BUILD_VERSION=$BUILD_VERSION HOME=${HOME:-?}"
echo " -> live measurement is possible; see measure.sh"
else
echo " not inside a running add-on (no BUILD_VERSION); source-only analysis"
fi
echo
echo "== repo =="
# git-aware check: in a worktree .git is a file, not a directory
if ! git -C "$REPO" rev-parse --git-dir > /dev/null 2>&1; then
echo " no git repo at $REPO"
exit 0
fi
cd "$REPO" || exit 0
branch=$(git branch --show-current 2> /dev/null || echo "(detached)")
echo " path=$REPO"
echo " branch=$branch"
# Another session may be mid-operation in this shared checkout.
echo " recent reflog (entries you did not make mean another session is active):"
git reflog --date=iso -3 2> /dev/null | sed 's/^/ /'
# The trap this exists for: a stale branch looks entirely normal.
if [ -n "${BUILD_VERSION:-}" ]; then
# Hostname is <8-hex>-<slug-with-dashes>. Anchor the hex to 8 chars: an unanchored
# [0-9a-f]* also eats real prefixes (dab-radio -> radio, cafe-monitor -> monitor).
# Slugs may legitimately contain dashes (birdnet-go), so try both forms.
if [ -z "$SLUG" ] && [ -n "${HOSTNAME:-}" ]; then
base=$(printf '%s' "$HOSTNAME" | sed 's/^[0-9a-f]\{8\}-//')
for cand in "$(printf '%s' "$base" | tr '-' '_')" "$base"; do
[ -f "$REPO/$cand/config.yaml" ] && { SLUG="$cand"; break; }
done
[ -z "$SLUG" ] && SLUG="$base"
fi
cfg="$REPO/$SLUG/config.yaml"
if [ ! -f "$cfg" ]; then
echo
echo " could not find $SLUG/config.yaml — pass the slug as \$2 to enable the"
echo " revision check (this is the check the script exists for)." >&2
exit 3
fi
if [ -f "$cfg" ]; then
here=$(grep -E '^version:' "$cfg" | head -1 | tr -d "\"'" | awk '{print $2}')
echo
echo " $SLUG/config.yaml version = $here"
echo " running image BUILD_VERSION = $BUILD_VERSION"
if [ "$here" = "$BUILD_VERSION" ]; then
echo " MATCH — this checkout corresponds to the running image."
else
echo " MISMATCH — this branch is NOT what is running."
git fetch origin master --quiet 2> /dev/null
master=$(git show origin/master:"$SLUG/config.yaml" 2> /dev/null |
grep -E '^version:' | head -1 | tr -d "\"'" | awk '{print $2}')
echo " origin/master version = ${master:-unknown}"
echo " -> work from origin/master; analysing this branch will mislead you."
echo
echo "== suggested isolated worktree (/tmp is noexec; use /data) =="
echo " git worktree add --detach /data/claude/.work/<task> origin/master"
echo " NOTE: never 'git stash' under /data/claude — refs/stash is shared."
exit 2
fi
fi
fi
echo
echo "== suggested isolated worktree (/tmp is noexec; use /data) =="
echo " git worktree add --detach /data/claude/.work/<task> origin/master"
echo " NOTE: never 'git stash' under /data/claude — refs/stash is shared across worktrees."

View File

@@ -0,0 +1,128 @@
#!/usr/bin/env bash
# Run every linter that CI will run and that works locally. The Docker build is deliberately not
# attempted: dockerd does not start in this environment, so CI is the only gate for it — say that
# rather than implying the build was checked.
#
# --vs-master re-lints each changed file at origin/master and prints only findings your diff
# ADDED. Without it you will chase warnings that were already in the file.
#
# Usage: validate.sh [addon-dir] [--vs-master]
set -uo pipefail
ADDON="${1:-}"
[ "${ADDON:-}" = "--vs-master" ] && { ADDON=""; set -- --vs-master; }
VS_MASTER=false
for a in "$@"; do [ "$a" = "--vs-master" ] && VS_MASTER=true; done
if [ -z "$ADDON" ]; then
mapfile -t _dirs < <(git diff --name-only origin/master...HEAD 2> /dev/null |
cut -d/ -f1 | sort -u | grep -vE '^\.' )
if [ "${#_dirs[@]}" -gt 1 ]; then
echo "several changed dirs: ${_dirs[*]}"
echo "pass one explicitly: validate.sh <addon-dir>"; exit 1
fi
ADDON="${_dirs[0]:-}"
fi
[ -z "$ADDON" ] && { echo "usage: validate.sh <addon-dir> [--vs-master]"; exit 1; }
git rev-parse --verify origin/master > /dev/null 2>&1 || {
echo "origin/master missing — run: git fetch origin master"; exit 1; }
export PYTHONDONTWRITEBYTECODE=1
echo "== validating $ADDON =="
fail=0
note() { printf ' %-13s %s\n' "$1" "$2"; }
# Shell: bash -n then shellcheck -x (follows sourced files, as CI does).
while IFS= read -r f; do
[ -f "$f" ] || continue
if ! out=$(bash -n "$f" 2>&1); then note "bash -n" "FAIL $f"; echo "$out" | sed 's/^/ /'; fail=1; fi
done < <(find "$ADDON" -type f \( -name '*.sh' -o -name 'run' -o -name 'finish' -o -name 'autostart' \) 2> /dev/null)
[ "$fail" -eq 0 ] && note "bash -n" "ok"
if command -v shellcheck > /dev/null 2>&1; then
sc=$(find "$ADDON" -type f \( -name '*.sh' -o -name 'autostart' -o -name 'run' -o -name 'finish' \) -print0 2> /dev/null |
xargs -0 -r shellcheck -x -f gcc 2>&1)
if [ -n "$sc" ]; then
note "shellcheck" "$(printf '%s\n' "$sc" | grep -c .) finding(s)"
printf '%s\n' "$sc" | sed 's/^/ /' | head -20
else note "shellcheck" "clean"; fi
fi
command -v hadolint > /dev/null 2>&1 && [ -f "$ADDON/Dockerfile" ] && {
hl=$(hadolint "$ADDON/Dockerfile" 2>&1)
[ -n "$hl" ] && { note "hadolint" "$(printf '%s\n' "$hl" | grep -c .) finding(s)"; printf '%s\n' "$hl" | sed 's/^/ /' | head -10; } || note "hadolint" "clean"
}
if [ -f "$ADDON/config.yaml" ]; then
# path passed as argv, never interpolated into Python source
python3 - "$ADDON/config.yaml" <<'PY' || { note "config.yaml" "FAIL parse"; fail=1; }
import yaml,sys
d=yaml.safe_load(open(sys.argv[1]))
print(' %-13s ok (version=%s, %d options)' % ('config.yaml', d.get('version'), len(d.get('options') or {})))
missing=[k for k in (d.get('options') or {}) if k not in (d.get('schema') or {})]
if missing: print(' %-13s options with no schema entry: %s' % ('WARN', missing)); sys.exit(0)
PY
command -v yamllint > /dev/null 2>&1 && {
yl=$(yamllint -f parsable "$ADDON/config.yaml" 2>&1 | grep -c .)
note "yamllint" "$yl finding(s) (compare with --vs-master)"
}
fi
while IFS= read -r f; do
python3 -m py_compile "$f" 2> /dev/null || { note "py_compile" "FAIL $f"; fail=1; }
done < <(find "$ADDON" -type f -name '*.py' 2> /dev/null)
$VS_MASTER && command -v npx > /dev/null 2>&1 && [ -f "$ADDON/CHANGELOG.md" ] && {
md=$(npx --yes markdownlint-cli2 "$ADDON/CHANGELOG.md" 2>&1 | grep -cE "CHANGELOG.md:[0-9]+")
note "markdownlint" "$md finding(s) in CHANGELOG (mostly pre-existing; lint is continue-on-error in CI)"
}
echo
echo "== CI requirements =="
if git diff --name-only origin/master...HEAD 2> /dev/null | grep -q "$ADDON/CHANGELOG.md"; then
note "CHANGELOG" "updated"
else
# This one IS gated: onpr_check-pr.yaml exits 1 without it.
note "CHANGELOG" "NOT UPDATED — this is the one CI hard-gate"; fail=1
fi
if git diff origin/master...HEAD -- "$ADDON/config.yaml" 2> /dev/null | grep -q '^+version:'; then
note "version" "bumped"
else
# Repo convention and required for the rebuild to be offered — but no workflow gates it,
# so this is a warning, not a failure.
note "version" "NOT bumped (convention; no rebuild will be offered) — not a CI gate"
fi
note "docker build" "NOT tested locally (dockerd unavailable) — CI is the only gate"
if $VS_MASTER; then
echo
echo "== findings ADDED by this diff (pre-existing ones filtered out) =="
tmp=$(mktemp -d); trap 'rm -rf "$tmp"' EXIT
git diff --name-only origin/master...HEAD -- "$ADDON" 2> /dev/null | while IFS= read -r f; do
git show "origin/master:$f" > "$tmp/base" 2> /dev/null || continue
# A missing linter must be a visible skip, not a silent "no new findings":
# its "command not found" error is identical for base and head, so comm would
# cancel it out and report a false clean.
case "$f" in
*.sh | *autostart | */run | */finish)
command -v shellcheck > /dev/null 2>&1 || { echo " $f: SKIPPED (shellcheck not installed)"; continue; }
cmd() { shellcheck -x -f gcc "$1" 2>&1 | sed 's/^[^:]*:[0-9]*:[0-9]*://'; } ;;
*.yaml | *.yml)
command -v yamllint > /dev/null 2>&1 || { echo " $f: SKIPPED (yamllint not installed)"; continue; }
cmd() { yamllint -f parsable "$1" 2>&1 | sed 's/^[^:]*//; s/^:[0-9]*:[0-9]*//'; } ;;
*Dockerfile)
command -v hadolint > /dev/null 2>&1 || { echo " $f: SKIPPED (hadolint not installed)"; continue; }
cmd() { hadolint "$1" 2>&1 | sed 's/^[^:]*//; s/^:[0-9]*//'; } ;;
*) continue ;;
esac
cp "$tmp/base" "$tmp/base_f"; b=$(cmd "$tmp/base_f" | sort)
a=$(cmd "$f" | sort)
new=$(comm -13 <(printf '%s\n' "$b") <(printf '%s\n' "$a") | grep -c .)
[ "$new" -gt 0 ] && { echo " $f: $new NEW finding(s)"; comm -13 <(printf '%s\n' "$b") <(printf '%s\n' "$a") | sed 's/^/ /' | head -5; }
done
echo " (nothing listed above = your diff introduced no new lint findings)"
fi
echo
[ "$fail" -eq 0 ] && echo "== local validation passed ==" || echo "== local validation FAILED =="
exit "$fail"

Binary file not shown.

Before

Width:  |  Height:  |  Size: 331 KiB

After

Width:  |  Height:  |  Size: 60 KiB

BIN
.github/stats.png vendored

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.9 KiB

After

Width:  |  Height:  |  Size: 1.8 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 10 KiB

After

Width:  |  Height:  |  Size: 4.1 KiB

View File

@@ -20,7 +20,7 @@ jobs:
pull-requests: write
steps:
- uses: actions/stale@v10
- uses: actions/stale@v11
with:
repo-token: ${{ secrets.GITHUB_TOKEN }}
stale-issue-message: 'This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.'

View File

@@ -91,7 +91,7 @@ jobs:
uses: actions/checkout@v7.0.1
- name: 🔎 Run Home Assistant Add-on Lint
uses: frenck/action-addon-linter@v2
uses: frenck/action-addon-linter@f6bef06a4cee6c67924b0a70be643aacb8500a43
with:
path: "./${{ matrix.addon }}"

View File

@@ -114,7 +114,7 @@ jobs:
steps:
- uses: actions/checkout@v7.0.1
- name: Run Home Assistant Add-on Lint
uses: frenck/action-addon-linter@v2
uses: frenck/action-addon-linter@f6bef06a4cee6c67924b0a70be643aacb8500a43
with:
path: "./${{ matrix.addon }}"

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.4 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.5 KiB

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.5 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.3 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.9 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.7 KiB

After

Width:  |  Height:  |  Size: 1.7 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

View File

@@ -1,3 +1,7 @@
## 2026.08.02 (02-08-2026)
- Fix: ingress returned "502 Bad Gateway" because Caddy never listened on :8082. `91-nginx_ingress.sh` hooked the ingress site into `update_caddyfile.sh` with a sed anchored on `sudo caddy fmt --overwrite`, but 2026.07.10-1 strips `sudo` from every BirdNET-Pi script at build time, so the anchor stopped matching. `update_caddyfile.sh` then rewrote the Caddyfile from scratch just before Caddy started, dropping the ingress site
- Fix: `caddy_ingress.sh` no longer appends a second `:8082` block when it runs twice (a duplicate site address makes Caddy refuse to start)
- Fix: `02-caddy.sh` re-adds the ingress site if it is missing from the Caddyfile just before starting Caddy
## 2026.07.22 (22-07-2026)
- Fix: health-check the WebUI port (8081) instead of port 80, so the standalone Docker container no longer reports "unhealthy" when ssl=false
- Fix: health-check now probes https when ssl is enabled, and no longer silently reports "healthy" regardless of the actual result (the previous check's `&>` redirection is a bash-ism that dash, the image's /bin/sh, parses as background + no-op, discarding curl's exit status)

View File

@@ -116,5 +116,5 @@ tmpfs: true
udev: true
url: https://github.com/alexbelgium/hassio-addons/tree/master/birdnet-pi
usb: true
version: 2026.07.22
version: 2026.08.02
video: true

View File

@@ -25,5 +25,24 @@ export TZ="${TZ_VALUE:-Etc/UTC}"
# Update caddyfile with password
"$HOME"/BirdNET-Pi/scripts/update_caddyfile.sh &> /dev/null || true
# update_caddyfile.sh rewrites the Caddyfile from scratch. 91-nginx_ingress.sh
# hooks the ingress site back into it, but if that hook ever fails to apply,
# caddy would start without a :8082 listener and ingress would answer 502.
# 91-nginx_ingress.sh writes /ingress_url when ingress is on and removes it when
# it is off, so this is a no-op in standalone mode.
# Require the Caddyfile to exist: if it is missing something went badly wrong
# earlier, and caddy failing on a missing config is easier to diagnose than an
# ingress-only Caddyfile conjured up here.
if [[ -f /ingress_url ]] && [[ -f /etc/caddy/Caddyfile ]] \
&& ! grep -qE '^[[:space:]]*:8082[[:space:]]*\{' /etc/caddy/Caddyfile 2> /dev/null; then
echo "Ingress site missing from the Caddyfile, re-adding it"
if ! /helpers/caddy_ingress.sh; then
# Start caddy anyway: ingress stays broken, but direct access on 8081
# keeps working. Exiting here would only make s6 restart this service in
# a loop and take the WebUI down completely.
echo "Failed to re-add the ingress site, the ingress panel will return 502" >&2
fi
fi
echo "Starting service: caddy"
exec /usr/bin/caddy run --config /etc/caddy/Caddyfile

View File

@@ -26,6 +26,10 @@ ingress_entry=$(bashio::addon.ingress_entry)
if [[ "$ingress_entry" != "/api"* ]]; then
bashio::log.info "Ingress entry is not set, exiting configuration."
sed -i "1a sleep infinity" /custom-services.d/02-nginx.sh
# This script re-runs from /etc/scripts-init on an add-on restart, so drop
# any marker left by an earlier run: it is what 02-caddy.sh reads to decide
# whether the ingress site belongs in the Caddyfile.
rm -f /ingress_url
exit 0
fi
@@ -68,9 +72,23 @@ sed -i "s|localhost|localhost:8082|g" "$HOME/BirdNET-Pi/scripts/utils/notificati
# Update the Caddyfile if update script exists
caddy_update_script="$HOME/BirdNET-Pi/scripts/update_caddyfile.sh"
if [ -f "$caddy_update_script" ]; then
sed -i "/sudo caddy fmt --overwrite/i /helpers/caddy_ingress.sh" "$caddy_update_script"
else
if [ ! -f "$caddy_update_script" ]; then
bashio::log.error "Caddy update script not found: $caddy_update_script"
exit 1
fi
# update_caddyfile.sh rewrites /etc/caddy/Caddyfile from scratch, which drops
# the ingress site added just above. 02-caddy.sh runs it right before starting
# caddy, so the hook below has to re-add the site from inside that script, just
# before it formats and reloads the config.
# The anchor must not require "sudo": the Dockerfile strips it from every
# BirdNET-Pi script at build time, so the shipped line is "caddy fmt --overwrite".
if ! grep -qF "/helpers/caddy_ingress.sh" "$caddy_update_script"; then
sed -i -E "/^[[:space:]]*(sudo[[:space:]]+)?caddy[[:space:]]+fmt[[:space:]]+--overwrite/i /helpers/caddy_ingress.sh" "$caddy_update_script"
fi
# sed is silent when the anchor is missing; make sure the hook is really there
if ! grep -qF "/helpers/caddy_ingress.sh" "$caddy_update_script"; then
bashio::log.warning "Could not anchor the ingress site in $caddy_update_script, appending it instead"
printf '\n/helpers/caddy_ingress.sh\n' >> "$caddy_update_script"
fi

View File

@@ -6,6 +6,13 @@ set +u
# shellcheck disable=SC1091
source /etc/birdnet/birdnet.conf
# Nothing to do if the ingress site is already there. This script runs both from
# cont-init and from update_caddyfile.sh, and a duplicate ":8082" site address
# makes caddy refuse to start.
if grep -qE '^[[:space:]]*:8082[[:space:]]*\{' /etc/caddy/Caddyfile 2> /dev/null; then
exit 0
fi
# Create ingress configuration for Caddyfile
cat << EOF >> /etc/caddy/Caddyfile
:8082 {

Binary file not shown.

Before

Width:  |  Height:  |  Size: 4.5 KiB

After

Width:  |  Height:  |  Size: 1.8 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.7 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.4 KiB

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.8 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.5 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

View File

@@ -1,3 +1,86 @@
## 2026.08.04 (04-08-2026)
- Fix: reverted the GPU acceleration added in 2026.08.03. It did not just fail to help — it was
what disabled the GPU. `--use-gl=angle --use-angle=gl-egl` forces Mesa's EGL X11 platform,
which offers no window-capable EGLConfig under this Xvfb, so the GPU process logged
`gl_surface_egl.cc:262 No suitable EGL configs found`, abandoned GL, and was relaunched with
`--use-gl=disabled` while every renderer got `--disable-gpu-compositing`.
Chromium already renders on the GPU here with no flags at all, because LSIO's Xvfb runs with
`-vfbdevice /dev/dri/renderD128` and its GLX is therefore backed by the real render node —
the premise that Xvfb offers only a software path was wrong for this base image. Measured
three times, including at the production 15360x8640 screen: with the flags the GPU process
loads `libEGL_mesa` and holds 1 fd on the render node; without them it loads `libGLX_mesa`,
holds 8, and no renderer carries `--disable-gpu-compositing`.
Removes `claude-gpu-probe` and the `gpu_acceleration` option. The probe was not wrong about
the hardware, it was answering the wrong question: it exercised ANGLE's default GLX path,
which works, and so it passed while the flags it gated disabled the GPU.
- Fix: `max_resolution` never did anything; the option is renamed to `MAX_RES`. It wrote
`MAX_RES` into the s6 `container_environment`, but the base image's `svc-xorg` starts
`#!/usr/bin/env bashio` rather than `with-contenv` and never reads that directory, so Xvfb
kept starting at 15360x8640. Naming the option `MAX_RES` makes the add-on env layer inject
it directly into every service `run` script, which is how `DRINODE` already reaches Xvfb.
`MAX_RES` ships with no default, so nothing changes for anyone until it is set: the screen
stays at the base image's 15360x8640. The old `max_resolution` default of `1920x1080` is
deliberately not carried over — it had never taken effect, so shipping it now would have
silently shrunk every existing desktop, including on 4K displays. Home Assistant drops
options that are no longer in the schema (it logs a warning), so a saved `max_resolution`
is discarded rather than migrated. The schema bounds each axis to 100-9999, which keeps the
worst case (~400 MB of framebuffer) below the 15360x8640 default it replaces.
Removes `22-display_tuning.sh`.
- Removed the amd64 GPU driver install, which has never run in any release. The 2026.08.03
build corrected its gate variable but the guard is `if [[ ... ]]` and there is no `SHELL`
directive, so it runs under dash, which has no `[[` — the condition was false and the `RUN`
still exited 0. Verified in the shipped 2026.08.03 image: no `vainfo`, no
`intel-media-va-driver-non-free`, and no matching install in its apt history.
It was deleted rather than repaired, because it had nothing to add. Hardware OpenGL already
works without it: the GPU process loads the Mesa gallium megadriver over DRI3 and holds 8
fds on `/dev/dri/renderD128`, with no swrast/llvmpipe in the GL path. `libgl1-mesa-dri` and
`mesa-vulkan-drivers` already come from the LSIO base at a newer backports version (Mesa
25.0.7) than reinstalling them would give. And `intel-media-va-driver-non-free`, the one
genuinely new payload, was installed live on the running add-on and changed nothing: VA-API
failed identically to the free driver (`iHD_drv_video.so init failed`) on both
`/dev/dri/renderD128` and `card0`. That failure is below the add-on, in the host i915 stack.
Repairing the guard would have turned dead code into an untested apt transaction that swaps
a base-image package for no measured benefit.
- A `[[` in the Dockerfile's legacy `/etc/services.d` shim is corrected to `[` for the same
dash reason; that directory does not exist on this s6-overlay v3 base, so it is
behaviour-neutral here.
## 2026.08.03 (03-08-2026)
- Performance: Claude Desktop now uses the GPU instead of rendering on the CPU.
Under Xvfb, Chromium probed GLX, found only Xvfb's software path, and fell back to
`--use-gl=disabled` + `--disable-gpu-compositing`, making the renderer the add-on's largest
CPU consumer. It is now launched with `--ozone-platform=x11 --use-gl=angle --use-angle=gl-egl`,
but only when the new `claude-gpu-probe` confirms Claude Desktop's own bundled ANGLE can
create a hardware GL context on this host. New `gpu_acceleration` option (`auto`/`on`/`off`,
default `auto`); any probe failure keeps the previous software rendering unchanged.
- Performance: new `max_resolution` option (default `1920x1080`) caps the virtual screen via the
base image's `MAX_RES`. Xvfb previously ran at 15360x8640, so Xvfb and the Selkies capture
loop tracked damage over a 133-megapixel area continuously, even with no browser connected.
Selkies still resizes dynamically below the cap.
- Performance: Claude Code now talks to the Home Assistant MCP server over its native HTTP
transport instead of through the `mcp-proxy` stdio bridge, removing one Python process per
Claude Code session (~45 MB of private resident memory each). Claude Desktop keeps the
bridge, as the config schema for a remote entry there is not confirmed.
- Performance: new `mcp_servers_desktop` / `mcp_servers_code` options select which MCP servers
each client registers. Every stdio MCP server is a separate process per client, and Desktop
starts another full set per Claude Code session it hosts. Defaults register all of them in
both clients, i.e. the previous behaviour.
- Fix: the Dockerfile's Intel graphics block was dead code. It was gated on `TARGETARCH`, which
the repo's builder does not pass, so it never ran: the shipped amd64 image has no `vainfo` and
no `intel-media-va-driver-non-free`. It now uses `BUILD_ARCH`, and its Vulkan ICD check no
longer names `intel_icd.x86_64.json`, a file Debian does not ship.
- Fix: stale Home Assistant MCP registrations (and the bearer token in them) are now removed
from Claude Code's config when `enable_ha_mcp` is turned off. Ownership of an HTTP entry is
recorded when the add-on writes it, so a manually configured `homeassistant` server is never
claimed, overwritten, or deleted — not even when it sits on the default URL.
- `--disable-dev-shm-usage` is now applied only when `/dev/shm` is actually small (under
256 MB). It is a workaround for Docker's 64 MB default, but Home Assistant ignores the add-on's
`shm_size` so the real size varies per install; the size is now read at startup, keeping the
crash workaround where it is needed and dropping it where it only pushed Chromium's shared
memory into ordinary files. If the size cannot be determined, the flag is kept.
## 2026.08.02 (02-08-2026)
- Minor bugs fixed
## ubunturesolute-version-3a10bef7 (2026-08-01)
- Update to latest version from linuxserver/docker-baseimage-selkies (changelog : https://github.com/linuxserver/docker-baseimage-selkies/releases)

View File

@@ -74,7 +74,7 @@ VOLUME [ "/sys/fs/cgroup" ]
# hadolint ignore=SC2015,DL4006,SC2013,SC2086
RUN \
usermod --home /data/data abc && \
if [[ -d /etc/services.d ]] && ls /etc/services.d/*/run 1> /dev/null 2>&1; then sed -i "1a set +e" /etc/services.d/*/run; fi
if [ -d /etc/services.d ] && ls /etc/services.d/*/run 1> /dev/null 2>&1; then sed -i "1a set +e" /etc/services.d/*/run; fi
ARG TEMPLATE_BASE_URL="https://raw.githubusercontent.com/alexbelgium/hassio-addons/master/.templates"
@@ -142,24 +142,29 @@ RUN install -d -m 0755 /etc/apt/keyrings && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
# The Intel N150/Twin Lake iGPU uses the host's i915 kernel driver through the mapped
# /dev/dri nodes. Explicitly install the amd64 userspace stack needed for accelerated
# OpenGL rendering, VA-API video encoding, and Vulkan, then fail the build if any driver
# payload is missing. Keep aarch64 unchanged because these Intel packages are amd64-only.
RUN if [[ "${TARGETARCH}" == "amd64" ]]; then \
apt-get update && \
apt-get install -y --no-install-recommends \
intel-media-va-driver-non-free \
libgl1-mesa-dri \
mesa-vulkan-drivers \
vainfo && \
test -f /usr/lib/x86_64-linux-gnu/dri/iHD_drv_video.so && \
test -f /usr/lib/x86_64-linux-gnu/dri/iris_dri.so && \
test -f /usr/share/vulkan/icd.d/intel_icd.x86_64.json && \
command -v vainfo > /dev/null; \
fi && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
# No bespoke GPU driver install. The LSIO base image already provides the userspace stack that
# hardware rendering actually uses, and it does so at a newer version than adding it here would:
# libgl1-mesa-dri and mesa-vulkan-drivers come from bookworm-backports Mesa 25.0.7, where a
# plain `apt-get install` of the same names is at best a no-op.
#
# There used to be an amd64 block here installing intel-media-va-driver-non-free, libgl1-mesa-dri,
# mesa-vulkan-drivers and vainfo. It never ran in any release: it was guarded by
# `if [[ ... ]]`, and with no SHELL directive RUN executes under /bin/sh — dash, which has no
# `[[`. dash printed `[[: not found`, the condition was false, and the RUN still exited 0, so
# every build silently skipped it and still passed. (An earlier fix corrected the gate variable
# from TARGETARCH to BUILD_ARCH but left the `[[`, so it stayed dead.) Confirmed in the shipped
# 2026.08.03 amd64 image: no vainfo, no intel-media-va-driver-non-free, no matching apt history.
#
# It was removed rather than repaired, because measurement showed it had nothing to add:
# - Hardware OpenGL already works without it. In the shipped image the Chromium GPU process
# loads the Mesa gallium megadriver over DRI3 and holds 8 fds on /dev/dri/renderD128, with
# no swrast/llvmpipe in the GL path.
# - The only genuinely new payload, intel-media-va-driver-non-free, does not fix anything
# here. Installed live alongside vainfo on the running add-on, VA-API failed identically to
# the free driver — `iHD_drv_video.so init failed` on both /dev/dri/renderD128 and card0,
# with both 23.1.1 builds. The failure is below the add-on, in the host i915/GuC stack.
# Repairing the guard would therefore have turned dead code into an untested apt transaction
# that swaps a base-image package for no measured benefit.
# Install the current upstream hadolint and actionlint releases for both supported
ARG HADOLINT_VERSION=v2.14.0

View File

@@ -97,6 +97,7 @@ Git synchronization hooks. A repository is indexed only when it is listed in
| `KEYBOARD` | | Optional Selkies keyboard layout. |
| `PASSWORD` | | Optional password for direct Selkies ports. |
| `DRINODE` | | Optional GPU device override for Selkies. |
| `MAX_RES` | _(unset)_ | Optional cap on the virtual screen, as `WIDTHxHEIGHT` (100-9999 per axis). Unset means the base image default, 15360x8640 — Selkies resizes dynamically below whatever the cap is, so this only sets the ceiling. Named `MAX_RES` because that is the environment variable the base image's Xvfb service reads. Setting it lowers the area Xvfb and the Selkies capture loop track for damage; the framebuffer itself is lazily populated, so this is a CPU saving, not a memory one. |
| `DNS_server` | `8.8.8.8` | DNS server used by the standard DNS module. |
| `permission_mode` | `auto` | Claude Code permission policy: `strict`, `auto`, or `bypass`. |
| `install_headroom` | `true` | Register Headroom MCP and run the supervised local proxy. |
@@ -106,6 +107,8 @@ Git synchronization hooks. A repository is indexed only when it is listed in
| `install_rtk` | `true` | Configure RTK's Claude Code `PreToolUse` Bash hook. |
| `install_tokensave` | `true` | Install TokenSave's complete global Claude integration. |
| `tokensave_project_paths` | `[]` | Explicit absolute Git repository paths to initialize or sync at startup. |
| `mcp_servers_desktop` | all | Which managed MCP servers Claude Desktop registers (`headroom`, `tokensave`, `homeassistant`, `codex`). |
| `mcp_servers_code` | all | Which managed MCP servers Claude Code registers. Each stdio server is a separate process per client, and Desktop starts another set per Claude Code session it hosts, so trimming this is the cheapest way to cut memory. |
| `install_caveman` | `false` | Install the third-party Caveman Claude Code plugin at startup. |
| `install_codex_cli` | `false` | Install the latest stable OpenAI Codex CLI at startup and register its native MCP server so Claude can delegate work to ChatGPT Codex. |
| `codex_sandbox_mode` | `workspace-write` | Filesystem scope Codex runs with: `read-only`, `workspace-write`, or `danger-full-access`. |

View File

@@ -65,6 +65,16 @@ options:
install_headroom: true
install_rtk: true
install_tokensave: true
mcp_servers_desktop:
- headroom
- tokensave
- homeassistant
- codex
mcp_servers_code:
- headroom
- tokensave
- homeassistant
- codex
permission_mode: auto
tokensave_project_paths: []
panel_admin: false
@@ -86,6 +96,7 @@ schema:
data_location: str?
DRINODE: list(/dev/dri/card0|/dev/dri/card1|/dev/dri/card2|/dev/dri/renderD128|/dev/dri/renderD129|)?
KEYBOARD: list(da-dk-qwerty|de-de-qwertz|en-gb-qwerty|en-us-qwerty|es-es-qwerty|fr-ch-qwertz|fr-fr-azerty|it-it-qwerty|ja-jp-qwerty|pt-br-qwerty|sv-se-qwerty|tr-tr-qwerty)?
MAX_RES: match(^[1-9][0-9]{2,3}x[1-9][0-9]{2,3}$)?
PASSWORD: str?
PGID: int
PUID: int
@@ -115,12 +126,15 @@ schema:
install_headroom: bool
install_rtk: bool
install_tokensave: bool
mcp_servers_desktop:
- list(headroom|tokensave|homeassistant|codex)
mcp_servers_code:
- list(headroom|tokensave|homeassistant|codex)
permission_mode: list(strict|auto|bypass)
tokensave_project_paths:
- str
slug: claude_desktop
tmpfs: true
udev: true
url: https://github.com/alexbelgium/hassio-addons
version: "ubunturesolute-version-3a10bef7"
version: "2026.08.04"
video: true

View File

@@ -20,4 +20,46 @@
# Headroom is intentionally not injected into the Desktop process: Claude Desktop
# force-overrides ANTHROPIC_BASE_URL (headroom #869), so Desktop uses the registered Headroom
# MCP tools instead.
exec claude-desktop --no-sandbox --disable-dev-shm-usage --password-store=basic
# GPU acceleration: deliberately no flags. Do not add ANGLE/EGL flags here.
#
# Chromium already renders on the GPU on this base image, because LSIO's Xvfb is started with
# `-vfbdevice /dev/dri/renderD128` — so GLX here is backed by the real render node, not the
# indirect/software path. Left alone, Chromium picks Mesa's GLX (libGLX_mesa), initialises the
# GPU process, and composites on the GPU.
#
# Passing `--use-gl=angle --use-angle=gl-egl` actively *broke* that. It forces Mesa's EGL X11
# platform, which offers no window-capable EGLConfig under this Xvfb, so the GPU process logged
# ui/gl/gl_surface_egl.cc:262 No suitable EGL configs found.
# gave up on GL entirely, and Chromium relaunched it with `--use-gl=disabled` while stamping
# `--disable-gpu-compositing` on every renderer — i.e. the flags caused the CPU rendering they
# were meant to remove. Measured three times, including at the production 15360x8640 screen:
# with the flags the GPU process loads libEGL_mesa and holds 1 fd on the render node; with no
# flags it loads libGLX_mesa and holds 8, and no renderer carries --disable-gpu-compositing.
#
# A standalone ANGLE probe is not evidence for any of this: Claude Desktop's bundled ANGLE
# happily creates a hardware context here via its default (GLX) path, which is why the probe
# that used to gate these flags passed while the flags themselves disabled the GPU.
#
# `--use-angle=gl` also works, but it only reproduces what Chromium already chooses by itself,
# so it is not passed either.
# Shared memory.
#
# --disable-dev-shm-usage exists because Docker's default /dev/shm is 64 MB, which is not
# enough for Chromium's renderers and produced a crash loop here (see the 1.3 changelog entry).
# Home Assistant ignores the add-on's shm_size, so the size cannot be set from this repo and
# genuinely varies between installs — it is 7.7 GB on some hosts and the 64 MB default on
# others. Hardcoding either answer is wrong, so ask the kernel: keep the workaround when
# /dev/shm is small, and drop it when there is plenty, where it would otherwise push Chromium's
# shared memory into ordinary files for no benefit. If the size cannot be determined, keep the
# flag — the crash it prevents is worse than the overhead it costs.
SHM_FLAGS="--disable-dev-shm-usage"
SHM_KB="$(df -k /dev/shm 2> /dev/null | awk 'NR==2 {print $2}')"
if [ -n "$SHM_KB" ] && [ "$SHM_KB" -ge 262144 ]; then
SHM_FLAGS=""
fi
# SHM_FLAGS must stay unquoted so it expands to a separate argument, or to nothing at all.
# shellcheck disable=SC2086
exec claude-desktop --no-sandbox --password-store=basic $SHM_FLAGS

View File

@@ -308,10 +308,13 @@ if bashio::config.true 'install_codex_cli'; then
fi
HA_MCP_ENABLED=false
HA_MCP_URL=""
HA_MCP_TOKEN=""
# Read unconditionally, even when enable_ha_mcp is off. Home Assistant keeps an option's value
# when its toggle is disabled, and the reconciliation below needs this URL to recognise the
# HTTP entry it previously wrote so that it can be removed — together with the bearer token
# inside it — rather than orphaned in ~/.claude.json.
HA_MCP_URL="$(bashio::config 'ha_mcp_url' 'http://homeassistant:8123/api/mcp')"
if bashio::config.true 'enable_ha_mcp'; then
HA_MCP_URL="$(bashio::config 'ha_mcp_url' 'http://homeassistant:8123/api/mcp')"
if bashio::config.has_value 'ha_mcp_token'; then
HA_MCP_TOKEN="$(bashio::config 'ha_mcp_token')"
fi
@@ -325,12 +328,73 @@ if bashio::config.true 'enable_ha_mcp'; then
fi
fi
# Which of the managed MCP servers each client gets.
#
# Every stdio MCP server is a separate process *per client*, and Claude Desktop starts another
# full set for each Claude Code session it hosts — so a server registered in both clients is
# paid for several times over. Measured on a live add-on with three sets running, the private
# (non-shared) resident cost was roughly 54 MB per extra `headroom mcp serve`, 45 MB per extra
# `mcp-proxy`, and only ~11 MB and ~2 MB for `codex` and `tokensave`, which share most of their
# pages. Registering a server only where it is actually used is therefore the cheapest lever
# available; these two options expose that choice.
#
# Defaults keep every enabled server in both clients, i.e. the pre-existing behaviour.
MCP_ALL_SERVERS="headroom tokensave homeassistant codex"
mcp_client_list() {
local option="$1"
local selected=() entry raw rc=0
# An option that is absent entirely — i.e. an existing install upgrading from a config that
# predates these options — keeps the previous behaviour of registering every enabled server.
# bashio distinguishes this from an explicitly empty list: an unset key yields the literal
# "null", while `[]` yields an empty string. Those must not be conflated, because an empty
# list is a legitimate way to say "no MCP servers in this client" and defaulting it back to
# all four would silently ignore the user.
if ! bashio::config.exists "$option"; then
echo "$MCP_ALL_SERVERS"
return 0
fi
# Capture first: reading a bashio list straight into `while read` via process substitution
# silently yields nothing under this script's errexit. The exit status is kept separately
# so that a failed read is not mistaken for a deliberate empty selection.
raw="$(bashio::config "$option" 2> /dev/null)" || rc=$?
if [ "$rc" -ne 0 ]; then
bashio::log.warning "Could not read '${option}'; registering every enabled MCP server for this client"
echo "$MCP_ALL_SERVERS"
return 0
fi
while read -r entry; do
[ -n "$entry" ] || continue
# Reconciliation deletes any managed server not named here, so an unrecognised value
# must never be treated as an authoritative selection.
case " $MCP_ALL_SERVERS " in
*" $entry "*) selected+=("$entry") ;;
*)
bashio::log.warning "Ignoring unknown MCP server '${entry}' in '${option}'; registering every enabled server for this client"
echo "$MCP_ALL_SERVERS"
return 0
;;
esac
done <<< "$raw"
echo "${selected[@]:-}"
}
MCP_SERVERS_DESKTOP="$(mcp_client_list 'mcp_servers_desktop')"
MCP_SERVERS_CODE="$(mcp_client_list 'mcp_servers_code')"
bashio::log.info "MCP servers for Claude Desktop: ${MCP_SERVERS_DESKTOP}"
bashio::log.info "MCP servers for Claude Code: ${MCP_SERVERS_CODE}"
HEADROOM_ENABLED="$HEADROOM_ENABLED" HEADROOM_BIN="$(command -v headroom || echo headroom)" \
HEADROOM_HF_HOME="${HOME}/.headroom/hf" \
TOKENSAVE_ENABLED="$TOKENSAVE_ENABLED" TOKENSAVE_BIN="$(command -v tokensave || echo tokensave)" \
CODEX_ENABLED="$CODEX_ENABLED" CODEX_BIN="$CODEX_BIN" CODEX_SANDBOX_MODE="$CODEX_SANDBOX_MODE" \
HA_MCP_ENABLED="$HA_MCP_ENABLED" HA_MCP_URL="$HA_MCP_URL" HA_MCP_TOKEN="$HA_MCP_TOKEN" \
MCP_PROXY_BIN="$(command -v mcp-proxy || echo mcp-proxy)" \
MCP_SERVERS_DESKTOP="$MCP_SERVERS_DESKTOP" MCP_SERVERS_CODE="$MCP_SERVERS_CODE" \
CLAUDE_DESKTOP_CONFIG="$CLAUDE_DESKTOP_CONFIG" CLAUDE_CODE_CONFIG="$CLAUDE_CODE_CONFIG" \
python3 - <<'PY' || bashio::log.warning "Unable to update the MCP server registrations automatically"
import json
@@ -385,6 +449,23 @@ if os.environ["HA_MCP_ENABLED"] == "true":
"env": {"API_ACCESS_TOKEN": os.environ["HA_MCP_TOKEN"]},
}
# Claude Code speaks Streamable HTTP MCP natively, so pointing it straight at Home Assistant
# removes the mcp-proxy bridge process entirely — it exists only to translate stdio to the HTTP
# transport Home Assistant already serves. That bridge was the most expensive duplicate
# measured (~45 MB of private RSS per copy, one per Claude Code session).
#
# Claude Desktop keeps the stdio bridge. Its bundled MCP SDK does contain a remote transport,
# but the shape `claude_desktop_config.json` accepts for a remote entry — and whether it
# persists a static bearer header — could not be confirmed, and a wrong guess would silently
# break Home Assistant access in Desktop. Revisit once that schema is verified upstream.
HA_MCP_CODE_ENTRY = None
if os.environ["HA_MCP_ENABLED"] == "true":
HA_MCP_CODE_ENTRY = {
"type": "http",
"url": os.environ["HA_MCP_URL"],
"headers": {"Authorization": "Bearer " + os.environ["HA_MCP_TOKEN"]},
}
# An entry is add-on-managed when its command is one of our binaries living outside the
# persistent home. Matching on the basename (rather than the exact path recorded at write
# time) keeps entries updatable when a base-image upgrade moves the binary, while commands
@@ -392,17 +473,77 @@ if os.environ["HA_MCP_ENABLED"] == "true":
HOME_PREFIX = os.path.expanduser("~") + os.sep
def is_managed(name, entry):
# Ownership record for the HTTP Home Assistant entry.
#
# The stdio entries can be recognised on sight, because their `command` points at a binary this
# image installs outside $HOME. An HTTP entry has no such tell: it is just a URL plus a bearer
# header, and a user who configured `homeassistant` by hand — very plausibly at the same default
# http://homeassistant:8123/api/mcp — would be indistinguishable from ours. Inferring ownership
# from shape or URL would let this script delete or overwrite that entry, including their token.
#
# So ownership is recorded rather than guessed: the URL of an entry this script actually wrote is
# remembered here, and only an entry matching that record is ever modified or removed. Anything
# this script did not write is untouchable, whatever it looks like. The file holds no secrets —
# just the endpoint — but is written 0600 to match the configs it describes.
STATE_PATH = Path(os.path.expanduser("~")) / ".config" / "claude_desktop_addon" / "managed-mcp.json"
def load_state():
try:
state = json.loads(STATE_PATH.read_text())
return state if isinstance(state, dict) else {}
except Exception:
return {}
def save_state(state):
try:
STATE_PATH.parent.mkdir(parents=True, exist_ok=True)
STATE_PATH.write_text(json.dumps(state, indent=2, sort_keys=True) + "\n")
STATE_PATH.chmod(0o600)
except Exception:
# Losing the record only costs us the ability to clean up later; never fail the boot.
pass
def is_managed_http_ha(entry, owned_url):
"""True only for an HTTP entry this script previously wrote."""
return (
bool(owned_url)
and entry.get("url") == owned_url
and set(entry) == {"type", "url", "headers"}
and entry.get("type") == "http"
and isinstance(entry.get("headers"), dict)
and set(entry["headers"]) == {"Authorization"}
)
def is_managed(name, entry, owned_url):
if not isinstance(entry, dict):
return False
command = entry.get("command")
if not isinstance(command, str) or command.startswith(HOME_PREFIX):
# A commandless entry is ours only when it is an HTTP Home Assistant registration this
# script recorded writing. Anything else — including a user's own remote server that
# reuses the name, even on the same URL — is left alone.
if command is None and name == "homeassistant":
return is_managed_http_ha(entry, owned_url)
return False
return os.path.basename(command) == MANAGED_BASENAMES[name]
SELECTED = {
"CLAUDE_DESKTOP_CONFIG": set(os.environ["MCP_SERVERS_DESKTOP"].split()),
"CLAUDE_CODE_CONFIG": set(os.environ["MCP_SERVERS_CODE"].split()),
}
state = load_state()
state_changed = False
for config_var, stdio_type in (("CLAUDE_DESKTOP_CONFIG", False), ("CLAUDE_CODE_CONFIG", True)):
path = Path(os.environ[config_var])
selected = SELECTED[config_var]
owned_url = state.get(str(path), {}).get("homeassistant_http_url", "")
try:
data = json.loads(path.read_text()) if path.exists() else {}
if not isinstance(data, dict):
@@ -417,17 +558,35 @@ for config_var, stdio_type in (("CLAUDE_DESKTOP_CONFIG", False), ("CLAUDE_CODE_C
changed = False
for name in MANAGED_BASENAMES:
existing = servers.get(name)
if name in desired:
entry = dict(desired[name])
if stdio_type:
entry["type"] = "stdio"
if existing is None or is_managed(name, existing):
if name in desired and name in selected:
if stdio_type and name == "homeassistant" and HA_MCP_CODE_ENTRY is not None:
# Claude Code talks to Home Assistant over HTTP directly; no bridge process.
entry = dict(HA_MCP_CODE_ENTRY)
else:
entry = dict(desired[name])
if stdio_type:
entry["type"] = "stdio"
# An entry we did not write is never overwritten, so a user's own HTTP
# `homeassistant` survives even when it sits on the configured URL.
claimable = existing is None or is_managed(name, existing, owned_url)
if claimable:
if existing != entry:
servers[name] = entry
changed = True
elif existing is not None and is_managed(name, existing):
if name == "homeassistant" and entry.get("type") == "http":
if owned_url != entry["url"]:
state.setdefault(str(path), {})["homeassistant_http_url"] = entry["url"]
owned_url = entry["url"]
state_changed = True
elif existing is not None and is_managed(name, existing, owned_url):
# Covers both "feature disabled" and "deselected for this client".
del servers[name]
changed = True
if name == "homeassistant" and state.get(str(path), {}).pop(
"homeassistant_http_url", None
):
owned_url = ""
state_changed = True
if changed:
if servers:
data["mcpServers"] = servers
@@ -440,6 +599,9 @@ for config_var, stdio_type in (("CLAUDE_DESKTOP_CONFIG", False), ("CLAUDE_CODE_C
# permissions between merges.
if path.exists():
path.chmod(0o600)
if state_changed:
save_state(state)
PY
# Guide Claude to actually use the Headroom compression tools so the MCP integration produces

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.5 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.5 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.6 KiB

After

Width:  |  Height:  |  Size: 1.7 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.9 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.8 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.5 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.6 KiB

After

Width:  |  Height:  |  Size: 1.7 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.9 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.4 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.4 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.3 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 4.1 KiB

After

Width:  |  Height:  |  Size: 1.8 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.9 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.8 KiB

After

Width:  |  Height:  |  Size: 1.8 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.4 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.8 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.7 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.6 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.7 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.9 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.4 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.5 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.6 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.3 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.7 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.6 KiB

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.9 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.7 KiB

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.1 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.8 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.2 KiB

After

Width:  |  Height:  |  Size: 1.6 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.4 KiB

After

Width:  |  Height:  |  Size: 1.5 KiB

Some files were not shown because too many files have changed in this diff Show More