Compare commits

...
Author SHA1 Message Date
Bailey Dixon 7dd1125686 chore(release): Android and plugin v1.4.0 (#181)
chore(release): Android and plugin v1.4.0
2026-07-09 23:06:03 -04:00
Bailey Dixon 522c4fe82a chore(release): finalize android-v1.4.0 and plugin-v1.4.0 2026-07-09 22:39:13 -04:00
Bailey Dixon c9b30dabae chore: sync main into dev before 1.4.0 release 2026-07-09 22:17:00 -04:00
Bailey Dixon e9e92d03f2 fix(voice): harden background session recovery 2026-07-09 22:16:48 -04:00
Bailey Dixon da8e23068a fix(voice): recover sessions after background route loss 2026-07-09 19:23:14 -04:00
Bailey Dixon aaee75e7fc docs: record realtime voice live verification 2026-07-09 17:26:20 -04:00
Bailey Dixon 8ebb21b16d fix(relay): dedupe background voice handoffs 2026-07-09 17:19:38 -04:00
Bailey Dixon 0700ac81c6 fix(relay): use exact xAI result delivery 2026-07-09 17:00:07 -04:00
Bailey Dixon 015298f90a fix(relay): make realtime session start idempotent 2026-07-09 16:29:22 -04:00
Bailey Dixon 3c0e51f664 fix(android voice): honor realtime model selection 2026-07-09 16:29:03 -04:00
Bailey DixonandClaude Fable 5 92f96831c4 fix(relay): let the voice agent recall an already-delivered result without re-running
After a background result was seeded into provider history, a pure-recall
follow-up ('what did that say?') still triggered a full hermes_run_task
round-trip instead of answering from history. Cause: _native_instructions
told the provider to re-route whenever context is 'tool-derived' -- which a
delivered background result is.

Rewrite the clause to separate recall from new work: a Hermes result already
delivered earlier in the conversation is in history, so recall/quote/reference
answers directly (no re-run); call hermes_run_task again only for new,
updated, deeper, or re-verified info. Drop the blanket tool-derived re-route,
keep 'fresh data or verification you don't already have -> re-route'.

Instruction-only. test_provider_native_instructions_include_recent_context
extended to assert the recall carve-out; route + promotion suites green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 15:15:33 -04:00
Bailey DixonandClaude Fable 5 033dcc37ba feat(relay): seed fallback-delivered background result into provider history
After a background result falls back to relay TTS (the provider deferred
instead of reading the answer), the provider's conversation history retained
only its own deferral -- so a follow-up ('what did that say?', 'can't you see
we ran the task?') had no record of the result and failed or re-ran. A
provider-VOICED delivery already becomes a history item; only fallbacks left
the gap. The existing native_pending_delivery_note is a one-shot correction
attached to the next Hermes-routed response and is skipped by follow-ups that
don't route through that branch.

Add RealtimeAgentConnection.append_context_item(role, text): a silent
conversation.item.create (assistant->'text', user/system->'input_text', no
response.create) implemented for xAI + OpenAI. On both fallback paths
(_finish_forced_summary_provider_response validator fallback,
_speak_fallback_answer provider-death/request-failed) the broker seeds the
delivered answer as an assistant turn, so any later follow-up finds it in
history durably, independent of routing. Best-effort (dead socket no-ops);
fallback-only, so a provider-voiced success is never double-recorded. Kept
the pending note as a belt-and-suspenders correction. Logs result_seeded /
result_seed_failed with a preview.

104 realtime tests green; test_filler_summary_triggers_fallback_delivery
extended to assert the seeded assistant history item; append_context_item
added to all connection fakes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 15:02:56 -04:00
Bailey DixonandClaude Fable 5 a144962e66 feat(relay): log provider-spoken delivery + preamble text for signoff diagnosis
Background-result delivery outcomes could not be told apart from the flight
recorder: the fallback event carried provider_text_preview, but the SUCCESS
paths (forced_summary_streaming early-commit, forced_summary_delivered
end-validated) logged only char counts and the pre-run acknowledgement
(hermes_forced_preamble.finished) logged only metadata. So a clean delivery
could not be confirmed verbatim, and a fallback could not be distinguished
from a validator false positive.

Add a _compact_status_text (<=120 char) preview of the actually-spoken text
to three existing _log payloads: transcript_preview (preamble),
prefix_preview (committed early-commit prefix), provider_text_preview
(end-validated delivery). Reuses the fallback path's existing compaction; no
new session state; bounded by the 14-day run-dir retention sweep.

Live payoff: confirmed think-fast spoke a genuine deferral ('One moment...
I'll let you know') rather than reading the answer -- a real fallback, not a
validator miss -- and the behaviour is model-agnostic across grok variants.

104 realtime tests green; test_realtime_summary_validation extended to assert
the delivered text rides prefix_preview / provider_text_preview.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 14:30:42 -04:00
Bailey DixonandClaude Fable 5 c683ad290c docs(todo): capture background-task UX asks + 2026-07-09 voice e2e findings
Records the owner's background-tasks-as-first-class-chat vision (titles,
kickoff/result chat entries, expand-to-detail, persisted results,
in-session provider context for follow-ups, concurrent multi-task) plus
today's live findings: first confirmed provider-voiced delivery (but
inconsistent on grok-voice-latest), think-fast untested until forced, a
duplicate-status bug, and a status-speak logging gap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 14:30:42 -04:00
Bailey DixonandClaude Fable 5 2968a173b1 fix(voice): render fallback deliveries in the voice overlay
A fallback-TTS delivery played audibly while the overlay sat on
"Thinking" — no waveform, no live text. Two client gates: the overlay
dropped every hermes-sourced voice.response.delta (a rule for mid-run
chatter, written before fallback TTS became a first-class delivery
mouth), and that handler was the only path flipping Thinking->Speaking,
so the Speaking-gated waveform envelope never fed.

- Relay: delivery responses now tagged on the wire (delivery:
  fallback/respeak/visual_only on started+delta events).
- App: hermes-sourced deltas with a delivery tag render (run chatter
  stays suppressed), and arriving output audio flips Thinking->Speaking
  so the waveform tracks any spoken response regardless of source.

116 realtime tests green with delivery-tag assertions on both fallback
paths; sideload debug build green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 22:14:34 -04:00
Bailey DixonandClaude Fable 5 5ff78da8e4 feat(relay): pre-RC observability hardening + realtime model updates
- Run-dir retention: sweep session JSONL logs + wav taps past
  realtime_voice.run_retention_days (default 14, 0 disables) at
  session-log creation — transcripts no longer accumulate indefinitely.
- Wav render tap is debug-only (debug_audio_tap, default off): artifact
  deleted after PCM streams, voice.response.done.audio_path blank.
- Delivery-outcome rollup: python -m plugin.relay.realtime_agent.report
  tallies provider-spoken vs fallback deliveries with reasons; new
  forced_summary_delivered marker makes end-validated deliveries countable.
- Models: OpenAI realtime default gpt-realtime-2 -> gpt-realtime-2.1
  (2.1-mini + rollback 2 selectable); xAI exposes the versioned
  grok-voice-think-fast-1.0 pin alongside the grok-voice-latest alias.
- Both providers surface the RESOLVED model id from session.created
  echoes (provider_model_resolved) so live rounds stay attributable
  across provider-side alias flips.

191 realtime/voice tests green, including new hygiene + report suites.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 20:48:10 -04:00
Bailey DixonandClaude Fable 5 ffe534454a docs: xAI voice platform re-baseline items (think-fast-1.0, alias flip, resumption)
grok-voice-latest now resolves to the new reasoning flagship
grok-voice-think-fast-1.0 (fast-1.0 deprecated); we default to the alias
everywhere, so live-round verdicts may predate the model change. Adds
re-baseline, lifecycle re-probe, and new-voices TODO items.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 20:26:37 -04:00
Bailey DixonandClaude Fable 5 49c183ef9a docs: OpenAI realtime next-RC roadmap, voice observability items, audit leftovers
- TODO: OpenAI realtime provider roadmap (model bump to gpt-realtime-2.1,
  live e2e verify, 60-min hard-cap handling, out-of-band exact delivery
  spike, async function-call delivery, tools guardrail) from the 2026-07-08
  research pass; full sourced findings in
  docs/plans/2026-07-08-openai-realtime-notes.md.
- TODO: voice observability pre-RC hardening (run-dir retention + wav-tap
  gating, delivery-outcome rollup, buffered flight-recorder writes).
- TODO: delivery-audit leftovers (respeak stays relay-TTS, exact-mode
  truncation cue) and post-audit hardening note on the live-verify item.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 19:42:55 -04:00
Bailey DixonandClaude Fable 5 906ce79a82 fix(relay): close five delivery-loss gaps found by adversarial audit
The provider-voiced delivery rework moved speak_verbatim from a synchronous
TTS render onto the async provider pipeline, opening failure windows the old
path couldn't have:

1. Foreground request_response was bare — a dead provider socket lost the
   answer and wedged native_forced_summary_active. Now falls back to the new
   shared _speak_fallback_answer relay-TTS mouth.
2. The delivery-confirm alarm only covered attached-background deliveries;
   foreground and deferred-resume injections now spawn it too.
3. A new user utterance mid-delivery wiped forced-summary state with no
   cancel and no record. New _preempt_pending_forced_summary cancels the
   stale response and lands a never-spoken answer as text.
4. Blocklist phrases present in the authoritative answer no longer flag a
   faithful exact reading; only model-added phrases count.
5. Structured JSON answers route to the summary prompt — no meaningful
   word-for-word reading exists for them.

Injection failure on the attached background path also now falls back to
spoken TTS immediately instead of waiting for the text-only alarm.

102 realtime tests green, five new covering each failure scenario.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 19:42:55 -04:00
Bailey DixonandClaude Fable 5 74f84d9492 docs: changelog + devlog for upstream-watch P1 batch
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 19:06:14 -04:00
Bailey Dixon c6bd418e30 Merge branch 'Codename-11/hrui-bootstrap-retire' into dev (HRUI-002) 2026-07-08 18:30:58 -04:00
Bailey Dixon 8b79c06922 Merge branch 'Codename-11/hrui-fallback-payloads' into dev (HRUI-001) 2026-07-08 18:30:57 -04:00
Bailey Dixon 7a63a0c6ea Merge branch 'Codename-11/hrui-prompt-submit-timeout' into dev (HRUI-016) 2026-07-08 18:30:44 -04:00
Bailey Dixon 83580cbf06 Merge branch 'Codename-11/hrui-manage-model-options' into dev (HRUI-022) 2026-07-08 18:30:43 -04:00
Bailey Dixon e636ca4715 Merge branch 'Codename-11/hrui-plugin-security' into dev (HRUI-014, HRUI-015) 2026-07-08 18:30:43 -04:00
Bailey DixonandClaude Fable 5 c1926f6434 fix(chat): align sessions/runs fallback payloads with upstream contract (HRUI-001)
Stop sending top-level messages/attachments fields the native upstream
session-chat and runs handlers never parse (silent data loss). Synthetic
phone-local history (voice intents, card dispatches, realtime voice
turns) now rides channels upstream actually consumes: tool-call pairs
render as a plain-text digest folded into the per-turn ephemeral system
prompt (system_message on sessions, instructions on runs, the system
message on completions); plain text turns splice into completions
messages and runs conversation_history. Attachments with no supported
channel are returned as ChatPayloadResult.droppedAttachments and logged
(the ChatViewModel user notice already existed) — never silently
discarded. Docs truth-up: HERMES-WEBAPI-REFERENCE session-chat body now
documents the native contract; decisions.md card-dispatch ADR gains an
HRUI-001 update note.

Verified: 18/18 HermesChatPayloadsTest unit tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 18:29:53 -04:00
Bailey DixonandClaude Fable 5 b78fe0244c feat(voice): provider-voiced exact result delivery
speak_verbatim no longer renders through relay TTS directly: every spoken
result_delivery mode now delivers through the realtime provider so
background/foreground Hermes answers keep the session's voice and tone.
Exact instructs a word-for-word reading of the authoritative answer
(_forced_hermes_exact_prompt); Summary keeps the natural-summary prompt.
The forced-summary validator, relay-TTS fallback, and delivery-confirm
alarm backstop both, so an off-script response degrades to the previous
TTS-direct behavior instead of losing the answer.

_speak_result_verbatim and the voice.response.verbatim_delivery event are
removed; foreground, background, and deferred-resume paths share the
injection pipeline. Voice Settings' delivery-mode info dialog, docs, and
the stale CHANGELOG keepalive bullet (superseded by idle-close recovery)
are updated to match.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 18:08:33 -04:00
Bailey Dixon d15f7860fc fix(voice): recover idle realtime sessions 2026-07-08 16:48:13 -04:00
Bailey DixonandClaude Fable 5 1660750b67 fix(cli): per-call RPC timeout override + long prompt.submit timeout
HRUI-016 (desktop half): RelayTransport.request() bounded every RPC by
the single env-tunable HERMES_RELAY_RPC_TIMEOUT_MS (default 120s), so a
legitimately long prompt.submit ack rejected mid-turn; the chat/voice
turn promises additionally wall-clock capped healthy turns at 10m/5m.

- RelayTransport.request() (and the Transport interface + GatewayClient
  pass-through) accept an optional per-call timeoutMs; the env-var
  default still covers every other call.
- Export PROMPT_SUBMIT_REQUEST_TIMEOUT_MS = 1_800_000 (mirrors upstream
  apps/desktop/src/hermes.ts, commit 164144183) and pass it at both
  prompt.submit call sites (chat.ts runOneTurn, voiceServer.ts) — the
  only two under desktop/src/.
- Convert the 10-min chat and 5-min voice turn caps into idle-progress
  watchdogs: the timer re-arms on every gateway event and only fires
  after that long with NO events at all, so streaming turns are never
  wall-clock capped.

Verified with npm run build (strict tsc).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:58:19 -04:00
Bailey DixonandClaude Fable 5 a1818579d9 fix(gateway): 30-min prompt.submit RPC timeout on Android, no fallback on slow ack
HRUI-016: upstream treats gateway prompt.submit as a long-running RPC —
the ack is effectively fire-and-forget (turn completion arrives via
stream events, not the RPC return) and can trail a MoA/deep-reasoning/
tool-heavy turn by minutes. Bounding it by the generic 15s rpc timeout
false-failed running turns into the SSE preflight fallback, resubmitting
the same prompt as a duplicate turn.

- Add PROMPT_SUBMIT_REQUEST_TIMEOUT_MS = 1_800_000 (mirrors upstream
  apps/desktop/src/hermes.ts, commit 164144183; matches the backend
  agent.gateway_timeout = 1800s ceiling) and pass it at the
  prompt.submit call site.
- Guard the submit-failure path: once this turn's own events are
  flowing (or it already ended), a slow/lost/socket-severed ack no
  longer fires onPreflightFailure — recovery stays with the idle
  watchdog and mid-turn rejoin. session.info is excluded from the
  "turn started" signal (connection-level, turn-independent).
- The 180s turn watchdog was already idle-progress (reset on every
  gateway event) — semantics unchanged, docs clarified.
- Expose rpc/submit/idle timeouts as constructor test seams (same
  pattern as midTurnRejoinWindowMs); harness gains a withheld-ack seam.
  4 new JVM tests: slow-ack survives past the generic timeout with no
  fallback, late ack timeout after completion does not resubmit, idle
  watchdog stays quiet while events trickle, and still fires on silence.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:57:54 -04:00
Bailey DixonandClaude Fable 5 f893330cfa docs(bootstrap): split supported-baseline vs compat-only surfaces in doctor/compat/TODO wording (HRUI-002)
Update every place that described the bootstrap as a sessions/skills
fallback to reflect the retirement split:

- plugin/doctor.py + plugin/compat.py docstrings, compat status
  recommendation, compat status text, and the legacy-bootstrap doctor
  check now state that compat covers only session search, memory, legacy
  skill detail/toggle, config, available-models, and slash middleware.
- TODO.md bootstrap-injection entry records the sessions + skills-list
  retirement as done (2026-07-08) and lists the still-gapped surfaces.
- CLAUDE.md bootstrap-maintenance bullet, compatibility-endpoints table,
  Key Files row, and Integration Points row updated to match.
- docs/upstream-surface-matrix.md, docs/upstream-integration-sync.md,
  docs/upstream-contributions.md, docs/decisions.md ADR 16 removal path,
  docs/HERMES-WEBAPI-REFERENCE.md, and docs/remote-access.md no longer
  claim the bootstrap injects sessions or the legacy skills list.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:55:30 -04:00
Bailey DixonandClaude Fable 5 16b16fd5ad refactor(bootstrap): retire native-upstream sessions + skills/toolsets injection (HRUI-002)
Remove the bootstrap handlers for surfaces current hermes-agent serves
natively: sessions CRUD/messages/fork (/api/sessions*, upstream PR #33134)
and the legacy read-only GET /api/skills list (superseded by /v1/skills +
/v1/toolsets, PR #33016). No pre-#33134 fallback remains; older core builds
degrade via the client capability probe to /v1/chat/completions or /v1/runs.

The bootstrap now injects only genuine compatibility gaps with no native
API-server replacement: GET /api/sessions/search, memory CRUD, legacy skill
detail (/api/skills/{name}) + the 501 toggle stub, config, available-models,
and the slash-command middleware. Registration stays method/path-aware so
native routes still win if any remaining surface lands in core.

Tests assert the split both ways: retired surfaces are never injected, kept
surfaces are, and native routes still win for kept surfaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:55:01 -04:00
Bailey DixonandClaude Fable 5 227748912d docs: keepalive final verdict — no protocol message resets xAI's 900s timer
Probe run 4 (valid): three session.update pings at 240/480/720s, each
acknowledged by the server, and the conversation still timed out at
exactly 900.0s. Combined with the silent-append runs: xAI's inactivity
timer counts only real conversation items. Remaining designs recorded
(scheduled reopen vs silent auto-reopen-on-next-turn with Hermes-session
context reseed; POC doc recommends the latter).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:47:16 -04:00
Bailey DixonandClaude Fable 5 9a1edf7d5a chore(deps): raise aiohttp floor to >=3.14.1 for 2026 CVE line
Upstream pinned all aiohttp paths to the patched 3.14.1 line
(CVE-2026-34993, CVE-2026-47265, and the earlier 2026 advisories). Raise
the relay floor to match in plugin/requirements.txt, pyproject.toml, and
the relay_server compat shim requirements, and update the version table
in docs/spec.md. README/AGENTS mention aiohttp without a version string,
so they need no change.

Refs HRUI-015.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:46:54 -04:00
Bailey DixonandClaude Fable 5 ab65ea7705 fix(manage): request include_unconfigured model options + keep provider setup rows (HRUI-022)
Upstream flipped the default of dashboard GET /api/model/options (and the
gateway model.options RPC) to configured-providers-only; unconfigured
provider skeleton rows now require an explicit include_unconfigured opt-in.
Relay Android called the route bare and parseModelOptions dropped
empty-models rows, so on new upstream every provider awaiting an API key
silently vanished from the Manage model picker along with its Keys-setup
affordance.

- DashboardApiClient.getModelOptions() always sends include_unconfigured=1
  (cached and refresh=1 paths); old upstream ignores the extra param.
- parseModelOptions keeps empty-models skeleton rows, resolves the provider
  id from the canonical slug first (what /api/model/set expects), defaults
  authenticated by model presence when picker hints are absent, and carries
  the upstream warning as a setupHint.
- ModelPickerDialog renders the setup hint (e.g. "paste X_API_KEY to
  activate") under empty skeleton providers, keeping the Keys guidance.
- Gateway-WS model.options callers audited: the only call site feeds the
  in-chat picker, which intentionally stays on the configured subset.

Verified: :app:testGooglePlayDebugUnitTest — DashboardApiClientTest (42)
and new ModelOptionsParserTest (5, incl. old-upstream back-compat fixture)
all pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:46:49 -04:00
Bailey DixonandClaude Fable 5 fde5030797 fix(relay): always-on credential denylist for /media/by-path permissive mode
Mirrors upstream hermes-agent media-delivery hardening
(gateway/platforms/base.py validate_media_delivery_path): even with
RELAY_MEDIA_STRICT_SANDBOX off, /media/by-path now refuses to serve
credential/system paths — ~/.hermes/.env, auth.json, config.yaml, OAuth
token stores, pairing/, mcp-tokens/, ~/.ssh and the other home credential
dirs, /etc and other system prefixes — plus the relay-specific
hermes-relay-qr-secret and hermes-relay-sessions.json secret stores.

The check runs after realpath resolution (a symlink can't launder a
denied target) and before the existence check (no 403-vs-404 existence
oracle for credential probes). Ordinary files keep serving in permissive
mode; strict-sandbox mode is unchanged except the denylist now outranks
the allowlist there too.

Refs HRUI-014.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:45:30 -04:00
Bailey DixonandClaude Fable 5 0173519183 fix(relay): hold result delivery while the user is mid-utterance
Live round-5 finding: a background task finishing while the user was
speaking delivered over them and ended their recording. The relay knows
the user is talking (live input_audio.append chunks now stamp
native_last_input_audio_at); _await_floor_idle_for_result additionally
requires the user quiet >= 1.5s before consuming the floor, bounded by
the existing wait deadline. Covers summary, fallback, and queued-start
transition deliveries. TODO logs the client half (don't end an active
recording on incoming audio), the end-of-response audio tail cut repro,
fallback path-speech polish, and the 4/4 grok delivery-instruction
failure stat elevating verbatim delivery to likely default.

Tests: input-quiet gate holds/proceeds cases; affected suites 42/42.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:44:34 -04:00
Bailey DixonandClaude Fable 5 c5d61bc427 test(probe): stamp stream-end time — a pre-ping socket death is not a mode verdict
The first session_update keepalive run died inside the initial 240s,
before any ping was sent (the pinger exits silently when the event stream
ends) — an unstamped clean stream end made that indistinguishable from a
keepalive failure. Stamp it so early deaths read as invalid runs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:29:05 -04:00
Bailey DixonandClaude Fable 5 f116d41295 docs: log live rounds 3-4 verdicts + keepalive negative + verbatim-delivery idea
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:27:13 -04:00
Bailey DixonandClaude Fable 5 4aedb9b839 test(probe): record silent-append verdict; add session_update keepalive mode
Empirical (relay host, live xAI, 2026-07-08): the 960s repro died at
exactly 900.0s, and the silent-PCM keepalive run ALSO died at exactly
900.0s — uncommitted input_audio_buffer.append does NOT reset xAI's
conversation-inactivity timer. POC doc revised; the silent-append
keepalive stays as harmless scaffolding until a working ping lands. Next
candidate: a session.update re-send (connection.configure()), now
available as --keepalive-mode session_update; if that fails too, the
remaining option is a scheduled provider-socket reopen.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:24:40 -04:00
Bailey DixonandClaude Fable 5 ba24b3e959 fix(relay): next-turn correction after system-side deliveries
Live round-4 finding: the fallback spoke the answer correctly, but the
provider never sees fallback/text-only deliveries — its conversation
history still read "running in background", so on the user's next turn it
claimed the task was still running.

Out-of-band deliveries (forced-summary fallback, text-only emit,
delivered-or-alarm force emit, respeak) now set a pending delivery note;
the next normal user turn's response carries it via per-response
instructions (composed WITH the session instructions, which per-response
instructions otherwise replace), then clears it: the task has ALREADY
COMPLETED, the answer was already spoken, don't re-deliver.

e2e filler test extended: after the fallback, the correction note must be
pending and carry the delivered answer. Realtime batch 85/85 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:21:08 -04:00
Bailey DixonandClaude Fable 5 5896d4c672 fix(relay): summary validation — whole-word overlap, 2-hit early-commit bar, queue-speak blocklist
Live round-3 regression: the forced summary for a completed answer came
back as ANOTHER queue acknowledgement ("It's queued and will start
automatically once the Minnesota check finishes. I'll let you know…") and
early-commit approved it at 45 chars — substring matching let "will START
automatically" count as evidence for an answer containing "starting", the
audio played, the response was marked delivered, and the real answer never
spoke (so the delivered-or-alarm stayed silent too).

- _summary_overlap_hits: WHOLE-WORD evidence matching (substring was the
  hole); overlap check now counts hits
- early commit requires >= 2 whole-word hits (irreversible once audio
  plays, so the early bar is higher than end-of-response validation's 1)
- blocklist gains the queue/deferral class a final answer must never
  contain: "i'll let you know", "it's queued", "is queued", "queued and
  will", "in the queue"
- _start_next_queued_run gains an explicit phase-1 wait for the summary
  injection to BEGIN (correct ordering previously held only by task
  scheduling luck) before the existing wait-for-finish
- the queued-start spoken transition now waits for floor idle so it can't
  overlap a fallback TTS render of the previous result

Tests: live regression strings pinned (queue-speak flagged end-of-response,
substring non-overlap, single-weak-hit no-commit); realtime batch 85/85
green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:13:36 -04:00
Bailey DixonandClaude Fable 5 3c24e82e5e docs: log the background-run A-E batch
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 14:47:48 -04:00
Bailey DixonandClaude Fable 5 381c62d6b6 feat(voice): queued-count chip, respeak on DONE-chip tap, compact-mode chip, exit breadcrumb
Client half of the background-run A-E batch:

- chip shows "+N queued" (hermes.run.queued + queued_count on promoted)
- tapping the settled (DONE) chip asks the relay to respeak the last
  delivered answer (hermes.result.respeak); chip stays up while it plays;
  taps on live phases no-op
- the chip now renders in compact (non-focus) voice mode too — it
  previously existed only in the focus layout, so a running task had no
  visible presence there
- exiting voice mode with a live background run posts a chat system
  notice ("Background voice task still running (+N queued) — Hermes will
  report back") via VoiceViewModel.chatNoticeSink, wired in RelayApp to
  the shared ChatHandler
- `_thinking` drafting deltas drive a "Drafting the answer…" chip status
  line (never a tool pill)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 14:47:47 -04:00
Bailey DixonandClaude Fable 5 ad0139ba94 feat(relay): streaming summary delivery, answer-overlap validation, task queue, respeak
Background-run A-E batch (relay half):

- forced summaries STREAM: buffered only until the prefix (>=40 chars)
  clears the bad-phrase check and shows content overlap with the Hermes
  answer (_maybe_commit_forced_summary_early), then flush + live stream —
  removes the "silence, then the whole answer in one burst" delivery gap;
  committed responses skip end-of-response validation (audio already
  played)
- positive validation: _summary_overlaps_answer requires the summary to
  share content tokens with the answer (vacuous for bare confirmations);
  no_answer_overlap joins the bad-summary reasons
- delivered-or-alarm: _confirm_background_delivery force-emits the answer
  as text (+ delivery_unconfirmed log) when no spoken delivery lands
  within 30s — a background answer can never be silently lost
- respeak: hermes.result.respeak client message replays
  last_background_result via relay TTS (DONE-chip tap client-side)
- task queue: a long second ask is queued (FIFO, cap 3, status "queued")
  instead of refused; starts automatically when the current run's task
  completes (_start_next_queued_run waits for the summary to settle, runs
  the task as durable with a spoken transition); cancel clears the queue;
  hermes.run.queued event + queued_count on promoted /
  background_completed / get_status; queue-full keeps the busy answer
- _thinking drafted text is the answer of last resort when the
  response-delta path yields empty (answer_from_thinking)
- fast lane reuses one Hermes side-session per voice session
  (fast_lane_session_id) instead of one session per quick ask
- idle probe injects the relay xAI OAuth token like the broker does
  (_probe_provider_options) so it runs on the relay host

Tests: 93 green across the realtime batch — 8 new overlap/early-commit
cases, 4 new queue/side-session cases, and a route-level misbehaving-
provider e2e (filler summary -> fallback carries the real answer; the
filler never reaches the client). The second-ask contract changed from
already_running to queued; existing tests updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 14:47:46 -04:00
Bailey DixonandClaude Fable 5 2abf9b000f docs: log the chip DONE-settle fix (sixth e2e finding)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 13:05:27 -04:00
Bailey DixonandClaude Fable 5 bc957ef641 fix(voice): background-run chip settles to DONE instead of vanishing mid-answer
The chip was nulled at the first summary-audio byte — it disappeared
exactly when the waveform/spinner returned to speak the answer, reading
as the background task being lost (second live e2e finding, same day).

- new BackgroundRunPhase.DONE: on first summary audio (or the 20s
  no-audio delivery watchdog) the chip settles to "Background task
  finished." — solid dot (no pulse), elapsed ticker frozen — lingers
  DONE_CHIP_LINGER_MS (10s), then auto-dismisses
- the chip ✕ on a DONE chip is a LOCAL dismiss, never a relay cancel
  (a late cancel used to overwrite a delivered answer); TalkBack label
  flips to "Dismiss"
- a newly promoted run replaces a lingering DONE chip and cancels its
  auto-dismiss timer so it can't clear the new chip's later DONE early
- progress / tool / reconnect handlers treat DONE like DELIVERING:
  a settled chip cannot be reanimated by stray late events

:app:assembleSideloadDebug green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 13:05:27 -04:00
Bailey DixonandClaude Fable 5 e3097682c1 docs: log e2e realtime forensics + the five voice fixes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 12:47:03 -04:00
Bailey DixonandClaude Fable 5 63a7a79d2d fix(voice): five chained e2e fixes — stuck thinking pill, late-cancel answer loss, spoken run IDs, phantom queue, deferral filler
Live e2e forensics (session event log): the gateway streams drafting text
as a `_thinking` pseudo-tool (deltas only, never tool.completed) -> the
client rendered it as a forever-"running" pill -> the user cancelled an
ALREADY-COMPLETED run -> the unguarded cancel marked it cancelled and the
completed answer was never spoken. Independently the model read the full
32-char run id aloud, claimed to "queue" a request (no queue exists), and
one delivery spoke "One moment while I look that up" filler the summary
validator didn't recognize.

- client: `_`-prefixed tool names are internal (upstream hidden-tool
  convention) — hermes.tool.delta/.started no longer create ToolCall pills
  for them; text still feeds the detailed thinking trace
- relay: response.cancel only cancels a Hermes run that is actually in
  flight; late cancel still stops speech but cannot flip a completed run
  to "cancelled" or emit hermes.run.cancelled for it
- relay: run/session ids removed from every model-visible payload
  (interim ack, forced-summary metadata); "never say run IDs, session
  IDs, or other identifiers aloud" added to interim-ack, handoff, and
  summary instructions (get_status/cancel default to the active run)
- relay: "there is no task queue" added to handoff/busy instructions
- relay: _bad_forced_summary_reason gains deferral-filler phrases (one
  moment / report back / looking into / i'll look / as soon as i have);
  summary prompt reworded to speak the answer NOW
- relay: pre-Hermes status lead no longer carries the previous run's
  run_id/tool-count into a new run's first progress event

Tests: new plugin/tests/test_realtime_summary_validation.py (5, pinning
the exact observed filler), cancel route test updated to the
no-active-run contract; realtime batch 69/69 green;
:app:compileSideloadDebugKotlin green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 12:47:03 -04:00
Bailey DixonandClaude Fable 5 789f32cd25 docs: log fast lane + stale voice-prefs TODO closure + device deploy
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 11:24:34 -04:00
Bailey DixonandClaude Fable 5 6f0357c2e8 feat(relay): fast lane — answer quick asks inline during a background run
Background-run v2 item 1. A second hermes_run_task while a detached
(promoted/durable) run holds the single background slot used to get an
unconditional busy answer — even for a two-second lookup. The broker now
first tries the request INLINE on a separate ephemeral Hermes session
(session_id=None) within the normal grace window:

- completes inside grace -> the tool result is returned (fast_lane: true)
  and spoken as usual
- grace elapses, a known-long tool starts (_long_tool_hints), the call
  asks mode=background, or promotion is off -> the attempt is abandoned
  (stream cancelled client-side) and the reworded busy answer falls
  through unchanged
- gate requires the in-flight run to actually be detached
  (hermes_run_tier promoted/durable)

The fast lane keeps every observation in locals and touches NONE of the
session's hermes_* run state — run_id, status, progress counters, and the
chip stay owned by the in-flight run — and emits no client events of its
own (bounded by grace, so no chip is needed; one would fight the detached
run's). Session-log events: voice.hermes_fast_lane.completed / .abandoned
/ .error.

Tests: plugin/tests/test_realtime_fast_lane.py (7 — inline answer + state
non-interference, grace/long-tool fall-throughs incl. client-side
cancellation, background-mode/promotion-off/foreground-tier skips, error
reporting). test_second_run_task_answers_busy_without_orphaning_first
updated to per-stream cancellation tracking: the abandoned fast-lane
stream is the designed fall-through; the first run's stream must stay
uncancelled. Full realtime batch 64/64 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 11:24:30 -04:00
Bailey DixonandClaude Fable 5 7569144cc4 docs: close stale voice-prefs connectionId TODO; correct the KDoc
The connectionId namespacing wiring already shipped in 0aa1b38 (2026-06-21):
RelayApp's (connection, profile) effect calls setVoicePrefsConnection before
onProfileChanged, and applyVoicePrefsScope pushes both into
VoicePreferencesRepository.setActiveScope. The deferred-list entry and the
"null until an integration wires this" KDoc paragraph described the
pre-wiring state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 11:13:13 -04:00
Bailey DixonandClaude Fable 5 5b25b9b154 docs: log #131 closure + demo composer; close audit + demo TODO items
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:39:45 -04:00
Bailey DixonandClaude Fable 5 79dd6b44fe feat(app): demo composer answers with a canned notice instead of a no-op
Typing + Send in offline Demo mode did nothing (sendMessage early-returned
on the null API client), which read as broken. sendMessage now intercepts
while isDemoMode: echoes the user bubble and appends
DemoContent.composerReply — an honest "offline demo, tap Connect in the
banner" assistant notice. Both bubbles are clientOnly, so demo-exit's
clearMessages() wipes them with the rest of the transcript.

Wired via setDemoModeWiring unconditionally in RelayApp: in demo there is
no API client, so the client-gated chat init never runs and ChatViewModel's
own handler stays null — the wiring supplies both the demo flag and the
shared ChatHandler. UUID-based ids so rapid sends can't collide on
LazyColumn keys.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:39:44 -04:00
Bailey DixonandClaude Fable 5 d96b68c794 fix(app): guard streaming URL builds against malformed base URLs (#131)
sendChatStream / sendCompletionsStream / sendRunStream built their Request
before any try/catch or listener existed, so a malformed apiServerUrl
(hand-edited connection, corrupt settings import) made
Request.Builder.url(String) throw IllegalArgumentException synchronously
out of the ViewModel — the last open group in the #131 "Invalid URL host"
crash-class audit.

- authRequestOrNull() chokepoint backed by top-level buildApiRequestOrNull
  (mirrors ConnectionManager's buildRelayRequestOrNull so the guard is
  unit-testable without instantiating the client)
- a bad URL fails the turn through the normal onError channel with a
  human message and returns an inert EventSource; no side effects fire
  before the guard
- tests: valid/malformed URL cases in HermesApiClientTest

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:39:43 -04:00
Bailey DixonandClaude Fable 5 e3ba358331 docs: record delegate_task async-delivery verdict; retract voice nudge records
Upstream verification (clone @ 5057f03bf): delegate_task(background=true)
never dispatches async on the api_server surface — every api_server route
binds async_delivery=False and tools/delegate_tool.py downgrades the batch
to synchronous execution (upstream issue #10760). All standard voice turns
ride SSE/api_server, so the background-delegation nudge could not work as
designed and was reverted (45c7ef4); the speak-delegated-result-on-overlay
follow-up is closed on the same finding (no delayed completion turn exists
on that surface). TODO records the verified mechanism with source
locations; the CHANGELOG entry is withdrawn; DEVLOG item corrected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:25:36 -04:00
Bailey Dixon 45c7ef49e2 Revert "feat(voice): nudge standard voice toward backgrounding long asks"
This reverts commit 5c214a2e6a.
2026-07-08 10:23:45 -04:00
Bailey DixonandClaude Fable 5 6003258c5d docs(devlog): log 2026-07-08 voice batch (keepalive, sync durability, nudge)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:15:16 -04:00
Bailey DixonandClaude Fable 5 a660b3825d docs: log realtime sync drain, provenance badge, and voice nudge
TODO: mark the durability + provenance-chip item shipped (gateway drain +
marker->badge + orphan dedupe), scope the remaining app-restart
persistence question, and mark the standard-voice delegate_task nudge
shipped with its on-device verify steps. CHANGELOG: user-facing entries
for the sync drain/badge and the background-delegation nudge under 1.4.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:13:33 -04:00
Bailey DixonandClaude Fable 5 5c214a2e6a feat(voice): nudge standard voice toward backgrounding long asks
Add a line to the ephemeral voice interface context telling the model the
user is waiting in a live voice session: clearly-long requests (builds,
research, multi-step tool work) should be delegated via
delegate_task(background=true) with a spoken "started it in the
background" acknowledgement, while quick questions keep answering
directly. Rides the per-turn SSE system_message — nothing is persisted
and text chat is unaffected. Deliberately hedged: false-positive
delegation is worse UX than a long turn (SSE recovery + the turn-complete
notification already make those survivable).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:13:33 -04:00
Bailey DixonandClaude Fable 5 c079a632ba fix(app): drain realtime turn sync on gateway + badge synced voice turns
Provider-answered realtime voice turns are folded into the Hermes session
as synthetic messages on the next chat/run request — but the gateway
prompt.submit can't carry them, and on a gateway-primary phone "wait for
the next SSE turn" meant never: the agent never learned what was said in
voice.

- force a gateway turn with unsynced synthetic sync messages (voice
  intents / card dispatches / realtime turns) onto the sessions SSE route
  so the traces land; guarded to an existing session id + the sessions
  fallback + the default profile (a non-default profile's gateway session
  is invisible to the shared api_server surface — the POST would 404 and
  fail the user's turn)
- mark traces synced based on the route the turn actually DISPATCHED on
  (effectiveEndpoint), fixing a latent duplicate re-send for voice turns
  forced onto SSE by their interface context
- RealtimeTurnSyncBuilder.stripProvenanceMarker(): recognize the synced
  "[Realtime Agent provider-native voice turn: ...]" marker in loaded
  history, strip the bracket noise, and restore the quiet "Realtime
  Agent" badge live turns get
- drop the superseded local clientOnly bubble when its synced copy loads
  from the server (the exchange rendered twice otherwise); unsynced
  traces stay preserved — they are still the only record of the turn
- tests: 4 new ChatHandlerTest load-path cases, 4 new
  RealtimeTurnSyncBuilderTest marker cases (incl. builder round-trip)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 10:13:13 -04:00
Bailey DixonandClaude Fable 5 c7de0da22d fix(relay): keep realtime provider sessions alive across long silence
xAI closes a realtime conversation after 900s of inactivity; with manual
turn-taking (turn_detection: None) the provider socket sees nothing while
the user is silent, so an open-but-quiet voice session — most commonly a
background-run wait — died with a raw provider error (observed live
2026-07-08).

- add _provider_keepalive_loop: per-connection broker task appends ~100ms
  of silent, never-committed PCM after RELAY_VOICE_PROVIDER_KEEPALIVE_MS
  of quiet (default 240s => 3 pings per 900s window; 0 disables); runs
  through detached periods; append-only so it can never race or clobber
  a user utterance
- stamp provider activity in two places only: client input_audio.append
  and once per provider event in _pump_provider_events
- classify residual provider idle-closes (_is_provider_idle_timeout) into
  a human-readable "voice session expired" error instead of raw provider
  text
- extend realtime-provider-idle-probe.py with --keepalive-ms for the
  relay-host repro (--windows 960) and fix verification
- revise ADR 33 Phase 0: xAI is needs-keepalive beyond ~900s (POC doc
  revision 2026-07-08 + decisions.md note)
- tests: 11 new in plugin/tests/test_realtime_keepalive.py; existing 54
  realtime tests green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 09:57:23 -04:00
Bailey Dixon 16133ff081 docs(todo): standard-voice research follow-ups + xAI 900s idle-timeout diagnosis 2026-07-07 22:51:15 -04:00
Bailey Dixon 3a12221185 docs: mark realtime injection-framing fix as deployed to relay 2026-07-07 21:00:16 -04:00
Bailey Dixon 699653bf87 docs: log voice chip fix, screen-wake-lock, and injection-framing fix 2026-07-07 20:55:02 -04:00
Bailey Dixon 3d354080dd fix(relay): stop faking user turns for realtime voice injections
The forced-Hermes preamble, background-task handoff ack, and
completed-background-task summary were all injected via
conversation.item.create role=user — the model's history contained
fake turns like "the user" saying "Hermes has already handled the
user's previous voice request...".

response.create supports a per-response instructions field that
overrides the session prompt for one response only, with no
conversation item created at all. Confirmed supported by both
providers (OpenAI docs; xAI's Voice Agent API docs show the same
shape) — conversation:"none" (OpenAI-only true out-of-band) is
deliberately not used since the spoken summary should remain real
history for follow-up turns to reference.

request_response() gained an optional instructions kwarg on both
provider adapters; the 4 broker-authored injection sites switched
from send_text(prompt) to request_response(instructions=prompt).
The one genuine passthrough (real client-supplied text) is
untouched.
2026-07-07 20:54:32 -04:00
Bailey Dixon 427145ccce fix(android): voice tool-call chip ordering + screen-wake-lock
Background-run chip pinned a finished tool's status line until the
next unrelated event overwrote it (no hermes.tool.completed/failed
handler); CompactTranscriptRow rendered the reply above the tool
calls that produced it. Also add KeepScreenOnWhile so voice mode
holds the screen on for the whole session and chat holds it only
while a reply streams, matching call/video-playback conventions
instead of relying on the OS default throughout.
2026-07-07 20:54:00 -04:00
Bailey DixonandClaude Opus 4.8 654663bc3e docs(todo): add compaction-safe "Active — next up" snapshot
Leads TODO with the current state so work can resume cleanly after a session
compact: the prepped android-v1.4.0 / plugin-v1.4.0 release act (sign-off →
notes → merge/tag → discard the 1.3.0 Play draft), the two open voice bugs
(tool-completion handler + PCM click, both needing a logcat repro), and the
Mizu triage queue. Trimmed the now-shipped duplicate-toast voice item.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 18:14:55 -04:00
Bailey DixonandClaude Opus 4.8 75d965c32e fix(android): finish relay URL-guard sweep + voice error-recovery UX
Relay URL guards (rest of the #131 relay class):
- RelayVoiceClient validates its base in resolveHttpBase() → null on a
  malformed URL, so all 7 voice endpoints fail via the existing Result.failure
  guards instead of a throwing .url() on the IO dispatcher.
- RelayHttpClient's two string-URL sites (fetchMedia, listSessions) now use
  toHttpUrlOrNull() → Result.failure. RelayProfileInspectorClient was already
  guarded (toHttpUrl + catch everywhere).

Voice error-recovery UX (from the on-device realtime test):
- VoiceModeOverlay no longer pipes errorEvents to the app-wide bottom snackbar
  while it's up — the inline top banner is the single surface, killing the
  duplicate bottom toast on a failed/timed-out turn.
- clearError() now resets Error→Idle so a dismissed/retried failure lands
  usable; the banner gained a Dismiss beside Retry (was retry-only, which
  trapped the user).

Also diagnosed in TODO (need a repro-with-logs before fixing): realtime
tool-call spinners run forever (no hermes.tool.completed/failed handler in
VoiceViewModel though the relay forwards them); tap/static click between
sentences in realtime PCM playback.

Build + install + launch-clean verified on device.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 18:07:51 -04:00
Bailey DixonandClaude Opus 4.8 37355974b4 fix(android): guard malformed relay URL in ConnectionManager (relay half of #131)
Play crash on 1.2.6 (Galaxy S25 Ultra / Android 16): IllegalArgumentException
from okhttp3.HttpUrl$Builder.parse via ConnectionManager.doConnectInternal →
Request.Builder.url(). doConnectInternal runs on a background coroutine, so a
malformed relay host (from a corrupt/edited pairing payload) made OkHttp's url()
throw uncaught → app crash. This is the relay-socket half of the #131 "Invalid
URL host" class the TODO flagged (the #131 fix only covered Manage/voice HTTP).

- Extracted a pure buildRelayRequestOrNull() (try/catch → null on
  IllegalArgumentException). doConnectInternal treats null as a connection
  failure: "Invalid relay URL" diagnostic + Disconnected + close-replaced-socket
  + backed-off reconnect — the same path onFailure uses. No happy-path change.
- ConnectionManagerUrlGuardTest: valid ws/wss build; empty-host / space-in-host
  return null. Green via :app:testSideloadDebugUnitTest; APK rebuilt + installed
  + launched clean on device.

TODO #131 audit: ConnectionManager marked fixed; RelayHttpClient /
RelayProfileInspectorClient / RelayVoiceClient remain for a defense-in-depth pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 17:26:37 -04:00
Bailey DixonandClaude Opus 4.8 380ad0b4bf release: prep android-v1.4.0 + plugin-v1.4.0 (versions + changelog + devlog)
- Android → 1.4.0 (versionCode 22); plugin → 1.4.0 (all metadata in sync).
- CHANGELOG [Unreleased] → [1.4.0] - 2026-07-07.
- DEVLOG entry for the CI-path-coverage + Android-14 crash-safety work.

Prep only — in-app What's New / RELEASE_NOTES / Play notes, the dev→main
merge, and the android-v1.4.0 / plugin-v1.4.0 tags are the release act,
pending an on-device smoke test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 16:31:24 -04:00
Bailey DixonandClaude Opus 4.8 13e747c6d9 fix(android): Android-14 crash-safety — removeFirst/removeLast + Tink pin
Built against SDK 35, Kotlin's removeFirst()/removeLast() resolve to Java 21's
List methods that don't exist below Android 15, crashing older devices.

- Replaced all 5 app-code removeFirst() calls (all on kotlin ArrayDeque, so
  members not the flagged MutableList extension — already safe, but converted
  per Google's guidance and for future-proofing) with removeAt(0). All sites
  are size-guarded, so behavior is identical.
- Pinned com.google.crypto.tink:tink-android:1.16.0 ahead of the transitive
  version security-crypto pulls, whose HybridConfig.<clinit> tripped the same
  Play pre-launch check. Our EncryptedSharedPreferences use is AEAD-only, so
  HybridConfig is almost certainly never loaded — this clears the static Play
  warning. Untestable without a build; on-device auth smoke-test queued in TODO.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 16:06:08 -04:00
Bailey DixonandClaude Opus 4.8 500387d3fa ci(plugin): trigger on all plugin/*.py, not just four named modules
The path list omitted doctor.py, compat.py, config.py, profiles.py, and 5
other top-level modules, so changes to them alone never ran plugin CI (my
doctor.py fix only got covered because it also touched plugin/tests/**). A
plugin/*.py glob covers every current and future top-level module.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 16:06:07 -04:00
Bailey Dixon 55904a600a Merge pull request #178 from Codename-11/fix/installer-stale-plugin-backup
fix(plugin): doctor + installer guard against stale duplicate plugin copies
2026-07-07 13:51:20 -04:00
Bailey Dixon 393a485f53 Merge branch 'dev' into fix/installer-stale-plugin-backup 2026-07-07 13:47:18 -04:00
Bailey Dixon 4fc1978df7 Merge pull request #176 from Codename-11/chore/todo-prune
chore(todo): prune shipped records + fix stale #8556→#33134 refs
2026-07-07 13:47:03 -04:00
Bailey Dixon 141a7560e7 Merge pull request #177 from Codename-11/revert/release-video-pipeline-repo-import
revert: remove release video pipeline import
2026-07-07 13:26:05 -04:00
Bailey DixonandClaude Opus 4.8 f965c205d7 fix(plugin): doctor + installer guard against stale duplicate plugin copies
The gateway plugin loader dedups discovered plugins by manifest name, so a
second directory declaring `name: hermes-relay` (a backup copy left by an
older installer, or a stray extra install) could win the dedup and make the
gateway load stale code — silently ignoring every later deploy. This was the
root cause of the 2026-06-29 phone-platform round-trip failure.

- doctor: new `_duplicate_plugin_dirs()` + `plugin-name-unique` check warns
  when >1 directory under the plugins dir declares the same plugin name
  (deduped by resolved real target); report gains `duplicate_dirs`/`plugins_dir`.
- install.sh: sweep the plugins dir after symlinking and remove any other
  entry declaring `name: hermes-relay`, so a stale duplicate can't linger.

Tests: 4 new doctor cases (14/14 green). install.sh grep-match validated in
isolation. Live-host verify queued in TODO.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 13:19:29 -04:00
Bailey Dixon a138165408 Revert "Merge pull request #175 from Codename-11/fix/release-video-pipeline"
This reverts commit 467e6722a2, reversing
changes made to 66728686b9.
2026-07-07 13:18:57 -04:00
Bailey Dixon 467e6722a2 Merge pull request #175 from Codename-11/fix/release-video-pipeline
feat: add release video pipeline
2026-07-07 13:12:50 -04:00
Bailey DixonandClaude Opus 4.8 b5c1d392fb chore(todo): prune shipped records + fix stale #8556 → #33134 refs
- Remove 11 shipped-and-released [x] records from "User-Added" (session
  delete, voice override, analytics/diagnostics, connections reframe,
  profile lock, etc.) — they live in DEVLOG; keep the one open [ ] item.
- Collapse the dot-matrix "thinking indicator" section: base + presets +
  colors shipped in android-v1.3.0; keep only the two real remainders
  (OS reduce-motion/TalkBack; optional avatar-style promotion).
- Fix the "Research / open questions" bootstrap notes: PR #8556 was closed
  as superseded; native upstream now covers sessions via #33134 and
  skill/toolset discovery via /v1/skills + /v1/toolsets (#33016). Bootstrap
  shrinks per surface, not one big delete. Stage 2 slash-preprocessor is
  unblocked (was "blocked on #8556").

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 13:10:22 -04:00
Bailey Dixon 5df941a179 feat: add release video pipeline 2026-07-07 13:02:26 -04:00
Bailey Dixon 66728686b9 Merge pull request #174 from Codename-11/chore/pr-batch-pre-minor
chore: batch 4 feature PRs onto dev for the next minor (#172, #170, #123, #171)
2026-07-07 13:00:36 -04:00
Bailey Dixon c05ee40db4 Merge PR #171 (bblicke1:feat/android-multi-device-bridge) into pr-batch — multi-device bridge targeting
# Conflicts:
#	DEVLOG.md
2026-07-07 12:44:17 -04:00
Bailey Dixon 66166b2c2a Merge PR #123 (feat/axi-26-notification-triggers) into pr-batch — notification triggers MVP
# Conflicts:
#	CHANGELOG.md
#	DEVLOG.md
2026-07-07 12:42:40 -04:00
Bailey Dixon c8a6534bec Merge PR #170 (feat/upstream-impact-cleanup-dashboard-serve) into pr-batch — upstream-impact Relay guardrails
# Conflicts:
#	CHANGELOG.md
#	DEVLOG.md
2026-07-07 12:39:49 -04:00
Bailey Dixon 8afbb6120d Merge PR #172 (feat/model-picker-refresh-parity) into pr-batch — model picker refresh parity
# Conflicts:
#	DEVLOG.md
2026-07-07 12:34:32 -04:00
Bailey DixonandClaude Fable 5 9c080f522c docs: changelog entry for #165 fix + devlog for the v1.3.0 release act
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 22:37:10 -04:00
Bailey DixonandClaude Fable 5 d2a2d01072 Merge branch 'fix/plugin-native-imports' into dev — native-loader imports + installer venv autodetect (#165, for plugin-v1.3.1)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 22:36:33 -04:00
dependabot[bot] fe297e2db9 chore(deps): bump io.github.takahirom.roborazzi:roborazzi-compose (#169)
Bumps [io.github.takahirom.roborazzi:roborazzi-compose](https://github.com/takahirom/roborazzi) from 1.64.0 to 1.66.0.
- [Release notes](https://github.com/takahirom/roborazzi/releases)
- [Commits](https://github.com/takahirom/roborazzi/compare/1.64.0...1.66.0)

---
updated-dependencies:
- dependency-name: io.github.takahirom.roborazzi:roborazzi-compose
  dependency-version: 1.66.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-07 02:22:41 +00:00
Bailey Dixon f5bb41e46d release: android-v1.3.0 + plugin-v1.3.0 (merge dev)
release: android-v1.3.0 + plugin-v1.3.0
2026-07-06 22:16:34 -04:00
Bailey DixonandClaude Fable 5 0b0e323b20 Merge branch 'main' into dev — sync dependabot bumps ahead of the release merge
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 22:16:08 -04:00
Bailey DixonandClaude Fable 5 d8cc3d7082 release(android): android-v1.3.0
Bump to 1.3.0 (code 21); cut the CHANGELOG [1.3.0] block; refresh
RELEASE_NOTES, whats_new, changelog.json, and Play notes; write
plugin-v1.3.0 release notes (with the #165 native-install known-issue
callout — that fix ships in v1.3.1).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 22:00:05 -04:00
Bailey DixonandClaude Fable 5 a8db3a2ed8 docs(todo): voice background-run v2 roadmap — fast lane, queue, tool-output surfacing, deferred concurrency
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 21:56:17 -04:00
Bailey DixonandClaude Fable 5 8dc874cbbf fix(voice): free the floor during background runs — chip-only progress, exit detaches, cancel never clobbers a delivered answer
Background-run progress events (tool.started/tool.delta/run.progress and the
shared emitStatus path) forced VoiceState.Thinking on every tick, pinning the
overlay in spinner+Stop and routing mic taps to the interrupt branch — the
relay's floor was free but the client never returned to Idle, so conversation
couldn't continue during a promoted run. Those paths now update only the
background chip while a run is active; inline (grace-window) turns keep
today's behavior.

Exit and Stop no longer kill a background task: exitVoiceMode detaches (the
relay delivers the result on the next session or as a proactive notification)
and interruptSpeaking silences audio only — the chip's ✕ remains the one
explicit cancel. hermes.run.cancelled in the chat sync now only replaces the
bubble with "Cancelled." when it holds no real content; a delivered answer
keeps its text and gets the Stopped badge, fixing the completed-task-shown-
as-cancelled transcript.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 21:45:21 -04:00
Bailey DixonandClaude Fable 5 ea8d09e7f0 fix(build): 2g test-worker heap — grown suite OOMs Roborazzi renders on the 512m default
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 20:26:47 -04:00
Bailey DixonandClaude Fable 5 5c6211e63f ci(android): add ChatStreamRecoveryTest to the focused test slice
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 20:18:41 -04:00
Bailey DixonandClaude Fable 5 8c8e24c4d0 Merge branch 'fix/chat-stream-recovery' into dev — sessions-stream answer recovery (#166)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 20:18:08 -04:00
Bailey DixonandClaude Fable 5 d16f477d81 fix(chat): settle streaming state on recovery abort + anchor by position (#166)
Aborting an in-flight answer recovery (session/profile switch, new
chat/thread, connection switch) called cancelAnswerRecovery(settleUi=false)
with no live stream, killing the poller but leaving ChatHandler._isStreaming
true and the "Reconnecting to your answer…" turn status frozen forever.
cancelAnswerRecovery now always settles the handler when a poller was
running: a silent clearStreamingStatus() on the abandon paths (no error
badge, per-message flags untouched so cancelStream's Stopped-badge findLast
still works) and the existing placeholder finalize on the new-send path.

The poller also anchored on the pending user message by trimmed-text
indexOfLast, so a short repeated prompt ("yes"/"continue") whose send never
reached the server could match a stale identical earlier row and adopt a
different turn's static (instantly "stable") answer. Recovery now captures
how many user rows existed before the send and requires the pending send to
be the (N+1)-th user row AND match its content; when it can't be established
it never adopts and fails fast to the error UI after confirming across two
polls, instead of polling to the 30-minute cap.

Also adds the clearPendingAsk(deny) parity the sibling stream-error branch
performs to the recovery give-up path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 20:16:55 -04:00
Bailey Dixon 71f3331f54 feat: add multi-device Android bridge targeting 2026-07-06 20:11:45 -04:00
Bailey DixonandClaude Fable 5 e9effeeb87 chore(batch): review nits, CI test-slice additions, changelog/devlog for the issue batch
Review-nit cleanup across the merged branches (KDoc placement, prefill KDoc
accuracy, synthetic test IP, RELEASE.md artifact wording, security.md plain-ws
gating description), ServerAddressTest + IssueReportAndDiagnosticsTest added to
the focused Android CI slice, and consolidated CHANGELOG/DEVLOG/TODO entries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 20:00:18 -04:00
Bailey DixonandClaude Fable 5 e9b00757d5 Merge branch 'fix/onboarding-scroll' into dev — scrollable compact onboarding (#145)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:55:30 -04:00
Bailey DixonandClaude Fable 5 891085ed61 Merge branch 'fix/diagnostics-report-noise' into dev — severity-gated Report flow (#155 #154 #146)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:55:30 -04:00
Bailey DixonandClaude Fable 5 b53cfbc906 Merge branch 'docs/freshness-pass' into dev — UI labels, anchors, stale claims
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:55:30 -04:00
Bailey DixonandClaude Fable 5 045f42386d Merge branch 'fix/release-assets' into dev — 2-asset android releases (#144)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:55:29 -04:00
Bailey Dixon 44e6e9b6dd feat: add model picker refresh parity 2026-07-06 19:51:48 -04:00
Bailey DixonandClaude Fable 5 c46560aeeb fix(onboarding): scroll + compact-height adaptation so slide content is never cut off
Onboarding slide content overflowed below the fold with no scroll
affordance on short viewports or raised font scale (#145).

- OnboardingPage: outer Column now verticalScroll(rememberScrollState())
  per page; Arrangement.Center dropped (the pager's centering Box handles
  short content and Center conflicts under verticalScroll).
- Hero adapts via BoxWithConstraints: 232dp -> 160dp below 620dp available
  height, hidden below 480dp, so typical devices don't need to scroll;
  scrolling remains the safety net for font-scale/foldable extremes.
- OnboardingScreen: bottom nav (pinned outside the pager) tightens its
  48dp bottom padding + indicator spacer under compact screen heights.
- OnboardingCompactScreenshotTest: Roborazzi renders at w320dp-h480dp @
  1.5x font scale and w360dp-h600dp, scrolling to the last body line to
  prove reachability; no golden PNGs, store screenshots untouched.

Fixes #145

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:45:43 -04:00
Bailey DixonandClaude Fable 5 9593787482 docs(user-docs): expand troubleshooting for endpoint and streaming issues
- "No reachable endpoint" subsection: what the diagnostic means and how to
  check each saved route from the phone (LAN vs Tailscale, port 8642).
- Callout: never use localhost/127.0.0.1 as the server address on a phone.
- Tailscale checklist, including that the relay Tailscale helper serves
  the relay + API ports but not the dashboard :9119 (Manage needs it
  reachable separately).
- "Long turns with local models" entry: mid-turn stream drops recover the
  finished answer automatically; screen-on/plugged-in reduces drops.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:30:16 -04:00
Bailey DixonandClaude Fable 5 44ba8ebb41 feat(util): ServerAddress.loopbackHostWarning advisory helper
Pure helper returning an advisory string when a typed server address
parses to a loopback / any-interface host (localhost, 127.x.x.x, ::1,
0.0.0.0) - such an address points at the phone itself and can never reach
the server. Accepts scheme-less input via the existing parseUserInput
normalization; bare "::1" is handled explicitly since it never parses
without brackets. Not wired into any UI yet - wiring is a queued
follow-up owned by the connection-flow workstream.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:30:15 -04:00
Bailey DixonandClaude Fable 5 d25f805fa7 fix(diagnostics): severity-gate the Report flow and fix issue prefill noise
Routine Info-severity probe lines ("Testing API connection") were one tap
away from becoming "[Bug]:"-titled GitHub issues with an unedited
boilerplate body (#155, #154, #146). Now:

- Info entries require a free-text "What were you expecting to happen?"
  answer before the GitHub link is offered; the answer replaces the
  boilerplate "What happened" line (secret-redacted via the shared
  DiagnosticsLog path). Error entries keep the direct flow.
- Info/Warning entries prefill as "[Diagnostic]: <title>" with the existing
  "question" label; Error entries keep "[Bug]:" + "bug".
- The "Connection mode" line now carries the actual route role
  (entry.endpointRole, else inferred from the entry URL, else "unknown")
  instead of the literal "LAN / Tailscale / public TLS / other" template.

Prefill logic extracted to the pure DiagnosticIssuePrefill object so the
title/label/body contract is unit-testable without Compose.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:30:01 -04:00
Bailey DixonandClaude Fable 5 3cef7c39e3 fix(chat): recover a dropped sessions-stream turn from the persisted transcript (#166)
Slow local models + delegating skills outlive the phone's SSE socket
(screen-off / Doze / Wi-Fi power-save — not an OkHttp timeout). Upstream
api_server keeps running the turn after the SSE writer dies and persists
the final answer, but the client finalized the turn as an error after one
immediate history reload that raced the still-running run — stranding the
"Still working…" placeholder forever.

Client-side recovery, standard-path safe (no server/plugin changes):

- ChatStreamRecovery: poll `/api/sessions/{id}/messages` after a
  transport-class (IOException-family) drop on the SESSIONS endpoint —
  5s cadence with exponential backoff to 30s, capped at 30 minutes.
  Finish = a new non-empty assistant message postdating the pending user
  message, stable across two consecutive polls; intermediate persisted
  rows reconcile progressively. A reachable transcript that lacks the
  pending user message fails fast (the run never started).
- Recovered turns finalize with normal-completion side effects
  (turn-complete notification, queued-send drain, session-list refresh)
  via finalizeTurnSideEffects, extracted from onCompleteCb and shared.
  Cap expiry falls back to the existing error UI.
- Exactly one poller per turn; aborted on new send, user Stop, session
  or profile or connection switch, new chat/thread, and VM clear; a late
  onComplete cancels it (double-finalize guard).
- Gateway/runs/completions error paths are unchanged; user cancel keeps
  the existing intentionallyCancelled discipline.
- HermesApiClient: shared streamFailureMessage() + TRANSPORT_ERROR_PREFIX
  so transport failures are distinguishable from server-reported errors.
- UI: streaming placeholder reads "Reconnecting to your answer…" during
  recovery; input caption mirrors it via turnStatus; one diagnostics
  Warning entry when recovery starts. ChatHandler.onStreamError now also
  clears the stale turn-status caption.

Tests: virtual-time poller coverage (backoff cadence, stability window,
fail-fast, cap expiry, cancel) plus Robolectric + MockWebServer
end-to-end coverage (mid-turn socket kill -> polls -> placeholder
completes with the recovered answer; user cancel and new send abort the
poller; cap expiry surfaces the error UI).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:11:52 -04:00
Bailey DixonandClaude Fable 5 c76e906b30 docs: freshness pass — UI labels, deep-link anchors, stale claims
- Quote the real Connect-page button label ("Hermes", renamed from
  "Vanilla Hermes" in v1.2.2) in getting-started's connect and
  manual-setup instructions; concept-term usage left untouched
- Add explicit VitePress heading anchors so existing deep links resolve:
  #sideload-apk, #relay-server-optional, #install-the-server-plugin
  (getting-started), #self-update-hermes-relay-update and
  #install-from-source-node-21 (desktop installation),
  #demo-native-paste-into-the-attached-tui (desktop index)
- Repoint pre-existing dead anchors found by a full built-site anchor
  sweep: features index -> chat #slash-commands (x2) and
  getting-started #_3-connect-chat; desktop index -> tools
  #desktop-open-in-editor-and-interactive-patches
- README: correct registered tool counts (35 android_*, 25 desktop_*)
- docs/security.md: drop prototype framing; describe shipped state —
  plain ws:// consent dialog + TOFU SPKI pinning, per-channel grants and
  Bridge safety rails, on-device Bridge activity log + relay service log
- Home page chat card: gateway WebSocket preferred, HTTP/SSE fallback
- Release tracks: plugin update via hermes plugins install or the
  hermes-relay-update shim (matches install.sh)
- Desktop troubleshooting: pin mismatch text now says SPKI sha256 pin
  (matches certPin.ts), not whole-cert hash

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 18:47:48 -04:00
Bailey DixonandClaude Fable 5 b56dec859d fix(release): attach only sideload APK + googlePlay AAB + SHA256SUMS to android releases
Four assets per android-v* release confused new users: GitHub sorts
assets alphabetically so the non-installable googlePlay .aab listed
first, and the two parity/testing artifacts had no meaning to
non-developers (#144, follow-up from #65).

- release-android.yml: still builds all four artifacts, but attaches
  only the sideload APK, the googlePlay AAB, and SHA256SUMS.txt;
  checksums now cover exactly the attached files. The parity twins
  remain reproducible from the tag via CI. The sideload APK filename
  is unchanged — the in-app UpdateChecker matches ".apk"+"sideload"
  in asset names.
- RELEASE_NOTES.md: replaced the 4-row download table with a lead
  install callout (sideload APK / Google Play) plus an explicit note
  that the .aab is a Play Console upload bundle, not tap-installable.
- RELEASE.md §2: codified the Download-block format and the 2-asset
  policy as the required RELEASE_NOTES format for future releases.

Closes #144

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 18:37:34 -04:00
Bailey Dixon 91763a286f feat: add upstream-impact relay guardrails 2026-07-06 18:05:51 -04:00
dependabot[bot] bc6f288b0b chore(deps): bump androidx.compose:compose-bom in the compose group (#167)
Bumps the compose group with 1 update: androidx.compose:compose-bom.


Updates `androidx.compose:compose-bom` from 2026.06.00 to 2026.06.01

---
updated-dependencies:
- dependency-name: androidx.compose:compose-bom
  dependency-version: 2026.06.01
  dependency-type: direct:production
  dependency-group: compose
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-06 11:54:41 +00:00
dependabot[bot] 96a6c963dc chore(deps): bump io.mockk:mockk from 1.14.9 to 1.14.11 (#162)
Bumps [io.mockk:mockk](https://github.com/mockk/mockk) from 1.14.9 to 1.14.11.
- [Release notes](https://github.com/mockk/mockk/releases)
- [Commits](https://github.com/mockk/mockk/compare/1.14.9...v1.14.11)

---
updated-dependencies:
- dependency-name: io.mockk:mockk
  dependency-version: 1.14.11
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 12:09:31 +00:00
dependabot[bot] 373b98f316 chore(deps): bump kotlinx-coroutines from 1.10.2 to 1.11.0 (#161)
Bumps `kotlinx-coroutines` from 1.10.2 to 1.11.0.

Updates `org.jetbrains.kotlinx:kotlinx-coroutines-android` from 1.10.2 to 1.11.0
- [Release notes](https://github.com/Kotlin/kotlinx.coroutines/releases)
- [Changelog](https://github.com/Kotlin/kotlinx.coroutines/blob/master/CHANGES.md)
- [Commits](https://github.com/Kotlin/kotlinx.coroutines/compare/1.10.2...1.11.0)

Updates `org.jetbrains.kotlinx:kotlinx-coroutines-test` from 1.10.2 to 1.11.0
- [Release notes](https://github.com/Kotlin/kotlinx.coroutines/releases)
- [Changelog](https://github.com/Kotlin/kotlinx.coroutines/blob/master/CHANGES.md)
- [Commits](https://github.com/Kotlin/kotlinx.coroutines/compare/1.10.2...1.11.0)

---
updated-dependencies:
- dependency-name: org.jetbrains.kotlinx:kotlinx-coroutines-android
  dependency-version: 1.11.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
- dependency-name: org.jetbrains.kotlinx:kotlinx-coroutines-test
  dependency-version: 1.11.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 12:07:50 +00:00
dependabot[bot] 83609a4608 chore(deps): bump com.meta.spatial:spatial-gradle-plugin-impl (#163)
Bumps com.meta.spatial:spatial-gradle-plugin-impl from 0.12.0 to 0.13.1.

---
updated-dependencies:
- dependency-name: com.meta.spatial:spatial-gradle-plugin-impl
  dependency-version: 0.13.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 12:07:15 +00:00
dependabot[bot] 2c1d7df764 chore(deps): bump gradle-wrapper from 9.6.0 to 9.6.1 (#159)
Bumps [gradle-wrapper](https://github.com/gradle/gradle) from 9.6.0 to 9.6.1.
- [Release notes](https://github.com/gradle/gradle/releases)
- [Commits](https://github.com/gradle/gradle/compare/v9.6.0...v9.6.1)

---
updated-dependencies:
- dependency-name: gradle-wrapper
  dependency-version: 9.6.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 12:03:12 +00:00
dependabot[bot] df536b308f chore(deps): bump markdown-renderer from 0.42.0 to 0.43.0 (#160)
Bumps `markdown-renderer` from 0.42.0 to 0.43.0.

Updates `com.mikepenz:multiplatform-markdown-renderer-m3` from 0.42.0 to 0.43.0
- [Release notes](https://github.com/mikepenz/multiplatform-markdown-renderer/releases)
- [Changelog](https://github.com/mikepenz/multiplatform-markdown-renderer/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/mikepenz/multiplatform-markdown-renderer/compare/v0.42.0...v0.43.0)

Updates `com.mikepenz:multiplatform-markdown-renderer-code` from 0.42.0 to 0.43.0
- [Release notes](https://github.com/mikepenz/multiplatform-markdown-renderer/releases)
- [Changelog](https://github.com/mikepenz/multiplatform-markdown-renderer/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/mikepenz/multiplatform-markdown-renderer/compare/v0.42.0...v0.43.0)

---
updated-dependencies:
- dependency-name: com.mikepenz:multiplatform-markdown-renderer-code
  dependency-version: 0.43.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
- dependency-name: com.mikepenz:multiplatform-markdown-renderer-m3
  dependency-version: 0.43.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 12:01:08 +00:00
dependabot[bot] 65f89d4c0b chore(deps): bump io.github.takahirom.roborazzi:roborazzi-compose (#158)
Bumps [io.github.takahirom.roborazzi:roborazzi-compose](https://github.com/takahirom/roborazzi) from 1.43.1 to 1.64.0.
- [Release notes](https://github.com/takahirom/roborazzi/releases)
- [Commits](https://github.com/takahirom/roborazzi/compare/1.43.1...1.64.0)

---
updated-dependencies:
- dependency-name: io.github.takahirom.roborazzi:roborazzi-compose
  dependency-version: 1.64.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 11:58:58 +00:00
dependabot[bot] e78602c9d8 chore(deps): bump androidx.core:core-ktx from 1.18.0 to 1.19.0 (#157)
Bumps androidx.core:core-ktx from 1.18.0 to 1.19.0.

---
updated-dependencies:
- dependency-name: androidx.core:core-ktx
  dependency-version: 1.19.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 11:54:46 +00:00
Bailey Dixon fd2a8b2546 Merge pull request #152 from Codename-11/dev
release: dev → main — activate dev-loop polish (github-script bump + deep-dive formatting)
2026-06-28 13:23:38 -04:00
Bailey Dixon 37ff218f33 Merge pull request #149 from Codename-11/dev
release: dev → main — activate Claude triage/review dev-loop (no version bump)
2026-06-28 12:07:00 -04:00
Bailey Dixon 45e7911d3d Merge pull request #143 from Codename-11/dev
release(android): android-v1.2.6
2026-06-27 23:54:22 -04:00
Bailey Dixon 1ae0627d92 Merge pull request #140 from Codename-11/dev
release(android): android-v1.2.5
2026-06-27 19:08:23 -04:00
Bailey Dixon de44059c8e Merge pull request #135 from Codename-11/dev
ci: activate issue triage on main (+ v1.2.4 devlog)
2026-06-27 12:17:23 -04:00
Bailey Dixon 0327012666 release(android): android-v1.2.4 (#130)
release(android): android-v1.2.4
2026-06-25 21:36:31 -04:00
Bailey Dixon 4160cb1f85 feat(android): add notification trigger MVP 2026-06-23 08:58:48 -04:00
155 changed files with 17266 additions and 2099 deletions
+3
View File
@@ -146,6 +146,9 @@ jobs:
--tests com.hermesandroid.relay.network.ArchitectureBoundaryTest \
--tests com.hermesandroid.relay.network.relay.RelayUrlDeriverTest \
--tests com.hermesandroid.relay.viewmodel.ConnectionSwitchTest \
--tests com.hermesandroid.relay.util.ServerAddressTest \
--tests com.hermesandroid.relay.util.IssueReportAndDiagnosticsTest \
--tests com.hermesandroid.relay.viewmodel.ChatStreamRecoveryTest \
--console=plain
# Upload reports only for failures. Successful PR report uploads add
+2 -8
View File
@@ -12,10 +12,7 @@ on:
push:
branches: [main, dev]
paths:
- "plugin/__init__.py"
- "plugin/android_tool.py"
- "plugin/cli.py"
- "plugin/pair.py"
- "plugin/*.py"
- "plugin/plugin.yaml"
- "plugin/relay/**"
- "plugin/tools/**"
@@ -31,10 +28,7 @@ on:
pull_request:
branches: [main, dev]
paths:
- "plugin/__init__.py"
- "plugin/android_tool.py"
- "plugin/cli.py"
- "plugin/pair.py"
- "plugin/*.py"
- "plugin/plugin.yaml"
- "plugin/relay/**"
- "plugin/tools/**"
+19 -11
View File
@@ -127,10 +127,12 @@ jobs:
# Flavor dimension adds an extra path segment to the AGP output layout.
# APKs live under `apk/<flavor>/release/`, AABs under `bundle/<flavor>Release/`
# (note the concatenated camelCase — AGP path quirk, documented but
# different between APK and AAB). The globs below match both flavors.
# different between APK and AAB). Checksums cover EXACTLY the files
# attached to the GitHub Release (see the 2-asset policy on the
# release step below) so SHA256SUMS.txt matches the assets 1:1.
run: |
cd app/build/outputs
sha256sum apk/*/release/*.apk bundle/*Release/*.aab > SHA256SUMS.txt
sha256sum apk/sideload/release/*.apk bundle/googlePlayRelease/*.aab > SHA256SUMS.txt
cat SHA256SUMS.txt
- name: Create GitHub Release
@@ -140,16 +142,22 @@ jobs:
tag_name: android-v${{ needs.validate.outputs.version }}
body_path: RELEASE_NOTES.md
prerelease: ${{ contains(needs.validate.outputs.version, '-') }}
# Attach all four flavored artifacts — users sideload the
# `hermes-relay-<version>-sideload-release.apk` for the full
# Phase 3 / Tier 3/4/6 feature set; the
# `hermes-relay-<version>-googlePlay-release.aab` is what gets
# uploaded to Play Console. APK twin of the googlePlay flavor
# and AAB twin of the sideload flavor are included for parity
# (useful for diff tooling, not primary downloads).
# Deliberate 2-asset policy (#144): attach ONLY
# `hermes-relay-<version>-sideload-release.apk` (the file users
# install by tapping — full Device Control feature set) and
# `hermes-relay-<version>-googlePlay-release.aab` (the Play Console
# upload bundle — NOT tap-installable on a phone), plus the
# SHA256SUMS.txt covering exactly those two files. GitHub sorts
# assets alphabetically, so extra files made the non-installable
# .aab list first and confused new users. The parity twins
# (googlePlay APK, sideload AAB) are still BUILT by the step above
# and reproducible from the tag via CI, just not attached.
# NEVER rename the sideload APK: the in-app update checker
# (update/UpdateChecker.kt) matches assets by ".apk" + "sideload"
# in the name, and user-docs verify steps cite the filename.
files: |
app/build/outputs/apk/*/release/*.apk
app/build/outputs/bundle/*Release/*.aab
app/build/outputs/apk/sideload/release/*.apk
app/build/outputs/bundle/googlePlayRelease/*.aab
app/build/outputs/SHA256SUMS.txt
- name: Upload to Play Console (production draft)
+61 -1
View File
@@ -6,6 +6,58 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this
## [Unreleased]
## [1.4.0] - 2026-07-09
### Added
- **Android model pickers can refresh the server catalog.** Chat's model sheet and Manage's main/profile model dialogs now expose upstream's explicit **Refresh Models** action, so dynamic/custom provider model lists can be reloaded on demand without making every picker open probe providers.
- **Server-backed session cleanup plumbing.** The dashboard client now supports single-session export, the upstream `/api/sessions/prune` route with a mandatory dry-run preview before destructive apply, plus soft archive/restore helpers and an `archived` session-list filter for the Manage surface.
- **Notification triggers MVP.** Settings → Notifications now has explicit opt-in proactive rules for the Notification companion: match by app package plus optional title/text filters, post a safe local "Ask Hermes?" prompt, show the latest trigger activity, and pause everything instantly with a kill switch.
- **Android bridge: multi-device targeting.** The relay can keep multiple Android bridge clients connected at once, route commands by `device` selector (`phone`, `pixel`, `fold`, `boox`, `note`, `notemax`, `tablet`, or device ID), expose `/bridge/devices` and `/bridge/select-active`, and advertise an optional `device` argument on the `android_*` tool schemas.
- **Voice: a second long request gets queued, not refused.** Ask for another long task while one is already running in the background and it's now queued (up to three) and starts automatically when the current one finishes — with a short spoken transition. The task card shows "+N queued", and cancelling the current task clears the queue.
- **Voice: background answers start speaking sooner and can never be silently lost.** The spoken summary now streams as it's generated (it used to be held until fully complete — a noticeable dead gap, then the whole answer at once). Delivery is verified two ways: the summary must actually reflect the answer's content (not just avoid known filler phrases), and if no spoken delivery lands within 30 seconds the answer is posted as text instead of vanishing.
- **Voice: tap the finished-task card to hear the answer again.** After a background task's card settles to "finished," tapping it replays the delivered answer. The card also now shows in the compact voice view (it previously existed only in the full-screen layout), a "Drafting the answer…" status appears as the reply is being composed, and leaving voice mode with a task still running leaves a note in chat so the work stays visible.
- **Voice: quick questions answered while a background task runs.** Realtime voice used to refuse *any* second request while a long task ran in the background — even a two-second lookup. A quick second ask is now answered inline on a side session (within the same few-second window that decides backgrounding); anything that turns out to be long still gets the "a task is already running" answer, and the running task is never disturbed.
- **Voice: the background-task card no longer vanishes mid-answer.** The card used to disappear the instant the spoken answer started (exactly when the waveform returned), reading as the task being lost. It now settles to a "Background task finished." state, lingers for a few seconds while the answer plays, then dismisses itself — and its ✕ during that settled state just dismisses the card instead of sending a cancel.
- **Voice: the "Thinking" pill no longer spins forever.** The server streams its drafting text as an internal pseudo-tool that never reports completion, and the app rendered it as a live tool pill — which then ran indefinitely in both chat and the voice overlay. Internal tool events no longer become pills (their text still feeds the thinking trace).
- **Voice: background-task answers can't be lost to a stray cancel.** Tapping cancel/stop after a background task had already finished used to mark the finished run "cancelled" — losing the answer that was about to be spoken. Cancel now only cancels a run that's actually still running; stopping the current speech works as before.
- **Voice: no more spoken run IDs or phantom queue state.** The realtime voice model no longer reads 32-character run IDs aloud after starting a background task (identifiers stay out of everything it's asked to speak), no longer claims a request was queued unless the relay accepted it, and a completed task's answer is spoken directly — deferral filler like "one moment while I look that up" in place of a finished result now triggers the fallback that speaks the real answer.
- **Voice: finished-task answers keep the realtime voice.** A completed background task's answer is now spoken by the same realtime voice you've been talking to — read word for word from the authoritative Hermes answer — instead of switching to the standard TTS voice mid-conversation. The answer always lands: if the realtime model goes off-script or the provider connection drops, standard TTS speaks it, and if you start talking mid-delivery it's posted as text instead of interrupting you. The "When the answer is ready" setting keeps its four modes (Exact / Summary / Notify / Show), now explained behind an info icon in Voice Settings.
- **Voice: realtime models refreshed.** OpenAI realtime now defaults to `gpt-realtime-2.1` (with the cheaper `gpt-realtime-2.1-mini` selectable), the versioned `grok-voice-think-fast-1.0` pin is available alongside xAI's `grok-voice-latest` alias, and session logs record which model the provider *actually* served — so provider-side alias moves no longer happen invisibly.
- **Voice: session logs clean up after themselves.** Realtime voice session logs are swept after 14 days by default (`realtime_voice.run_retention_days`, 0 disables), and the per-response TTS audio capture is now opt-in debug tooling (`debug_audio_tap`) instead of an always-on multi-MB tap.
- **Voice: one-command delivery health report.** `python -m plugin.relay.realtime_agent.report` summarizes recent voice deliveries — how many were spoken by the realtime voice vs fell back to TTS or text, and why — for quick health checks after live testing.
### Changed
- **Bootstrap compatibility layer slimmed to true gaps.** The optional compatibility hook no longer injects session CRUD/messages or the legacy skills list — current Hermes serves those natively; it now covers only surfaces with no native replacement yet (session search, memory, legacy skill detail/toggle, config, available-models, and the slash-command middleware). Older pre-session-API Hermes builds degrade to the standard completions/runs chat paths.
- **Dependency floor: aiohttp ≥ 3.14.1.** Raised from 3.9 across plugin requirements and package metadata to the patched line covering the 2026 aiohttp security advisories.
### Fixed
- **Realtime voice recovers after background route loss.** A recorded turn now waits for a relay-confirmed resumed socket, retains unacknowledged follow-up PCM for replay, and reports transport rejection instead of sitting on a dead persistent connection. Resume handshakes are coalesced, and the relay requires a valid resume claim before replacing the active phone socket, so a slower stale connection cannot detach background-result delivery. Long-lived sessions start their bounded retry window when the route actually drops instead of at voice-mode entry, and a bare socket open cannot reset it. Late callbacks from a retired session are ignored. Exiting voice mode clears its detached reconnect and confirmation state before another session opens; rejected or unacknowledged cancels no longer leave an undismissable background-task pill. Provider transcription no longer impersonates active microphone capture, Stop settles the local turn even when the route is gone, and provisional `Listening...` / `Still working...` rows cannot remain stuck in chat.
- **xAI exact background answers bypass model deferral.** Non-structured **Exact** deliveries now use xAI's provider-native forced speech event, preserving the selected realtime voice and normal assistant history while speaking the authoritative Hermes answer without asking the model to follow a read-verbatim prompt. Structured results and summary modes still use natural model summarization, and the validator plus standard-TTS fallback remain as safety nets.
- **Background voice handoffs no longer repeat themselves.** If the realtime provider already spoke an acknowledgement before calling Hermes, promotion keeps that first line and suppresses the redundant "running in the background" follow-up; silent tool calls still receive the configured spoken handoff. Provider protocols that report both response creation and output-item creation now also produce one client `response.started` event instead of two.
- **Realtime voice model and voice picks now apply to the next session.** Voice Settings persists the selected Realtime Agent model and voice per connection/profile and sends both when opening a session, so choosing a pinned model immediately controls the next session instead of requiring **Save realtime agent** to rewrite the relay config. The active voice UI reflects the override, changing it retires any prewarmed session, and the choice survives an app restart.
- **Fresh realtime sessions emit one ready event.** Android's required `session.start` acknowledgement no longer causes the relay to send a second `voice.session.ready`, avoiding duplicate event IDs and duplicate session-ready telemetry on every new voice conversation.
- **Relay media can no longer serve credential files.** `/media/by-path` now always blocks paths that resolve into credential or system locations (`~/.hermes/.env`, `auth.json`, `config.yaml`, OAuth/MCP token stores, `pairing/`, `~/.ssh`, and similar) even in the default permissive mode — mirroring upstream Hermes' media-delivery hardening — so a prompt-injected `MEDIA:` marker can't deliver live secrets to a paired phone. Symlinks are resolved before the check, and the relay's own QR-signing secret and session-token store are covered too.
- **Long agent turns no longer die or duplicate at the transport.** Gateway chat (Android and the desktop CLI) now gives `prompt.submit` up to 30 minutes to acknowledge — matching upstream desktop and the server's own turn ceiling — instead of short generic RPC timeouts that could falsely fall back to SSE (duplicating the turn on Android) or kill a legitimately long deep-reasoning turn. Turn liveness is governed by idle-progress watchdogs (no events at all for a stretch), never a hard cap while output is still streaming.
- **Manage → Models keeps providers that still need keys.** Newer Hermes hides unconfigured providers from the model catalog unless a management UI opts in; Android Manage now opts in and keeps rendering greyed provider rows with their key-setup guidance on both old and new servers. In-chat model picking is unchanged (configured providers only).
- **Phone-local context actually reaches the server on fallback chat paths.** The sessions/runs streaming payloads carried voice-intent traces, card dispatches, and attachments in fields the server never reads — silently dropping them. That context now rides channels the server actually consumes (a per-turn context digest, real history fields where they exist, inline images on the completions path), and any attachment with no supported channel is reported instead of silently discarded.
- **Relay plugin works under the native `hermes plugins install` path.** The plugin's runtime imports assumed the repo's editable layout, so upstream's native installer (which loads plugins under its own package namespace) broke `hermes relay start` and `hermes pair` with `ModuleNotFoundError: No module named 'plugin'`. All runtime imports are now package-relative, the dashboard module boots correctly when the upstream web server loads it standalone, and `hermes relay doctor` now exercises the real import chain so this class of breakage can't pass doctor again. (#165)
- **Installer handles modern venv layouts.** `install.sh` now autodetects the classic venv, uv-managed `.venv`, and containerized layouts — and everything it generates (the systemd unit and all four command shims) points at the interpreter it actually detected instead of a hardcoded classic path. On immutable container images it steers to the native install path with a clear message instead of dying mid-run. (#165)
- **Doctor catches dashboard URLs pointed at the wrong Hermes surface.** `hermes relay doctor` now distinguishes the dashboard/Manage surface from an API-server/headless backend URL and tells operators to use `hermes dashboard` when a configured dashboard URL is actually pointing at `hermes serve` / the API server.
- **Doctor and installer catch stale duplicate plugin copies.** The gateway plugin loader picks a discovered plugin by manifest name, so a second directory declaring `name: hermes-relay` (a leftover backup copy or a stray extra install) could win and make the gateway load stale code — silently ignoring every later deploy. `hermes relay doctor` now warns when more than one directory under the plugins dir declares the same plugin name, and `install.sh` removes any such duplicate so only the canonical plugin symlink remains.
- **Crash-safety on Android 14 and earlier.** Built against SDK 35, Kotlin's `removeFirst()`/`removeLast()` resolve to the new Java `List` methods that don't exist below Android 15, crashing older devices. All such calls in the app are now `removeAt(...)`, and Tink (pulled in by encrypted storage) is pinned ahead of the transitive version whose `HybridConfig` tripped the same Google Play pre-launch check.
- **No crash when a relay address is malformed.** A corrupt or hand-edited pairing address with an invalid host could crash the app the moment it opened the relay connection (the connection is built on a background thread, so the error escaped uncaught). A bad relay address is now handled as a normal connection failure — shown as disconnected with a "re-pair to refresh" note — instead of crashing. The same guard now also covers the relay's media, session, and voice HTTP calls. (relay half of #131)
- **Voice: cleaner error recovery.** A failed or timed-out voice turn no longer shows the same error twice (the top overlay banner and a duplicate bottom banner) and can now be **dismissed**, not just retried — so a stuck error state can't block the screen.
- **Voice: fallback-spoken answers no longer play into a frozen overlay.** When an answer is delivered by the standard TTS fallback (or replayed from the finished-task card), the voice screen now shows the waveform and the answer text while it speaks — previously it sat on "Thinking" with no visuals even though audio was playing.
- **Voice: a quiet realtime session no longer dies with a raw provider error.** xAI ends a realtime conversation after 900 seconds of inactivity, and no keepalive traffic resets that timer — so a voice session left open through a long background task (or simply left open) died with a raw provider error. That provider timeout is now treated as routine expiry: the session ends cleanly with no error banner, and your next voice turn transparently opens a fresh provider conversation that picks up from the same durable Hermes chat session.
- **No crash when a malformed server address reaches a chat send.** The three streaming chat paths built their HTTP request before any error handling, so a corrupt or hand-edited API URL could throw instead of failing the turn gracefully. They now surface "Invalid server address — edit the connection's API URL or re-pair" through the normal in-chat error channel (closes the remaining #131 crash-class gap).
- **Demo mode: typing a message now gets an honest reply.** Sending a message in the offline demo used to do nothing (the composer silently ignored it, reading as broken). The demo now echoes your message and answers with a short notice explaining it's an offline sample, pointing at the Connect action to chat for real.
- **Voice: realtime conversations reliably reach your chat history.** Turns the realtime voice model answers directly (without calling Hermes) are folded into the chat session on your next message — but on the default gateway connection that hand-off could be deferred indefinitely, so the agent never learned what was said in voice. The turn that carries them now routes so the sync actually lands. Synced voice turns also render cleanly when a chat reloads: a quiet "Realtime Agent" chip instead of a raw provenance footnote, and no more duplicated voice exchange after the sync.
## [1.3.0] - 2026-07-06
### Added
- **Voice settings: edit your server's voice engine.** Voice settings now has a **Server voice config** section that reads and writes the host's text-to-speech and speech-to-text settings — provider, voice, model, language, and per-provider options — over the dashboard, the same config the official desktop app edits. It includes an **ElevenLabs voice picker** that lists the voices available on your server's ElevenLabs key (and tells you when no key is set). Works on the no-plugin (Standard) path; sign in to Manage to use it.
@@ -24,6 +76,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this
### Changed
- **Reporting a diagnostic now files the right kind of issue.** The Report button on a diagnostics entry used to turn routine log lines into "[Bug]" GitHub issues with an empty template. Now informational entries first ask "what were you expecting to happen?" and file as a "[Diagnostic]" question, error entries keep the direct bug flow, and every report carries the connection mode you were actually on instead of a placeholder line. (#155, #154, #146)
- **Simpler release downloads.** Each Android release on GitHub now attaches just two files — the tap-to-install sideload APK and the Play Store upload bundle — plus checksums, with the release notes leading with the one file most people want. The extra "parity/testing" artifacts are gone from the release page (still reproducible from the tag via CI). (#144)
- **Clearer, snappier voice capture and playback.** Voice now engages the device's echo-cancellation and noise-suppression while recording (matching the desktop's microphone setup), and requests audio focus before the first reply so the opening words aren't clipped on a cold start. Listening timing also matches the official desktop: auto-stop ~1.25s after you stop speaking (was 3s), give up after 12s with no speech, and cap a turn at 60s.
- **Refreshed chat look.** Message bubbles are wider and denser, each assistant turn shows a small Hermes avatar to its left (once per group), and code blocks are richer — a language label, a copy button, and a clearer inset so fenced code and inline `code` no longer blend into the bubble.
- **Desktop CLI: visual + ergonomics refresh.** A single color theme across the CLI, aligned tables for `devices`/`sessions`, status dots for on/off states, and progress spinners for slow operations (the multi-endpoint pairing probe and the gateway connect) so nothing looks hung. Errors now suggest the fix (e.g. re-pair on auth failure).
@@ -43,6 +97,11 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this
### Fixed
- **Realtime voice: you can keep talking while a background task runs.** Progress updates from a background task were flipping the voice UI back into "Thinking" with a Stop button on every tick, so the mic never came back until the task finished. Progress now feeds only the task chip; the conversation stays open the whole time.
- **Realtime voice: leaving voice mode no longer cancels a running task.** Exiting (or tapping Stop to interrupt speech) used to kill an in-flight background task and could overwrite its already-delivered answer with "Cancelled." in the chat. Exit now detaches — the task keeps running and the result arrives on your next session or as a notification — and a delivered answer always keeps its text (a Stopped badge marks a genuine cancel). The chip's ✕ remains the one deliberate way to cancel.
- **Long answers are no longer lost when the connection drops mid-turn.** On slow local models (or skills that delegate long background work), the phone could drop the stream mid-turn — the server finishes and saves the answer, but the chat sat on "Still working…" forever. The app now detects the dropped stream and quietly re-checks the conversation until the finished answer arrives, then completes the turn normally (with the usual done-notification if you've backgrounded the app). Switching chats or sending something new cancels the wait. (#166)
- **Onboarding slides fit every screen.** Intro slide text could run past the bottom of the screen with no way to scroll on short displays or large font sizes. Slides now scroll when needed and compact their artwork on short viewports, so no setup guidance is unreachable. (#145)
- **Docs: fixed stale setup labels and broken links.** The setup guide referenced a "Vanilla Hermes" button the app hasn't shown since v1.2.2 (it's labeled "Hermes"), several deep links into the getting-started page were dead, and the README under-counted the available phone tools. (docs site)
- **Back button on Manage and Bridge now works.** The back arrow on the Manage ("Hermes management") and Bridge screens did nothing — it tried to jump to Chat in a way that silently no-op'd. Back now reliably returns to the screen you opened it from.
- **Dropped relay connections from a status-report race.** The phone's periodic device-status report could occasionally be sent to the relay *before* the connection had finished authenticating, which made the relay reject the whole connection and forced a reconnect. The app now holds every message until the connection is authenticated, so the handshake always completes first.
- **Fewer needless connection re-checks when switching apps.** Returning to the app after a quick glance at another app no longer triggers a full connection re-probe (and the brief "checking…" flash) when the connection was already healthy — it only re-checks after a longer absence or if something actually looks off.
@@ -1444,7 +1503,8 @@ MVP release — native Android companion app for Hermes agent with direct API ch
- **Dev scripts** — build, install, run, test, relay via scripts/dev.bat
- **ProGuard rules** — okhttp-sse, markdown renderer, intellij-markdown parser
[Unreleased]: https://github.com/Codename-11/hermes-relay/compare/android-v1.0.0...HEAD
[Unreleased]: https://github.com/Codename-11/hermes-relay/compare/android-v1.4.0...HEAD
[1.4.0]: https://github.com/Codename-11/hermes-relay/compare/android-v1.3.0...android-v1.4.0
[1.0.0]: https://github.com/Codename-11/hermes-relay/compare/android-v0.8.0...android-v1.0.0
[0.8.1]: https://github.com/Codename-11/hermes-relay/compare/android-v0.8.0...android-v0.8.1
[0.8.0]: https://github.com/Codename-11/hermes-relay/compare/v0.7.0...android-v0.8.0
+7 -7
View File
@@ -46,20 +46,20 @@ The Vanilla Hermes path must stay upstream-only. API-server bearer auth and dash
Upstream main now contains the focused session-control API (`#33134`) and read-only skills/toolsets (`#33016`). The original broad PR [#8556](https://github.com/NousResearch/hermes-agent/pull/8556) was closed as superseded. Keep these distinctions straight:
1. **Native upstream** — `/api/sessions`, `/api/sessions/{id}/messages`, `/api/sessions/{id}/chat`, `/api/sessions/{id}/chat/stream`, `/v1/capabilities`, `/v1/skills`, and `/v1/toolsets` exist in current `gateway/platforms/api_server.py`.
2. **Bootstrap compatibility** (`plugin/hermes_relay_bootstrap/`) — monkey-patches aiohttp on startup via `.pth` file for older or partial core builds. It skips native routes per method/path and should be retired per surface, not treated as the preferred path. The repo-root `hermes_relay_bootstrap/` package is a legacy import shim.
2. **Bootstrap compatibility** (`plugin/hermes_relay_bootstrap/`) — monkey-patches aiohttp on startup via `.pth` file, injecting only compatibility-only surfaces (session search, memory, legacy skill detail/toggle, config, available-models, slash middleware). Sessions CRUD/messages/fork and the legacy skills list are **retired** — native upstream owns them (#33134/#33016) and the bootstrap carries no fallback for old builds. Native routes still win per method/path for the remaining set. The repo-root `hermes_relay_bootstrap/` package is a legacy import shim.
3. **Legacy fork branches** — useful as lineage only. Do not cite `feat/session-api` / `#8556` as the current upstream contract.
| Endpoint | Purpose | Provided by |
| -------------------------------------- | -------------------------------------- | -------------------------------------------------------------------------------------- |
| `GET /api/sessions` (CRUD) | Session list/create/rename/delete/fork | Native upstream (#33134); bootstrap only for old builds |
| `GET /api/sessions/{id}/messages` | Conversation history | Native upstream (#33134); bootstrap only for old builds |
| `GET /api/sessions` (CRUD) | Session list/create/rename/delete/fork | Native upstream (#33134); bootstrap injection retired |
| `GET /api/sessions/{id}/messages` | Conversation history | Native upstream (#33134); bootstrap injection retired |
| `POST /api/sessions/{id}/chat` | Synchronous session chat | Native upstream (#33134) |
| `POST /api/sessions/{id}/chat/stream` | Session-based SSE chat | Native upstream (#33134); bootstrap does NOT inject |
| `GET /v1/skills`, `GET /v1/toolsets` | Read-only skill/toolset discovery | Native upstream (#33016) |
| `GET /api/sessions/search` | Full-text message search | Bootstrap/fork legacy; not in current upstream main |
| `GET /api/config`, `PATCH /api/config` | Personalities + model config | Bootstrap/fork legacy or dashboard web-server surface; not current API-server upstream |
| `GET /api/skills`, `/{name}` | Legacy skill discovery/detail | Bootstrap/fork legacy; prefer native `/v1/skills` for lists |
| `GET /api/skills/{name}` | Legacy skill detail | Bootstrap compat; list (`GET /api/skills`) retired — use native `/v1/skills` |
| `PUT /api/skills/toggle` | Enable/disable installed skill | `hermes_cli/web_server.py` dashboard surface; bootstrap stub returns 501 |
| `GET/POST/PATCH/DELETE /api/memory` | Memory CRUD | Bootstrap/fork legacy; not current API-server upstream |
| `GET /api/available-models` | Provider model list | Bootstrap/fork legacy; not current API-server upstream |
@@ -84,7 +84,7 @@ Current upstream supports two auth modes on this surface. Loopback dashboards st
- **Vanilla Hermes path = upstream-only.** The default (no-plugin) connection path — gateway/API chat, Manage, and Vanilla Hermes voice via the dashboard surface — must work against **unmodified upstream hermes-agent**: no fork patches, no bespoke server config as a dependency. The app ships on Google Play to users whose servers we don't control. Features that need server-side changes go through upstream PRs (with graceful degradation until merged) or live behind the opt-in relay plugin.
- **Always verify upstream before assuming an endpoint exists.** Check `gateway/platforms/api_server.py` in hermes-agent. If an endpoint isn't there, document whether bootstrap injects it or it requires the fork.
- If we use a non-standard endpoint, ensure `probeCapabilities()` covers it and the auto-resolver degrades gracefully.
- **Bootstrap maintenance:** Retire `plugin/hermes_relay_bootstrap/` per surface. Sessions and read-only skills/toolsets now have native upstream replacements; config, memory, legacy skill detail/toggle, available-models, and slash middleware still need explicit replacement decisions before full removal.
- **Bootstrap maintenance:** Retire `plugin/hermes_relay_bootstrap/` per surface. Done: sessions CRUD/messages/fork and the legacy skills list are retired from the bootstrap (native upstream #33134/#33016, no old-build fallback kept). Remaining: config, memory, legacy skill detail/toggle, available-models, session search, and slash middleware still need explicit replacement decisions before full removal.
## Repository Layout
@@ -282,7 +282,7 @@ This is a **public, distributed repo** — every committed file (CHANGELOG, DEVL
| `plugin/pair.py` | QR payload builder + CLI; `build_payload(sign=True)`; `--register-code` fallback |
| `plugin/doctor.py` | `hermes relay doctor`; checks standard upstream API/dashboard reachability, Relay loopback state, plugin layout, and compat hook state |
| `plugin/compat.py` | `hermes relay compat status/install/remove`; owns the optional `hermes_relay_bootstrap.pth` lifecycle |
| `plugin/hermes_relay_bootstrap/` | Plugin-owned runtime compatibility patch; skips native routes per method/path; retire only after remaining config/memory/legacy skill/slash gaps are handled |
| `plugin/hermes_relay_bootstrap/` | Plugin-owned runtime compatibility patch — compat-only surfaces (session search, memory, skill detail/toggle, config, available-models, slash middleware); sessions + skills-list injection retired (#33134/#33016) |
| `install.sh` | Canonical installer — 6 steps; idempotent; drops `hermes-relay-update` shim |
| `uninstall.sh` | Canonical uninstaller; reverses install.sh; never touches `.env` or `state.db` |
| `hermes_relay_bootstrap/` | Legacy import shim for old `.pth` files and editable installs |
@@ -473,7 +473,7 @@ See [RELEASE.md](RELEASE.md) for the full recipe.
| Chat streaming | `POST /v1/runs` → `GET /v1/runs/{id}/events` | Structured tool events; async run-control path |
| Chat (sessions) | `POST /api/sessions/{id}/chat/stream` | Native upstream session-persisted SSE; preferred when capability probe finds it |
| Chat (compat) | `POST /v1/chat/completions` (stream=true) | Inline tool annotations only |
| Session CRUD | `GET/POST/PATCH/DELETE /api/sessions` | Native upstream (#33134); bootstrap fallback only for old builds |
| Session CRUD | `GET/POST/PATCH/DELETE /api/sessions` | Native upstream (#33134); bootstrap fallback retired |
| Manage | Dashboard `/api/status`, `/api/auth/me`, `/api/config`, `/api/profiles/*`, `/api/env`, `/api/model/*`, `/api/mcp/*` | Vanilla Hermes dashboard surface; do not proxy through Relay |
| Vanilla Hermes voice | Dashboard `POST /api/audio/transcribe`, `POST /api/audio/speak` | Vanilla Hermes no-plugin voice; uses dashboard session from Manage |
| Pairing (QR) | `POST /pairing/register` (loopback only) | Via `/hermes-relay-pair` or `hermes-pair` shim; accepts optional `endpoints` for multi-endpoint QRs |
+1094
View File
File diff suppressed because it is too large Load Diff
+38 -11
View File
@@ -1,29 +1,56 @@
# Hermes-Relay-Plugin v__VERSION__
**Release Date:** June 22, 2026
**Since the previous plugin release:** Reliability fixes for the Realtime Agent voice path — brokered Hermes turns no longer drop with `session_not_found`, and long-running Hermes work no longer times out a live voice session.
**Release Date:** July 9, 2026
This is a focused patch for the relay's Realtime Agent. When a spoken turn reached back into Hermes for context or tool work, a session-namespace mismatch could make the API Server reject the turn, and long background tasks could let the voice session lapse mid-run. Both paths are now resilient. Provider-native voice turns and vanilla upstream (no plugin) are unaffected.
**Since v1.3.0:** Realtime Agent background work gains queued long requests, quick side-session answers, deterministic exact xAI delivery, stronger resume ownership, and a delivery-health report. The plugin now installs through upstream Hermes' native plugin path, handles modern virtual-environment layouts, targets multiple Android devices, protects credential paths in media delivery, and adds sharper doctor checks.
Pairs with Hermes-Relay-Android v1.4.0 for the matching background-task, model/voice selection, resume, and task-chip behavior. Standard chat and Vanilla Hermes voice remain upstream-owned and do not require this plugin.
## What's changed
### Fixed
- **Brokered Hermes turns no longer fail with `session_not_found`.** When the Realtime Agent reached back to Hermes for context or tool work, it could hand the API Server a session id from a different session namespace (the gateway/client store), which the API Server rejected. The broker now mints a valid API Server session and retries the turn once when that happens, reuses an existing API Server session when the id is already valid, and reads the API Server's current nested `{"session": {"id": …}}` create-session response (previously only the legacy flat shape) so session creation no longer errors with "created a session without an id."
- **Realtime voice survives long Hermes runs.** A heartbeat now keeps the realtime voice session alive while a long-running Hermes task is in flight, so the turn no longer times out before the work finishes.
### Added
## Install
- **Queued background voice work.** Up to three additional long requests can wait behind an active Hermes task and start automatically in order; cancelling the active run also clears its queue.
- **Quick side-session answers.** A short follow-up can be answered while a background run continues, without disturbing the durable task or its eventual delivery.
- **Provider-native exact xAI delivery.** Exact non-structured results use xAI's forced speech event so the selected realtime voice reads the authoritative Hermes answer without another model inference step.
- **Multi-device Android Bridge.** Multiple Android clients can remain connected and tools can target a named device class, alias, or explicit device ID. `/bridge/devices` and `/bridge/select-active` expose current routing.
- **Delivery health report.** `python -m plugin.relay.realtime_agent.report` summarizes recent realtime-voice delivery modes and fallback reasons.
### Changed
- **Compatibility bootstrap covers only true gaps.** Current Hermes owns native session CRUD/messages and skill discovery; the optional hook now limits itself to legacy surfaces with no upstream replacement.
- **aiohttp 3.14.1 or newer.** Plugin/package requirements move to the patched dependency line covering the 2026 aiohttp security advisories.
- **Long gateway turns use liveness, not a short RPC cap.** Prompt submit can wait up to the server's long-turn ceiling while idle-progress watchdogs determine whether a turn has actually stalled.
### Fixed
- **Native `hermes plugins install` compatibility.** Runtime imports are package-relative, dashboard loading works under the upstream plugin namespace, and doctor exercises the real import chain.
- **Modern install layouts.** The installer detects classic, uv-managed, and containerized environments and points generated services/shims at the interpreter it actually found.
- **Doctor catches wrong dashboard surfaces and duplicate plugin copies.** Operators get an actionable correction instead of silently loading a stale directory or pointing Manage at a headless API server.
- **Resume ownership is generation-safe.** A stale candidate cannot detach an active phone route; confirmed replacements reject old failure/close/fatal callbacks, and failed opening candidates are never activated after their terminal callback.
- **Background results survive route loss.** Resumable sessions retain unacknowledged input and replay missed output, retry budgets start when a route is lost, and a detached durable run can still deliver by resume or notification.
- **One handoff and one ready event.** Duplicate spoken background acknowledgements and duplicate fresh-session ready telemetry are suppressed.
- **Credential files cannot be served as media.** Resolved paths under auth, token, pairing, SSH, relay-secret, and system-config locations are blocked even when general media delivery is permissive.
## Install / update
```bash
pip install hermes-relay==__VERSION__
# Native upstream plugin path:
hermes plugins install Codename-11/hermes-relay/plugin --enable
# Classic install / update on a systemd host:
curl -fsSL https://raw.githubusercontent.com/Codename-11/hermes-relay/main/install.sh | bash
# or, if already installed:
hermes-relay-update
```
## Verify
```bash
python -m relay_server --help
hermes relay doctor
python scripts/check-plugin-version-sync.py --expect __VERSION__
```
---
Tag prefixes: Android releases use `android-v*`, CLI releases use `cli-v*`. Historical
relay/plugin releases used `relay-v*` tags.
Tag prefixes: Android releases use `android-v*`, plugin releases use `plugin-v*`, and CLI releases use `cli-v*`.
+1 -1
View File
@@ -311,7 +311,7 @@ docker build -t hermes-relay relay_server/ && docker run -d --network host --nam
ln -s "$PWD/plugin" ~/.hermes/plugins/hermes-relay
```
Then restart hermes and run `hermes pair` to verify. The 18 `android_*` and 9 `desktop_*` tools register regardless of hermes-agent version. See [docs/relay-server.md](docs/relay-server.md) for TLS, systemd, and full setup.
Then restart hermes and run `hermes pair` to verify. The 35 `android_*` and 25 `desktop_*` tools register regardless of hermes-agent version. See [docs/relay-server.md](docs/relay-server.md) for TLS, systemd, and full setup.
</details>
+22 -8
View File
@@ -403,12 +403,25 @@ the new app version and a higher `appVersionCode`.
- `RELEASE_NOTES.md` — body of the GitHub Release for this version
(rewritten each release; the workflow uses this as-is). This is the
operator-facing summary, not the CHANGELOG mirror. Keep the
**Download** section near the top — it should spell out which file
to grab by its `-sideload-release.apk` / `-googlePlay-release.aab`
suffix (every artifact is version-tagged as
**Download** section near the top, in the required format (#144):
1. A lead callout naming the **one file most people want** —
"Installing on your phone? Download
`hermes-relay-<version>-sideload-release.apk` and tap it"
(full feature set), with the Play Store link for the
conservative build.
2. One explicit line that the `.aab` is a Play Console upload
bundle and **cannot** be installed by tapping it on a phone.
3. The `SHA256SUMS.txt` verify line + sideload-guide link.
No download table, no parity/testing artifacts: releases attach
exactly **two** app artifacts — the sideload APK and the googlePlay
AAB — plus `SHA256SUMS.txt` covering exactly those two (the 2-asset
policy in `.github/workflows/release-android.yml`; the parity twins
stay reproducible from the tag via CI but are not attached).
Every artifact is version-tagged as
`hermes-relay-<version>-<flavor>-<buildType>` via `archivesName`
in `app/build.gradle.kts`) and link to the sideload guide.
The v0.3.0 body is a good template.
in `app/build.gradle.kts`. Never rename the sideload APK — the
in-app update checker matches assets by `.apk` + `sideload` in the
name, and user-docs verify steps cite the filename.
- `app/src/main/assets/whats_new.txt` — in-app "What's New" content
shown in the settings/about screen. Update with the version number
and a brief feature summary. Gets stale silently if forgotten
@@ -651,9 +664,10 @@ On every push of a tag matching `android-v*`, `.github/workflows/release-android
regression slice with explicit timeouts.
3. Decodes `HERMES_KEYSTORE_BASE64` into `$RUNNER_TEMP/release.keystore`
and exports `HERMES_KEYSTORE_PATH` (skipped if the secret is unset).
4. Builds both Android release artifacts:
`./gradlew bundleRelease assembleRelease`.
5. Generates `SHA256SUMS.txt` covering both.
4. Builds all four flavored release artifacts
(`./gradlew bundleRelease assembleRelease`); only the sideload APK and
googlePlay AAB are attached (see §Release assets).
5. Generates `SHA256SUMS.txt` covering the two attached files.
6. Creates a GitHub Release named `Hermes-Relay-Android v<version>` with `RELEASE_NOTES.md` as
the body. Attaches the APK, AAB, and `SHA256SUMS.txt`. Tags any version
containing a dash (e.g. `android-v0.2.0-beta.1`) as a prerelease automatically.
+36 -18
View File
@@ -1,22 +1,18 @@
# Hermes-Relay-Android v1.2.6
# Hermes-Relay-Android v1.4.0
**Release Date:** June 27, 2026
**Since v1.2.5:** A fix for chats stuck showing "Untitled", and a calmer way to surface connection status. Chats now keep your first message as a stand-in title until the server names them (and titles reconcile after a turn / via a new refresh button in the session drawer), renaming sticks on non-default agent profiles, and the connection-status card no longer floats over your chat — transient states slide the screen down as a thin top banner, with the floating alert reserved for persistent errors.
**Release Date:** July 9, 2026
v1.2.6 is recommended for everyone.
**Since v1.3.0:** Realtime voice can keep a long task moving while you ask a quick follow-up, queue another long request, and deliver the finished answer in the selected realtime voice. Recovery is substantially stronger across backgrounding and route changes, model choices apply to the next session, and stale listening, thinking, reconnecting, and cancellation states no longer strand the voice screen. This release also adds model-catalog refresh, proactive notification rules, multi-device Bridge targeting, session-cleanup plumbing, and broad chat, startup, and security fixes.
v1.4.0 is recommended for everyone. Realtime Agent remains experimental and pairs with relay plugin v1.4.0; the no-plugin Standard chat and Vanilla Hermes voice paths remain upstream-compatible.
---
## Download
v1.2.6 ships in two Android build flavors. APK and AAB filenames are version-tagged:
**Installing on your phone?** Download **`hermes-relay-1.4.0-sideload-release.apk`** and tap it — that's the direct-install build with the full feature set (installs as `com.axiomlabs.hermesrelay.sideload`). Prefer the conservative build (no Device Control surface)? Get it from [Google Play](https://play.google.com/store/apps/details?id=com.axiomlabs.hermesrelay).
| Flavor | File | Who it's for |
|---|---|---|
| Google Play | `hermes-relay-1.2.6-googlePlay-release.aab` | Upload this Android App Bundle to Play Console. It has no AccessibilityService, screen reading, screenshots, gestures, SMS/calls, contacts/location, overlays, or unattended phone control. |
| sideload | `hermes-relay-1.2.6-sideload-release.apk` | Direct-install APK for full Device Control. Installs as `com.axiomlabs.hermesrelay.sideload`. |
| googlePlay APK | `hermes-relay-1.2.6-googlePlay-release.apk` | Parity/testing artifact. |
| sideload AAB | `hermes-relay-1.2.6-sideload-release.aab` | Parity/testing artifact. |
The other file, `hermes-relay-1.4.0-googlePlay-release.aab`, is an Android App Bundle for uploading to Play Console — it **cannot** be installed by tapping it on a phone.
Verify integrity with `SHA256SUMS.txt` from the same release. See the [Sideload guide](https://codename-11.github.io/hermes-relay/guide/getting-started.html#sideload-apk) for APK install steps.
@@ -24,15 +20,37 @@ Verify integrity with `SHA256SUMS.txt` from the same release. See the [Sideload
## Highlights
### Fixed
- **Chats stuck showing "Untitled".** The session drawer treated the server's session list as fully authoritative for the title, so a re-list that arrived before (or without) the server auto-naming a chat overwrote the optimistic first-message preview with a blank title. The drawer now keeps a known local title when the server returns a blank one, re-pulls shortly after a turn settles, and offers a manual refresh button — so chats stop reading "Untitled". The api_server SSE path never auto-titles, which is why the preview is now the durable fallback there. (#133)
- **Rename on a non-default agent profile.** A non-default profile's chats live in that profile's own store, but rename went through the shared path — so the new title never landed. Renaming is now profile-scoped (the write twin of the earlier session-delete and list fixes).
### Realtime voice that finishes the job
### Changed
- **Calmer connection status.** Transient/active/warning connection status — reconnecting, checking, LAN↔Tailscale handoffs — now renders as a thin banner at the top that takes its own space (content slides down) instead of a card floating over the chat. A persistent **error** keeps the floating alert so it still demands attention. Frequent confirmations (copied, profiles updated, profile/personality switches) moved to the same top banner instead of a bottom pop-up.
- **Keep talking while work runs.** Quick follow-ups can be answered while one Hermes task runs in the background, and another long request can wait in a bounded queue instead of being discarded.
- **Hear the authoritative answer.** Exact xAI delivery uses provider-native forced speech, finished-task answers can be replayed from the task chip, and TTS/text/notification fallbacks keep a result from disappearing when the realtime floor is unavailable.
- **Stronger route recovery.** Recorded turns wait for relay-confirmed resume, unacknowledged audio is replayed without starting a second Hermes run, and long-lived sessions get a fresh bounded retry window when the route actually drops. Retired sockets and sessions cannot overwrite a newer connection.
- **Clean lifecycle state.** Provider transcripts no longer impersonate active microphone capture; Stop settles local placeholders; exit detaches durable work while clearing session-owned UI; rejected, unacknowledged, or terminal cancels cannot leave an undismissable reconnecting task chip.
- **Your model and voice selection sticks.** Realtime Agent model and voice choices are scoped to the active connection/profile, survive restart, and apply when the next session opens.
### Chat and model management
- **Long turns stay alive.** Gateway submits use the server's long-turn window and idle-progress checks, avoiding premature transport fallback and duplicate turns.
- **Phone context reaches Hermes.** Voice-intent traces, card actions, and supported attachments now use payload channels the upstream server actually consumes; unsupported attachment paths report the gap instead of dropping it silently.
- **Refresh model catalogs on demand.** Chat and Manage can explicitly reload dynamic/custom provider models, while Manage keeps unconfigured providers visible with key-setup guidance.
- **Session cleanup groundwork.** The dashboard client supports export, prune preview/apply, archive, restore, and archived-session filtering for the Manage surface.
### Phone automation
- **Notification triggers.** Opt-in rules can match app notifications and show a safe local "Ask Hermes?" prompt, with recent activity and a global pause switch.
- **Multi-device Bridge targeting.** Relay tools can select a paired phone, foldable, tablet, or explicit device ID instead of assuming one Android client.
### Reliability and security
- **Older Android crash safety.** Collection calls that require Android 15 were removed from lower-API paths, and encrypted-storage dependencies are pinned to the compatible line.
- **Bad server addresses fail safely.** Malformed relay, media, session, voice, and chat URLs surface a normal connection error instead of closing the app.
- **Credential paths stay private.** Relay media delivery resolves symlinks and blocks credential, token, pairing, SSH, and system-config locations.
- **Cleaner voice failures.** Duplicate error surfaces are gone, fallback speech animates the voice UI, routine provider idle expiry opens fresh on the next turn, and fresh sessions emit one ready event.
---
## Upgrade notes
- This is an app-side release on **both** flavors — no Device Control or server changes needed.
- `appVersionCode` is **20**.
- App-side release on **both** flavors. Realtime Agent background/recovery features require relay plugin **v1.4.0**; Standard chat and Vanilla Hermes voice continue to work against unmodified upstream Hermes.
- `appVersionCode` is **22**.
- Realtime Agent is still an experimental engine. Stable assistant speech remains available through **Hermes Chat + Voice Output**.
+535 -31
View File
@@ -6,6 +6,420 @@ For shipped work, see `DEVLOG.md`. For architectural decisions, see `docs/decisi
---
## Active — next up (2026-07-07)
Compaction-safe snapshot of where we are; details in the linked sections below.
- **RELEASE IN PROGRESS — cut android-v1.4.0 + plugin-v1.4.0 (owner direction 2026-07-09).** Android is **1.4.0 / versionCode 22**, plugin is **1.4.0**, public release notes and store copy are synchronized, focused realtime recovery tests and Android lint are green, both signed release flavors build, the plugin package builds, and the current sideload APK is installed. Extended on-device recovery stress testing and the force-stop persistence check are explicitly deferred rather than release blockers:
1. Push `dev`, wait for its release-facing CI, merge `dev` -> `main` with a merge commit, then tag the shared merge tip as `plugin-v1.4.0` and `android-v1.4.0`. Plugin is a **MINOR** (it carries #165 native-loader + installer-venv, #170 doctor guardrails, #171 multi-device bridge, #178 dedup guard — not the 1.3.1 patch originally queued).
2. Verify both GitHub releases, their checksums/artifacts, and the Android signing summary.
3. Discard the 1.3.0 Play Console draft, inspect the uploaded 1.4.0 production draft, and start rollout deliberately.
- **Voice bugs being worked now** — see "Voice — on-device findings" below for full detail:
1. Background/resume turn stuck on `Listening...` / `Still working...` — **fixed in code; current APK installed; extended live stress test deferred** (2026-07-09).
2. Tool-call status pills/ordering + stuck "Thinking" — fixed in code; final visual ordering re-check remains.
3. Tap/static click between sentences (`RealtimePcmPlayer` boundary) — still needs an on-device audio repro before fixing.
- **Deferred post-release voice validation.** Repeat long-idle prewarm → record → background/foreground → route-change recovery, terminal retry exhaustion, repeated reopen/exit, cancel-without-ack, and force-stop persistence on physical devices. Capture both Android and relay traces for any recurrence; further recovery hardening or UX refinement may be required from those results.
- **Screen-wake-lock — SHIPPED (2026-07-07).** See "Voice — on-device findings" below.
- **Owner / Mizu — GitHub triage.** Close #64 as superseded, plus the queued open-issue comment/close/label batch (see "Open-issue resolution batch" below).
- **Voice exact-mode signoff — PASSED (2026-07-09 e2e).** Both `grok-voice-latest` and the pinned `grok-voice-think-fast-1.0` deferred on model-generated exact delivery, so xAI exact mode now bypasses inference through provider-native `force_message`. The full on-device background path spoke the authoritative answer through xAI with no fallback, and a pure-recall follow-up repeated it from history without a second Hermes route/run. OpenAI's separate out-of-band delivery spike remains on its next-RC roadmap. Full detail + the background-tasks-as-chat UX asks + remaining audio/UI gaps are below.
---
## Voice background-tasks — live findings + UX vision (2026-07-09 e2e realtime test)
Live on-device e2e (relay through `8ebb21b`, app `1.4.0-sideload` build 22, provider
`xai_realtime`). The delivery-report tooling from `5ff78da` was confirmed working
against live data during this test.
### Findings
- **Background route loss could strand a recorded turn — FIXED IN CODE; EXTENDED LIVE STRESS TEST DEFERRED (2026-07-09).** Initial logs showed valid PCM accepted by the persistent turn channel after foregrounding, but the socket had failed during background route retries and no new relay event arrived. The first fixed APK restored submission and let the background Hermes run finish, then exposed the delivery race: a slower overlapping resume handshake connected 250 ms after the valid resume, claimed relay ownership before the Android generation check, and detached the session just before the forced answer. The second installed reproduction completed the run and delivered its notification fallback, but voice stayed on `Waiting for route`: the periodic retry deadline had been created when the session was prewarmed, so its coroutine had already expired after five minutes of healthy uptime. Exiting voice mode also retained the session-owned `RECONNECTING` run; reopening rendered that orphaned pill and its close action targeted the new session instead of the detached task. Android now coalesces pending handshakes, waits for a relay-confirmed resumed socket, retains unacknowledged chunks for atomic replay, and returns a per-turn delivery result. Its retry worker lives for the session, starts a fresh bounded budget only when a route is lost, and clears that budget only after `voice.session.resumed`; bare WebSocket opens cannot reset it. Voice sessions carry a generation fence so late handoff, run, playback, and completion callbacks cannot repopulate or act on a newer session. Voice exit atomically drops detached handoff/run/confirmation UI before another session can prewarm; offline cancel rejection dismisses immediately, and a queued cancel without acknowledgement dismisses after a bounded wait. The relay requires a valid resume claim before changing ownership and isolates invalid/stale candidate failures from the active phone socket. Provider STT stays in `Transcribing` unless `VoiceRecorder.isRecording()` is true; Stop/failure settles local placeholders, late terminal deltas are ignored, and stale capture state is reconciled on resume. Route, promotion, ownership, UI-state, and chat-terminal regressions are green. The current APK is installed; repeated long-idle, background/foreground, route-churn, and terminal-exhaustion coverage remains a post-release follow-up and may drive further hardening.
- **Model-generated exact delivery is inconsistent and deferral is model-agnostic; xAI now has a deterministic path.**
A background turn ("what do you
think about our notes so far?") delivered `forced_summary_streaming`
(provider-voiced, early-commit) — grok read the answer in its own voice and
passed validation. BUT the same session's earlier turn ("check Hermes for what
we know about Minnesota") fell back to relay TTS (`acknowledgement_not_summary`).
A later forced-summary round on `grok-voice-think-fast-1.0` also spoke a genuine
deferral ("one moment ... I'll let you know") rather than the completed answer.
Validator fallback is therefore correct; model choice alone does not solve the
delivery-voice problem. xAI's provider-native `force_message` now handles
non-structured Exact deliveries without model inference. Its raw live event
stream and the full Android background path are verified; a recall follow-up
also answered from that provider history without re-running Hermes.
- **think-fast selection bug fixed + live-verified.** The app's session POST omitted
model/voice, so the settings dropdown was only a transient server-config editor
until **Save realtime agent** was tapped. Model/voice now persist per
connection/profile and ride every new session. On-device verification selected
think-fast without Save, saw the relay request it and the provider's final
resolution report it, then force-stop/relaunch restored the selection.
- **Duplicate "background task is running" — FIXED + LIVE-VERIFIED (2026-07-09).** The signoff trace captured both lines and disproved the suspected TTS mismatch: the provider first said it would check Hermes, then the broker requested a second provider response after promotion. Promotion now suppresses that second handoff when the original tool-calling response already emitted audio; silent calls still get one handoff. A deployed on-device round recorded `provider_acknowledged: true` and `spoken_handoff: false`, with only the original acknowledgement spoken. The same round confirmed the forced delivery emits one client response-start event after deduplicating xAI's `response.created` + `response.output_item.added` pair.
- **Status-speech logging gap — CLOSED / premise disproved (2026-07-09).** The raw signoff log contains both provider utterances as `voice.response.delta` text, plus the progress events; relay TTS did not speak either line. The flight recorder can reconstruct what the user heard. The real defect was redundant provider response generation, fixed above.
### Background-tasks-as-first-class-chat vision (owner ask 2026-07-09)
Theme: stop treating a background run as an ephemeral voice-only side effect —
surface it in chat like any other turn and keep its result. Overlaps the "Voice
background-run v2" chip roadmap below (items 3/4/7) but reframed around
chat/history rather than the voice chip; unify rather than build twice.
- **Titled background tasks.** Give each run a short title/label (first-line- or
model-derived) so it's identifiable in a list and in chat.
- **Chat entry on kickoff + result.** Drop a chat entry when a background task
starts ("Background task: <title> — running") and settle the result into the same
thread when it finishes. Don't leave it voice-only.
- **Results persisted in chat/history.** Show the background result cleanly in chat
history instead of discarding it after it's spoken — especially valuable for
follow-ups ("what did that say again?").
- **Detail view (expand on tap).** Tapping a background-task chat entry expands to
the full run detail like a normal chat message / tool timeline (reuse
`SubagentLane`). Same intent as v2 items 3+4 — build once.
- **Realtime agent retains background-result context in-session — FALLBACK PATH
DONE + SEEDING LIVE-VERIFIED (2026-07-09); NO-RERUN VERIFY PENDING.** On a FALLBACK delivery the broker now
seeds the delivered answer into the provider's history as an assistant turn
(`append_context_item` → silent `conversation.item.create`, no `response.create`),
so a follow-up ("what did that say?", "expand on that") finds it durably — fixing
the live "can't you see we ran the task?" failure; live follow-up confirmed the
provider knew the delivered context. Provider-VOICED success already
had its own turn in history, so it's untouched (no double-record). **Remaining:**
(a) live on-device verify that a pure-recall post-fallback follow-up is answered
without a re-run after the `92f9683` instruction fix; (b) the detached/promoted delivery (`_deliver_pending_background_result`)
and the DONE-chip respeak weren't in scope — confirm whether they leave the same
gap; (c) decide if the one-shot `native_pending_delivery_note` is now redundant
with durable seeding or still earns its keep as an explicit correction.
- **Proper concurrent multi-task.** True N-way parallel background runs — see v2
item 7 (deferred: needs session-per-run topology, run-id-targeted cancel,
multi-run chip/list). Owner is now explicitly asking for it; re-rank against the
queue rather than leaving deferred.
---
## Voice background-run A–E enhancement batch — SHIPPED in code (2026-07-08 PM)
Owner-approved full batch from the gap review; relay 93/93 realtime tests
green. Needs relay deploy + APK install + live verify.
- **A1 — positive summary validation + early-flush streaming.** The forced
summary must content-overlap the Hermes answer (`_summary_overlaps_answer`;
vacuous for bare confirmations) — blocklists chase phrasings, overlap
doesn't. And the summary response now STREAMS: buffered only until the
prefix (≥40 chars) clears the blocklist + shows answer overlap
(`_maybe_commit_forced_summary_early`), then flushes and streams live —
kills the observed "silence, then the whole answer in one burst" delay.
Uncommitted responses still get full end-of-response validation.
- **A2 — delivered-or-alarm.** `_confirm_background_delivery`: within 30s of
injection the summary must be done or committed-streaming, else
`delivery_unconfirmed` is logged and the answer is force-emitted as text.
A background answer can no longer be silently lost.
- **A3 — respeak.** `hermes.result.respeak` client message → relay respeaks
`last_background_result` via relay TTS. Client: tapping the settled (DONE)
chip requests it; chip stays up while it plays.
- **B — task queue (+N queued).** A long second ask is queued (FIFO, cap 3)
instead of refused (`status: "queued"`); starts automatically when the
current run's delivery settles (`_start_next_queued_run`, waits for the
summary, runs as durable, spoken transition via `_queued_start_prompt`).
Cancel clears the queue. `hermes.run.queued` event + `queued_count` on
promoted/background_completed/get_status; chip shows "+N queued". Queue
full → the old busy answer.
- **C1 — chip in compact mode.** The chip previously rendered ONLY in the
focus layout; compact mode now shows it above the bottom controls
(`bottom = 120.dp` — eyeball on device).
- **C2 — exit breadcrumb.** Exiting voice mode with a live background run
posts a chat system notice ("Background voice task still running (+N
queued) — Hermes will report back") via `VoiceViewModel.chatNoticeSink`
(wired in RelayApp to the shared ChatHandler).
- **D — `_thinking` drafting signal + answer redundancy.** Relay: the
drafted `_thinking` text is the answer of last resort when the
response-delta path yields empty (`answer_from_thinking` log). Client:
`_thinking` deltas drive a "Drafting the answer…" chip status line.
- **E — hygiene.** Fast lane reuses ONE side-session per voice session
(`fast_lane_session_id`); the idle probe now injects the relay xAI OAuth
token (`_probe_provider_options`) so it actually runs on the relay host;
new e2e test where the provider answers the summary request with filler →
fallback must carry the real answer
(`test_filler_summary_triggers_fallback_delivery`).
- **Live verify list:** summary starts speaking promptly (streaming, no
burst); filler → fallback speaks the answer; queue: two long asks →
"queued" spoken + "+1 queued" on chip → auto-starts with spoken
transition; DONE-chip tap respeaks; compact-mode chip visible; exit
leaves the chat breadcrumb; probe run completes (repro + keepalive).
- **VERIFIED LIVE (rounds 3–4, 2026-07-08 PM):** queue flow end-to-end
(queued ack → auto-start → both answers), chip +1-queued/finished states,
fallback delivery + audibility (user's own follow-up confirmed), and two
new gaps found + fixed same-day (see DEVLOG: whole-word/2-hit validation,
next-turn delivery note).
- **KEEPALIVE FINAL VERDICT — no protocol message resets xAI's 900s timer
(empirical 2026-07-08, 4 probe runs).** Repro died at 900.0s; silent-PCM
pings died at 900.0s; server-ACKNOWLEDGED `session.update` pings
(240/480/720s) died at 900.0s. The timer counts only real conversation
items. **SHIPPED IN CODE (2026-07-08 PM):** picked design (b): treat
idle-close as routine, close the Android websocket cleanly while idle,
and let the next user turn open a fresh provider conversation seeded
from the synced Hermes session. `_provider_keepalive_loop` is retired.
**Remaining:** relay deploy + live >15 min idle probe to verify silent
next-turn recovery on device.
- **Delivery input-quiet gate — SHIPPED (2026-07-08 PM, round-5 finding).**
A background task finishing while the user was mid-utterance delivered
over them and ended their recording. The relay now knows the user is
speaking (live `input_audio.append` chunks stamp
`native_last_input_audio_at`) and `_await_floor_idle_for_result` holds
delivery until they've been quiet ≥1.5s (bounded by the existing floor
timeout). Covers summary/fallback/queued-transition. **Client half shipped:**
`VoiceViewModel` suppresses realtime response/audio/done only while
`VoiceRecorder.isRecording()` is actually true. Provider STT uses
`Transcribing`, not the capture-owned `Listening` state, so a partial
transcript cannot wedge the mic controls or suppress its own response.
- **Audio tail cut at end of response (round-5 repro) — MITIGATED IN CODE.**
Final word ("you?") cut hard instead of finishing smoothly. The client
output resume tail guard is raised from 350ms to 650ms so the final
buffered PCM has more time to drain before capture resumes. **Remaining:**
verify on device; if the final syllable still snaps, inspect
`RealtimePcmPlayer` drain/fade-out behavior.
- **Fallback speech says file paths (round-5 polish) — FIXED IN CODE.**
The fallback spoke "Source: 1. Personal/Household/Househol…"; TTS-safe
answer extraction now strips `Source:` / `Sources:` / citation lines and
source-list path lines before relay TTS.
- **grok-voice fails the delivery instruction ~always (4/4 live rounds) —
DEFAULT CHANGED, then REWORKED same-day.** Every observed forced summary
was deferral filler; the validator+fallback carried every delivery.
`speak_verbatim` was first made a direct relay-TTS default, then reworked
to provider-voiced exact delivery (below) to keep voice continuity.
- **Provider-voiced exact delivery — xAI direct path live-verified.**
Model-generated word-for-word instructions were not reliable. Non-structured
`speak_verbatim` now supplies the authoritative answer to xAI's `force_message`,
which synthesizes it in the selected realtime voice without inference and
records a normal assistant turn. A raw live probe confirmed the full transcript,
audio, history, and completion lifecycle. Structured results and summary modes
remain model-generated; relay TTS remains the validator fallback. The on-device
background path produced a clean `forced_summary_streaming` event and recall
reused the resulting provider history without another Hermes run. Post-audit hardening remains:
provider-death TTS fallback on all three delivery paths, confirm alarm on all
three, barge-in preemption-as-text, blocklist answer-exemption, and
structured-answer prompt routing.
- **Audit leftovers (deliberate, small).** (1) DONE-chip respeak always
renders via relay TTS — intentional determinism, but it voice-mismatches
the exact mode's promise; candidate: provider-voiced respeak with TTS
fallback. (2) Exact-mode answers >1400 chars are truncated with an
appended "…" (and machine-looking text gets "…" even under the cap) —
silent for a mode promising completeness; consider a visual "full answer
in chat" cue on truncation.
## Voice observability (2026-07-08 assessment) — pre-RC hardening
The realtime flight recorder (per-session JSONL under
`realtime-agent-runs/`, decision-point events with reasons, task-failure
wrappers, Android `DiagnosticsLog` Voice category) is in good shape — it
carried every live-round forensics session. Three gaps before the release
candidate:
- **Run-dir retention + wav tap gating — DONE (2026-07-08).**
`run_retention_days` (default 14, 0 disables) sweeps JSONL + wav
artifacts at session-log creation; the render wav is a debug-only tap
(`debug_audio_tap`, default off) deleted after PCM streams.
- **Delivery-outcome rollup — DONE (2026-07-08).**
`python -m plugin.relay.realtime_agent.report [--days N] [--json]`
tallies provider-spoken vs fallback deliveries with reasons; new
`forced_summary_delivered` marker makes clean deliveries countable.
- **Buffered flight-recorder writes (minor).** `_log` open/appends per
event on the event loop, including one line per audio chunk. Fine so
far; switch to a buffered writer if voice sessions ever stutter under
load — measure before optimizing.
## OpenAI realtime provider — next-RC roadmap (2026-07-08 research)
Full findings with sources in
`docs/plans/2026-07-08-openai-realtime-notes.md`. Headline: the OpenAI
provider already exists and is broker-wired
(`plugin/relay/realtime_agent/providers/openai.py`) but has never had a
live round and defaults to a superseded model. Key provider contrasts vs
xAI: hard 60-min wall-clock session cap (not an inactivity timer),
out-of-band responses (`conversation:"none"` + explicit `input`), async
function calls, per-token pricing (2.1 audio $32/$64 per 1M; mini $10/$20)
vs grok's flat $0.05/min.
- **Bump OpenAI realtime default to `gpt-realtime-2.1` — CODE DONE
(2026-07-08).** Default bumped, `2.1-mini` + rollback `2` in the model
options. Remaining: live connect on 2.1 (covered by the live-verify
item below).
- **Live-verify the OpenAI provider end-to-end.** Code-complete but no
recorded live round (all forensics are grok-voice). Run the xAI
on-device battery (pair → voice turn → `hermes_run_task` →
exact-delivery → queue → respeak) on 2.1. Success bar: a
`realtime-agent-runs/` log shows a clean OpenAI session reproducing the
flows with provider-voiced Hermes delivery.
- **Handle OpenAI's 60-min hard cap.** Distinct failure mode from xAI's
900s inactivity close — it can cut an ACTIVE session. First confirm how
a cap-close currently surfaces (idle-close handling is xAI-shaped, e.g.
`_PROVIDER_IDLE_CLOSE_WS_REASON`), then add wall-clock-aware proactive
reconnect/reseed. Success bar: a >60-min OpenAI session survives the
cap with a proactive reseed, no user-visible break.
- **Spike out-of-band exact delivery on OpenAI
(`conversation:"none"` + answer as `input`).** Supply the Hermes answer
as explicit input context instead of an instructions injection the
model may ignore. Success bar: measurably lower deferral/filler rate
than grok forced-summary in repeated live deliveries, demoting the
validator to a safety net.
- **Async function-call delivery on OpenAI.** OpenAI GA allows the
session to continue while a function call is pending — a promoted
`hermes_run_task` could complete with a real late
`function_call_output` instead of interim-ack + synthetic
instructions, retiring `native_pending_delivery_note`. Success bar:
provider history reads "done" (never "still running") after a promoted
run, verified live.
- **Guardrail test: only `hermes_*` tools advertised on OpenAI
realtime.** Assert `session.update` never advertises hosted-MCP or
non-Hermes tools. Success bar: test fails if any such tool appears.
- **(Defer/eval-only) provider `semantic_vad` vs relay-owned floor.**
Better turn-taking naturalness but moves barge-in ownership off
`RealtimeFloor` — re-architecture, not RC scope.
## xAI voice platform moved (2026-07) — re-baseline items
xAI shipped `grok-voice-think-fast-1.0` (reasoning voice model, built for
tool-calling precision) as the new flagship; `grok-voice-fast-1.0` is
deprecated and the `grok-voice-latest` ALIAS NOW RESOLVES TO THINK-FAST.
We default to the alias everywhere (`config.py:106`,
`providers/xai.py:31`), so the live model may have changed under us —
xAI's docs explicitly say to pin versioned models in production. July also
added 21 multilingual voices, speech tags, voice cloning, session
resumption (30-min inactivity history retention), and a
`turn_detection.idle_timeout_ms` re-engagement knob.
- **Decide pin-vs-alias, then re-baseline the live delivery rounds.** The
4/4 deferral-filler verdicts may predate the alias flip — a reasoning
voice model may comply with the exact-reading instruction where fast-1.0
didn't. Resolved-model logging is DONE (2026-07-08):
`provider_model_resolved` records the session.created echo, the delivery
report prefers it, and `grok-voice-think-fast-1.0` is a selectable pin.
Remaining: run the live rounds, read the resolved ids, and decide
pin-vs-alias for production. Success bar: we know which model each live
round actually ran on, and the default is a deliberate choice.
- **Re-probe session lifecycle on think-fast.** The 900s
conversation-inactivity close and the keepalive-negative verdict were
measured pre-think-fast; xAI now documents session resumption and
`idle_timeout_ms`. Re-run `scripts/realtime-provider-idle-probe.py`;
if resumption is real, the idle-close-and-reseed handling can become
reconnect-and-resume. Success bar: fresh empirical timeout/resume
verdicts recorded in the POC doc.
- **Surface the new voices + speech tags.** `provider_options.py` carries
a static grok voice list; refresh or fetch dynamically, and evaluate
speech tags against the enhanced-voice config contract. Success bar:
new voices selectable in Voice Settings against a live relay.
## Voice — on-device findings (2026-07-08 e2e realtime test)
Live e2e test (phone on 1.4.0 dev APK, relay at `789f32c`) surfaced a chained
failure — full forensics from the session event log
(`realtime-agent-20260708-122613`). **All five fixes below are in code
(2026-07-08 PM); need relay redeploy + app rebuild + a repeat of the same
test.**
- **Stuck "Thinking" pill (root of the chain) — FIXED.** The gateway streams
drafting text as a `_thinking` pseudo-tool (`hermes.tool.delta` only, never
`tool.completed`), and `ChatViewModel.applyRealtimeAgentEvent` created a
ToolCall pill from the first delta of ANY tool name → a pill that spins
"running" forever (chat + voice overlay transcript). Fix: `_`-prefixed tool
names are internal (upstream's own hidden-tool convention) — never become
pills; their text still feeds the detailed thinking trace. Defensive same
guard on `hermes.tool.started`.
- **Cancel on an already-finished run killed the delivered answer — FIXED
(relay).** `response.cancel` unconditionally flipped `hermes_run_status` to
"cancelled" and emitted `hermes.run.cancelled` even with no run in flight
(observed: user cancelled 10s after completion — invited by the stuck pill —
and the Tokyo answer was never spoken). Now the Hermes-run half of cancel
only fires when a run is actually active; speech-stop always happens.
- **Model read the 32-char run ID aloud — FIXED (relay).** The interim ack
and the forced-summary prompt both handed the model `run_id`
(payload/metadata). Removed everywhere model-visible (get_status/cancel
default to the active run; the client gets ids via events) + explicit
"never say run/session IDs aloud" in all three instruction sites.
- **Model claimed "I'll add that to the queue" — FIXED (relay, instruction).**
No queue exists (v2 item 2 not built). All handoff/busy instructions now
state "there is no task queue — do not offer to queue or claim to have
queued anything." True multi-task chip stacking remains the v2 queue item.
- **Delivery spoke deferral filler instead of the answer — FIXED (relay).**
The forced-summary validator caught run-id speech (that saved the Minnesota
answer via fallback) but not "One moment while I look that up. I'll report
back as soon as I have the info." — Tokyo's answer was lost behind that
filler. Added deferral phrases (one moment / report back / looking into /
i'll look / as soon as i have) to `_bad_forced_summary_reason`; summary
prompt reworded to "speak the answer NOW". Tests:
`plugin/tests/test_realtime_summary_validation.py` (5) + updated cancel
route test; realtime batch 69/69 green.
- **Stale pre-lead — FIXED (relay).** A new run's "I'll check Hermes"
progress event carried the PREVIOUS run's run_id + completed_tool_count
(fires before the per-run reset). Now sends null/zero identity when no run
is in flight; keeps the active run's identity during a fast-lane attempt.
- **Background-run chip vanished the instant the waveform came back — FIXED
(client, second finding same day).** The chip was nulled at the first
summary-audio byte ("the DELIVERING chip has done its job"), so it
disappeared exactly when speech started — reading as the task being lost.
New `BackgroundRunPhase.DONE`: on first summary audio (or the 20s
no-audio watchdog) the chip settles to "Background task finished." — solid
dot, frozen ticker — lingers 10s (`DONE_CHIP_LINGER_MS`), then
auto-dismisses; ✕ on a DONE chip is a local dismiss (never a cancel); a
new promoted run replaces a lingering DONE chip and cancels its timer;
progress/tool/reconnect handlers can't reanimate a settled chip. Verify:
chip visibly settles + lingers while the answer is being spoken, ✕ during
DONE doesn't emit a relay cancel.
## Voice — on-device findings (2026-07-07 realtime test)
Surfaced during a live realtime-voice test with a long, many-tool-call background run. (The duplicate-error-toast + no-dismiss issue from the same test shipped this session — see DEVLOG 2026-07-07.)
- **Tool-call status pills stuck / ordering wrong — FIXED, needs on-device re-verify (2026-07-07).** After the recent background-run-chip work (`8dc874c`/`9554c7c`), the owner found on-device that the "Thinking" indicator can get stuck and that the relative order of tool-call pills vs. the agent's reply doesn't cleanly track what actually happened. Root cause was narrower than first suspected — `VoiceUiState.responseText` is write-only for the realtime path (nothing renders it), so the actual stuck surface was the `BackgroundRunChip`: no `hermes.tool.completed`/`hermes.tool.failed` branch in `VoiceViewModel`'s event handler meant a finished tool's `statusLine` stayed pinned at `phase=RUNNING` until the next unrelated event overwrote it. Fixed (`VoiceViewModel.kt:2619`): clears the finished tool's status line, advances `completedToolCount`, leaves `DELIVERING` alone. The ordering half was `CompactTranscriptRow` (`VoiceModeOverlay.kt`) rendering reply text above the tool rows that produced it — reordered to tool-rows-first (chronological). The per-message `ToolCall` transcript rows were already correct (untouched). `:app:compileSideloadDebugKotlin` green. **Needs on-device re-verify** (long multi-tool background run: chip never shows a stale finished-tool name; reply reads below its tool calls, not above) before the release resumes.
- **Tap/static click between sentences (realtime PCM playback) — NEEDS on-device audio investigation.** Suspected discontinuity at TTS chunk/sentence boundaries in `RealtimePcmPlayer` (a buffer underrun between segments, or a pop when a new segment's `AudioTrack` write starts). Capture head-position / underrun logs during a multi-sentence reply to confirm before touching the buffer sizing or adding a boundary crossfade/fade. Related to the existing "Realtime-PCM waveform output gating" note.
- **Screen-wake-lock for chat/voice — SHIPPED (2026-07-07).** The app previously relied entirely on the OS screen-timeout during both chat and voice mode. Added `KeepScreenOnWhile(enabled)` (`ui/components/OrientationOverride.kt`, `Window.FLAG_KEEP_SCREEN_ON` via `DisposableEffect` — the same Android-recommended visible-surface mechanism `power/WakeLockManager.kt`'s doc comment already pointed at for a background/no-window case), wired at the `ChatScreen` root as a single call site: `enabled = voiceUiState.voiceMode || isStreaming`. Rationale (matches other apps): voice mode is a call-like continuous session (Assistant/phone-call convention) so it holds the flag for the whole time the overlay is open, regardless of Idle/Listening/Thinking/Speaking sub-state; chat only holds it while a reply is actively streaming (video-playback convention) — idle reading/scrolling falls back to the OS default, matching WhatsApp/Telegram/Signal norms rather than pinning the screen on for a static transcript. Deliberately a single owner of the window flag (not ref-counted) — see the function's doc comment before adding a second caller. **Needs on-device confirmation**: screen stays on for the whole voice session incl. silent gaps, screen stays on only during active streaming in chat (not while idle), and the flag is correctly released on exiting voice mode / when a stream ends.
## Voice background-run v2 (2026-07-06 roadmap — post plugin-v1.3.0)
The v1 shape shipped in plugin-v1.3.0 (single durable run, free floor during
background work, busy answer, deliver-on-reattach, exit-detaches / chip-✕-
cancels). Ranked next increments, in value-per-complexity order:
1. **Fast lane — SHIPPED in code (2026-07-08; needs relay deploy + live voice
verify).** `_run_fast_lane_task` in `broker.py`: while a detached
(promoted/durable) run holds the background slot, a second
`hermes_run_task` first runs INLINE on a separate ephemeral Hermes session
(`session_id=None`) within the normal grace window; grace-elapse, a
known-long tool start (`_long_tool_hints`), explicit `mode=background`, or
promotion-off all abandon it and fall through to the (reworded) busy
answer. Touches NONE of the session's `hermes_*` run state — run_id/
status/progress/chip stay owned by the in-flight run — and emits no client
events of its own (bounded by grace; a chip would fight the detached
run's). Events: `voice.hermes_fast_lane.completed/abandoned/error` in the
session log. Tests: `plugin/tests/test_realtime_fast_lane.py` (7) +
updated `test_second_run_task_answers_busy_without_orphaning_first`
(per-stream cancellation tracking). **Residuals:** (a) context injection —
the ephemeral session gets only the task text + interface context, not
rolling conversation context (broker keeps no per-turn transcript; the
model is instructed to pass self-contained task text); (b) an abandoned
attempt may still finish server-side into the ephemeral session
(at-least-once, unread) — same property as promotion; (c) live verify:
during a long background run, ask a quick second question → answered
inline; ask a second long thing → busy answer unchanged.
2. **Task queue** — upgrade the busy answer from refusal to offer ("want me
to queue it?"): small FIFO in the broker session, start-next-on-completion
with a spoken handoff, chip shows "+1 queued". Pairs with (1).
3. **Chip tap-through to the transcript** — the run executes on a real
gateway session, so full tool calls/outputs already live in that session's
history; make the chip (or the finished turn) open it. Cheapest "see tool
output" step.
4. **Live tool-output sheet** — chip expands to a run timeline (tool name,
status, capped ~500-char output snippet). Relay adds a truncated output
field to `hermes.tool.*` events; client renders a lane (reuse the
`SubagentLane` pattern).
5. **Injection framing (recorded earlier, still open)** — on providers with
native async function calling, leave the tool call pending and deliver the
real `function_call_output` late instead of interim-ack + synthetic
instruction text. Needs a live xAI parity check first.
6. **Pending-result FIFO** — `pending_background_result` is a single slot
(correct for one run); generalize to an ordered list the day (1)/(2) land
so two results delivered during a detach don't race.
7. **Full N-way concurrent background runs — deliberately deferred.** Needs
session-per-run topology (a gateway session serializes turns), which
fragments conversation context, multiplies delivery/floor/failure modes,
and needs run-id-targeted cancel + a multi-run chip. Only worth it when
two *long* tasks genuinely need parallel wall-clock; revisit if the queue
feels slow in practice.
## Open-issue resolution batch (2026-07-06) — owner GitHub actions + deferrals
Plan: `docs/plans/2026-07-06-open-issue-resolution.md` (13 open issues triaged;
@@ -69,6 +483,37 @@ Deferred from the batch (coordination / decisions):
create a dedicated relay venv under a writable path so the full installer
works in-container.
Implementation-batch follow-ups (from the per-branch reviews):
- **#166 recovery: empty-session fail-fast.** `HermesApiClient.getMessages()`
maps fetch failures to `emptyList()`, so the recovery poller can't distinguish
"server unreachable" from "session genuinely empty" — a `Result`-returning
history read would let the never-landed-send fail-fast also cover a dropped
FIRST message of a fresh session (today that case polls to the cap).
- **#166 recovery cap.** Recovery gives up after 30 minutes; longer turns still
land in session history but only surface after a manual reload. Consider a
"keep waiting" affordance if real turns exceed the cap.
- **CI android slice.** `ServerAddressTest` + `IssueReportAndDiagnosticsTest`
added to the focused `--tests` slice; the Robolectric/MockWebServer recovery
tests and the compact-onboarding Roborazzi test stay local-only (same
precedent as `StoreScreenshotTest`) until the broad-suite hang (#32) is fixed.
- **Skills docs still cite editable-only fixes.** `skills/devops/hermes-relay-pair/SKILL.md`
and `skills/android/SKILL.md` document `python -m plugin.pair` + `pip install -e`
as the ModuleNotFoundError fix — add the native-layout equivalent when the
#165 branch ships.
- **Dashboard API tests not CI-visible.** `plugin/dashboard/test_plugin_api.py`
isn't discovered by `unittest discover -s plugin/tests` and needs
fastapi/httpx — wire into a CI runner or move under plugin/tests with skips.
- **Desktop tool-count drift.** `user-docs/desktop/index.md` counts client-side
handlers (clipboard/screenshot/open_in_editor) that have no server-side
`desktop_*` registration in `plugin/tools/desktop_tool.py` — reconcile the
advertised set; also `user-docs/desktop/pairing.md` wrongly says Android uses
`~/.hermes/remote-sessions.json` (it's Keystore/EncryptedSharedPrefs; the file
is shared with the Ink TUI). CLAUDE.md Key Files also still says 18/24 tools.
- **Info-report button label.** The diagnostics Report button reads "Report"
even when the first tap only reveals the expectation field — a "Continue"
label would make the two-step flow clearer.
## Connections UI / status banner (2026-06-30 restructure follow-ups)
The Connections screen was split into a scannable list + a tabbed detail screen
@@ -185,17 +630,83 @@ gated to pre-first-token. Deferred:
The deliver-on-reattach / adaptive-promotion / milestone-speech / resume-retry /
prewarm batch shipped (see DEVLOG 2026-07-01). Deferred:
- **Result injection framing (needs xAI parity check).** The completed background
summary is injected as a synthetic *user* message (`send_text` →
`conversation.item.create` role=user). Cleaner per current realtime-API practice:
inject as a function-call output / out-of-band response so the model can't mistake
it for the human speaking. OpenAI realtime supports this; xAI support unverified —
requires a live parity test before switching. Keep the user-message path as the
fallback.
- **Result injection framing — FIXED in code, deployed, needs live e2e voice verify (2026-07-07).** The completed background summary, the background-handoff acknowledgement, and the forced-Hermes preamble were all injected as a synthetic *user* message (`send_text` → `conversation.item.create` role=user) — the model saw a fake turn where "the user" said things like "Hermes has already handled the user's previous voice request..." Research turned up a cleaner mechanism than the one originally guessed at: `response.create` supports a per-response `instructions` field that overrides the session system prompt for one response only, **without creating any conversation item at all** — confirmed supported by both providers (OpenAI's own docs; xAI's Voice Agent API docs explicitly show the same `response.create.response.instructions` shape). `conversation: "none"` (true out-of-band, not in history) is OpenAI-only and was deliberately NOT used — we want the spoken summary to land in real conversation history so follow-ups like "what was that again" still work; only the injection *transport* changed, not where the turn ends up. Implementation: `RealtimeAgentConnection.request_response()` (`providers/base.py`) gained an optional `instructions: str | None` kwarg; both `providers/openai.py` and `providers/xai.py` implement it identically (`{"type": "response.create", "response": {"instructions": ...}}` only when instructions are given, else the original bare `response.create`); all 4 broker-authored injection call sites (`broker.py:1244, 2113, 2352, 2560`) switched from `send_text(prompt)` to `request_response(instructions=prompt)`. The one genuine passthrough site (`broker.py:699`, real client-supplied text) is untouched. `python -m unittest discover -s plugin/tests` — 1073/1074 green (the one failure is the pre-existing, already-documented `test_reads_hermes_xai_oauth_credential_pool` fixture gap, unrelated). **Deployed to the relay (2026-07-07) — still needs a real on-device voice session** confirming the model still speaks a natural summary when driven by `instructions` alone (no preceding fake user turn); watch for a background-task delivery in particular since that's the highest-traffic call site. **Confirmed live-verified (2026-07-08)** via the raw event log on the relay: a background run (~4min, terminal tool ×9-10) delivered its spoken summary correctly through the new `request_response(instructions=...)` path (`voice.response.started` → `voice.output_audio.delta` ×N → `voice.response.done`, clean).
- **xAI closes the realtime session after 900s of true silence — SETTLED (2026-07-08).** Live logs showed the provider closing after ~900s of zero conversation activity. Four probe runs proved no keepalive works: the repro, silent-PCM appends, and acknowledged `session.update` pings all died at exactly 900.0s. **Current code path:** idle-close is routine provider-session expiry; the broker closes Android cleanly with no `voice.error`, the old keepalive loop is gone, and the next user turn opens a fresh provider conversation seeded from the durable Hermes session. **Remaining:** relay deploy + on-device >15 min idle recovery verify.
- **Realtime voice: provider-answered turn durability — gateway drain + provenance badge SHIPPED (2026-07-08); app-restart persistence still open.** Shipped in code (needs on-device verify with the rest of the voice batch): (a) **gateway trace drain** — a gateway-configured turn with unsynced synthetic sync messages (voice intents / card dispatches / provider-answered realtime turns) now forces itself onto the sessions SSE route so the traces actually reach the server (previously "leave them for the next SSE turn" meant *never* on a gateway-primary phone). Deliberately narrow: only with an existing session id + the sessions fallback route (a stateless completions/runs detour would drop the turn itself from the transcript) and only on the default profile (a non-default profile's gateway session lives in its own state.db — the shared api_server POST would 404 and fail the user's turn; that residual defer case is accepted). The synced-mark guard now checks the route the turn actually *dispatched* on (`effectiveEndpoint`), also fixing a latent duplicate-resend for forced-SSE voice turns. (b) **provenance badge on reload** — `RealtimeTurnSyncBuilder.stripProvenanceMarker()` recognizes the synced `[Realtime Agent provider-native voice turn: …]` marker in loaded history, strips the bracket noise, restores the quiet "Realtime Agent" badge (same chip live turns get), and drops the superseded local clientOnly bubble so the exchange doesn't render twice. **Still open — app-restart loss:** unsynced traces are in-memory only; a restart before the next Hermes turn loses them. A fix needs a client-side pending-trace store (DataStore) plus answers to: which session should late traces sync into (voice binds per-session; the next turn may be a different session/profile), and restore-as-bubbles vs builder-side-only. A true flush-on-voice-exit is NOT implementable without an upstream append-messages API (every chat POST runs the agent); the drain above narrows the exposure window to "restart before the very next turn." Deliberately NOT a separate relay transcript store (forks the conversation).
- ~~**Realtime voice: subtle "Voice" provenance chip (2026-07-08).**~~ **Done via the durability item above** — turned out message-level "Realtime Agent"/"Voice" badges already rendered for live turns (`MessageBubble.kt` VolumeUp chips); the actual gap was reloaded history showing raw bracket provenance instead of the badge, now fixed by the marker → badge restore.
- **Pre-existing test failure:** `test_realtime_voice_routes.py::
test_reads_hermes_xai_oauth_credential_pool` fails at HEAD too (`token is None`) —
looks like an environment/fixture dependency on a local xai oauth pool, not a code
regression. Diagnose or gate on the fixture.
- **Standard voice `delegate_task(background=true)` nudge — SHIPPED then
REVERTED same-day (2026-07-08); premise disproven by the VERIFY-FIRST
check.** The nudge (a `STABLE_VOICE_INTERFACE_CONTEXT` line telling the
model to background long voice asks) was implemented, then the companion
verify-first item below was actually checked against upstream source and
killed it: **`delegate_task(background=true)` never dispatches async on the
api_server surface at all.** Upstream downgrades it to synchronous
execution (issue #10760): every api_server route binds
`async_delivery=False` (`gateway/platforms/api_server.py` ~4000), and
`tools/delegate_tool.py` (~2775) checks
`gateway.session_context.async_delivery_supported()` and runs the batch
inline with a "ran SYNCHRONOUSLY" note — "the adapter's send() is a no-op,
so a background dispatch would silently never re-enter the conversation."
Since ALL standard voice turns are forced onto SSE (ephemeral prompt slot),
the nudge would have made the model block just as long (plus subagent
overhead) while claiming it backgrounded. Reverted in `45c7ef4`. If a
"don't hold the voice floor" behavior is ever wanted on the standard path,
it needs the upstream async-delivery gap fixed first (a poll/webhook
delivery channel for stateless sessions — upstream contribution), or the
Relay realtime engine, which already has real background runs (ADR 33).
- ~~**Standard voice: speak a delegated result if the overlay is still open when
it lands.**~~ **CLOSED 2026-07-08 — premise gone.** There is no delayed
`delegate_task` completion turn on the standard voice path: the api_server
surface downgrades `background=true` to synchronous execution (see the
reverted-nudge entry above), so the "delegated result landing later" case
this wanted to speak cannot occur on SSE. On the gateway transport a
background completion does re-enter as a new turn — whether the phone's
gateway client renders an unsolicited idle-time turn is a separate
(text-chat) question, tracked nowhere yet; add it if gateway background
delegation becomes a used flow on phone text chat.
- **VERIFIED 2026-07-08 — a `delegate_task` completion turn can NEVER reach an
api_server-sourced session, because upstream never dispatches one there.**
Answered by reading current upstream source (clone @ `5057f03bf`): the
question is moot one layer earlier than expected. Every api_server route
binds the session context with `async_delivery=False`
(`gateway/platforms/api_server.py` ~4000, "the stateless HTTP path");
`tools/delegate_tool.py` (~2775) consults
`gateway.session_context.async_delivery_supported()` and, when false, runs
the whole batch SYNCHRONOUSLY with an explanatory note (issue #10760) —
there is no detached child, no completion event, no forged turn. The
`_async_delegation_watcher` → `_inject_watch_notification` →
`adapter.handle_message()` path only ever fires for sessions whose origin
routes to a real push-capable platform adapter (gateway chats, Discord,
etc.). Consequences applied same-day: the voice delegate nudge was reverted
and the speak-on-overlay item closed (entries above).
## Relay-enhanced standard voice for background tasks — research (2026-07-08)
**Verdict: NO — don't build it.** Full owner ask + Fable 5 agent research (cross-
checked against hermes-desktop's actual source, found in the local upstream
monorepo clone). Three lanes already cover "a long voice request survives and
reports back": (1) standard voice isn't a blocking call — a long turn just keeps
streaming, and the #166 SSE-recovery poller + `TurnCompleteNotifier` already
recover + notify on a dropped socket, zero relay involvement; (2) upstream's own
`delegate_task(background=true)` is the standard-path equivalent of the realtime
broker's `hermes_run_task` promotion — the model can detach a long task itself;
(3) hermes-desktop's own voice hook (`apps/desktop/src/app/chat/composer/hooks/
use-voice-conversation.ts` in the upstream monorepo — verified, zero mentions of
background/promotion) is the same thin synchronous record→transcribe→submit→speak
loop with NO background awareness; their background-task UX lives entirely in the
chat/composer surface (a status stack + native OS notification, never spoken) —
convergent with Android's existing background-run chip / `SubagentLane` /
`TurnCompleteNotifier`, not a gap to fill. Building a relay-side background layer
for standard voice would mean proxying an upstream-only surface through the relay
or monkey-patching deeper than the accepted `plugin/enhancements/` seam — against
the standard-path rule — to duplicate machinery ADR 33 itself calls the most
fragile code in `broker.py`, for an audience realtime already serves better.
Action items from this research are above (prompt nudge, speak-on-overlay-open
polish, the api_server-routing verify-first gate).
- **Prewarm cost watch.** Voice-mode entry now opens the provider session before the
first utterance. If users habitually open+close voice mode without speaking, idle
provider sessions cost connect/teardown churn — consider a short "no utterance in
@@ -229,7 +740,7 @@ Phase 1 (end-to-end spine) shipped on `Codename-11/phone-platform` — `send_mes
- End-to-end: with the app paired + "Let Hermes message me" on, run `send_message target=phone text=...` (and a cron `deliver=phone`) and confirm a notification on the device. Verify 503 (no phone) and the off-by-default gates.
- **Phase 2c reply round-trip — ✅ DONE (verified on-device 2026-06-29).** Confirmed: agent → phone notification → inline reply → drained through the relay's loopback `GET /phone/replies` (different process) → `handle_message` (`role_authorized=True`, no `PHONE_ALLOW_ALL_USERS`) → agent answer back in the *same* thread. Both fixes required (see DEVLOG / the Phase 2c bullet above).
- **FIX: cron `deliver=phone` / standalone send is broken.** Live testing: `hermes send --to phone` returns `{"error": "Unknown platform: phone"}`. The standalone (non-gateway) send path doesn't run a `kind=standalone` plugin's programmatic `ctx.register_platform`, so it never learns `phone` — only the running gateway (which loads `register()` at startup) does. The agent path (`send_message target=phone` in the gateway) works and was verified end-to-end on-device; the standalone/cron path needs the platform discoverable there too (declare it so the standalone loader picks it up, or route cron through the gateway). Until then `cron deliver=phone` won't work.
- **FIX: installer leaves stale plugin backup copies in the plugins dir (root cause of the 2026-06-29 round-trip failure).** `install.sh`'s plugin-clone rebuild backs the old copy up *inside* `~/.hermes/plugins/` (e.g. `hermes-relay.copy-backup-…`). Because the loader dedups discovered plugins by manifest `name` and both copies declare `name: hermes-relay`, the backup can win the dedup and the gateway loads stale code — so every later deploy is silently ignored. Fix: back up *outside* the plugins dir (or delete the old copy), and have `hermes relay doctor` warn when more than one directory under `~/.hermes/plugins/` resolves to the same plugin `name`.
- **FIX SHIPPED (2026-07-07) — installer + doctor guard against stale duplicate plugin copies; live-host verify pending.** Root cause of the 2026-06-29 round-trip failure: the gateway loader dedups discovered plugins by manifest `name`, so a second directory declaring `name: hermes-relay` (an old-installer backup copy, or a stray native install) could win the dedup and make the gateway load stale code — silently ignoring every later deploy. `plugin/doctor.py` now emits a `plugin-name-unique` warning when more than one directory under `~/.hermes/plugins/` declares the same plugin name (distinct real targets only — two links to the same target are deduped), and `install.sh` sweeps any such duplicate so only the canonical `hermes-relay` symlink survives. (Current `install.sh` already `rm -rf`s the old link rather than backing it up inside the plugins dir, so the original "back up outside the plugins dir" half is moot.) **Verify on the live host:** `hermes relay doctor` reports the `plugin-name-unique` check, and a reinstall leaves exactly one `hermes-relay` entry under `~/.hermes/plugins/`.
## Phone platform — usability roadmap (post device-verification, 2026-06-29)
@@ -283,9 +794,13 @@ The gateway-platform model is the *correct + sufficient architecture* (the phone
## Crash-class follow-ups
- **Audit remaining throwing URL-build sites for the "Invalid URL host" class (#131).** The #131 fix guarded the two clients that take a user-entered base URL on the Manage/voice path (`DashboardApiClient`, `StandardHermesVoiceClient`) and validates input at entry, but two lower-risk site groups still call okhttp's throwing `url(String)` / `.toHttpUrl()`:
- `HermesApiClient` streaming methods (`sendChatStream` / `sendCompletionsStream` / `sendRunStream`) build `authRequest("$baseUrl/…")` *outside* the surrounding `try`. Latent only — the non-streaming methods (incl. `checkHealth`) already `try/catch`, so a bad `apiServerUrl` is caught and marks the connection unreachable before streaming is reached. Consider a non-throwing `authRequestOrNull()` chokepoint → `onError`.
- Relay clients (`RelayHttpClient`, `RelayProfileInspectorClient`, `RelayVoiceClient`, `ConnectionManager`) use `.toHttpUrl()` on `$httpBase/…`. These ride post-pairing relay URLs (from a signed QR / pairing payload), not free-text fields, so the input-validation layer doesn't cover them — route them through `ServerAddress`/`toHttpUrlOrNull` for defense-in-depth.
- **Verify the Tink pin didn't break EncryptedSharedPreferences (owner, on-device).** The Android-15 `removeFirst`/`removeLast` crash lint flagged `com.google.crypto.tink.hybrid.HybridConfig.<clinit>` in the Tink dependency. Our app pulls Tink transitively via `androidx.security:security-crypto` for `SessionTokenStore`'s `EncryptedSharedPreferences`, which uses the AEAD path (not Hybrid), so the flagged `<clinit>` is very likely never reached at runtime — but we pinned `com.google.crypto.tink:tink-android:1.16.0` (ahead of security-crypto's transitive Tink) to clear the Play warning. **This is untestable without a build:** a too-new Tink can break `EncryptedSharedPreferences` at *runtime* (a `NoSuchMethodError`, not a compile error, so `./gradlew build` won't catch it). On-device smoke: launch the app, pair/sign in, force-stop + relaunch, and confirm the stored session survives (no re-pair prompt) and no startup crash. If it breaks, the blast radius is one line — revert the `tink-android` pin (catalog + `app/build.gradle.kts`) and the token store falls back to security-crypto's transitive Tink; then either try a lower Tink (1.15.0) or leave the (unreached) warning.
- **Bridge screenshots: regrant UX.** Multi-device live smoke found that a device can report `screen_capture_granted=false` because the MediaProjection grant was revoked and needs an in-app/user-consent regrant. The e-ink timeout path has been hardened with a longer configurable wait and one capture-pipeline rebuild retry; remaining polish is to surface the regrant action more prominently in Bridge status.
- **Audit remaining throwing URL-build sites for the "Invalid URL host" class (#131).** The #131 fix guarded the two clients that take a user-entered base URL on the Manage/voice path (`DashboardApiClient`, `StandardHermesVoiceClient`) and validates input at entry. Remaining site groups:
- **`HermesApiClient` streaming methods — DONE 2026-07-08.** `sendChatStream` / `sendCompletionsStream` / `sendRunStream` now build via the non-throwing `authRequestOrNull()` chokepoint (backed by top-level `buildApiRequestOrNull`, unit-tested like `buildRelayRequestOrNull`); a malformed base URL fails the turn through the normal `onError` channel ("Invalid server address …") and returns an inert EventSource instead of throwing out of the ViewModel. The whole #131 audit list is now closed.
- **`ConnectionManager` WSS connect — FIXED 2026-07-07** (this was the confirmed crasher: Play 1.2.6 on a Galaxy S25 Ultra / Android 16, `IllegalArgumentException` from `HttpUrl$Builder.parse` via `doConnectInternal` → `Request.Builder.url()` on the IO coroutine). Now routed through `buildRelayRequestOrNull()` → graceful Disconnected + diagnostic instead of a throw. `ConnectionManagerUrlGuardTest` covers it.
- **Remaining relay HTTP clients — DONE 2026-07-07 (defense-in-depth).** `RelayVoiceClient` now validates its base in `resolveHttpBase()` (returns null on a malformed URL → the existing `Result.failure` guards fire), and `RelayHttpClient`'s two string-URL sites (`fetchMedia`, `listSessions`) use `toHttpUrlOrNull()` → `Result.failure`. `RelayProfileInspectorClient` was already fully guarded (every `.toHttpUrl()` wrapped in `catch (IllegalArgumentException)`). The whole #131 relay class is now covered; `HermesApiClient` streaming (the other lower-risk group above) remains the only open item.
## Session titles (#133) — follow-ups beyond the client fixes
@@ -310,32 +825,21 @@ The client-side mitigations shipped (see DEVLOG 2026-06-27): the `updateSessions
## User-Added:
- [x] **Clean-chat: taller scrollable text viewport** *(impl 2026-06-22, orchestration batch — unbuilt; verify in Studio.)* Replaced the fragile `screenHeightDp*0.34f` cap with a weight split (sphere `weight(1f)` / flow `weight(1.1f)` ≈ 52% of the vertical slack); kept the internal scroll + top-fade + `min=96.dp` floor. `AgentTextFlow.kt` (`1dca285`).
- [ ] Verify profile selection retains voice config selections in all voice modes/configuration combinations - enhance UI/configurability/management for this.
- [x] **Session delete on a non-default profile now persists** *(impl 2026-06-22, orchestration batch — unbuilt; verify in Studio.)* Root cause: a non-default profile's sessions live in that profile's own `state.db`, but the delete went through the unscoped api_server `DELETE /api/sessions/{id}` (shared DB) so the row survived and the next profile-scoped list resurrected it. Fix routes gateway deletes through the dashboard profile-scoped surface (write twin of the list path) + `refreshSessions()` after success. `DashboardApiClient`/`ConnectionViewModel`/`ChatViewModel`/`RelayApp` (`6552566`).
- [x] **Voice-settings profile override in 'auto' mode** *(impl 2026-06-21, orchestration batch — unbuilt; verify in Studio. See DEVLOG + "Orchestration batch (2026-06-21)" below.)* Root cause: `VoiceViewModel.shouldPreferRealtimeVoice()` gated on `.route` (configured) not `.effectiveRoute` (resolved), so 'auto'+relay never engaged the override-capable relay path and fell back to host-global Standard `/api/audio/speak` (no override slot). Fixed + wired `connectionId` for per-profile voice-prefs namespacing. Original note: *Look into the voice-settings profile specific capabilities - in 'auto' mode the user-override voice wasn't applied (system default used) despite being displayed; only 'Relay' applied it.*
- [x] **Analytics + Diagnostics overhaul** *(impl 2026-06-22, orchestration batch — unbuilt; verify in Studio.)* Diagnostics is now a full-screen `DiagnosticsScreen` (new `Screen.Diagnostics` route, replacing the modal sheet) led by a vertical status-check timeline — Network, API server, capabilities, chat transport, pairing/auth, relay, voice — each a green/amber/red/gray dot on a connecting rail with an inline failure reason; checks backed by a logged error are tappable into `DiagnosticDetailDialog`. Derived read-only from existing `ConnectionViewModel` flows + recent `DiagnosticsLog` via a pure `buildStatusChecks()`; recent-activity log kept below. Analytics hierarchy tidied. `c3098a9`. See follow-ups below.
- [x] **Realtime voice stall + over-chatty status** *(client half impl 2026-06-21, orchestration batch — unbuilt; server half deferred, see below.)* Client now relaxes the 90s idle watchdog on promoted/long runs (5-min backstop kept) and throttles spoken status (≥22s gap, ≤3/turn); realtime waveform now gates on real playback-start. Original note: *Realtime voice mode stalls/times-out when calling a background Hermes task and repeatedly reports status vocally when not necessary.*
- [x] **Connections reframe: "Vanilla/Standard Hermes" → "Hermes"** *(impl 2026-06-22, orchestration batch — unbuilt; verify in Studio.)* 28 user-facing display strings across 10 connection/voice/permissions files; "Hermes-Relay plugin" → "Relay plugin" where it reads naturally. Display text only — no enum names, sealed types, when-branches, or stored route values touched. `c9fa8f7`.
- [x] **Lock app to a specific profile** *(impl 2026-06-21, orchestration batch — unbuilt; verify in Studio.)* Per-connection lock: new `ProfileLockStore`, `ProfileController` lock flows + enforcement, `ConnectionInfoSheet` collapses the picker to a static "Locked to <name>" row, `SettingsScreen` adds the lock card + dialog (the one surface still listing all profiles). Original note: *Allow locking app to a specific profile, hiding all other profiles except from this setting - cleanly hide profile specific UI elements based on this gate.*
- [x] **Profile icon in the floating voice overlay** *(impl 2026-06-21, orchestration batch — unbuilt.)* `VoiceModeOverlay` header pill now shows the per-profile icon (`LocalAgentIconPath`); sphere/pet stays the fallback.
- [x] **Voice dropdown state mixes + label overflow** *(impl 2026-06-21, orchestration batch — unbuilt.)* Invalid engine/route combos made unreachable (RealtimeAgent disabled without relay, unavailable routes disabled, `coerceAudioRoute` auto-corrects); long dropdown/provider labels get `maxLines=1`+ellipsis. Original note: *Fix the voice dropdown mode toggles to not allow weird state mixes - labels need overflow control to prevent 2 lines or crunching.*
### Thinking indicator — post-v1.3.0 follow-ups
- [x] **Per-profile agent icon + static-image avatar (shipped 2026-06-20 —** `d827e46`**, see DEVLOG).** Per-profile icon: client-side `ProfileIconStore` (per `(connection, profile)`, never sent to Hermes; stores a copied-file path) → small Coil image beside the agent name in `MessageBubble` via `LocalAgentIconPath`; picker is `AgentIconRow` under the local-name row in `ConnectionInfoSheet`. Static image: "Add a pet" accepts a single image (magic-byte detect → one-frame static pet). Scope shipped: small name-adjacent icon only; big avatar stays global. Follow-ups: on-device smoke (import an image as a pet; set a profile icon, confirm it shows by the name + persists across restart); optionally also show the icon in the profile picker.
The animated dot-matrix "thinking" indicator shipped in **android-v1.3.0** (Wave/Pulse/Bounce/Sparkle motions + Auto/accent colors, live preview in Chat settings; static when animations are off). Remaining:
- [ ] **Dot-matrix "thinking" indicator** *(prototype impl 2026-06-28 — unbuilt; verify in Studio.)* New `DotMatrixIndicator` (`ui/components/DotMatrixIndicator.kt`): a Compose-`Canvas` dot grid with a brightness wave sweeping left→right — the dot-anime-react concept reimplemented natively (not a port). Swaps the in-bubble `StreamingDots` working indicator via `LocalThinkingIndicator` (provided in `ChatScreen` around the message `LazyColumn`), behind a new **Chat settings → "Thinking indicator" (Dots / Matrix)** selector with a live preview (`thinkingIndicatorStyle` pref on `ConnectionViewModel`, default "matrix"). Brand-themed (uses the bubble `textColor`), frame-throttled via `rememberAmbientPhase` (not `rememberInfiniteTransition`), and renders a static frame when `animationEnabled` is off. Follow-ups once the base motion is approved:
- [x] **Preset frame patterns** *(impl 2026-06-28)* — `ThinkingMatrixPattern` (Wave/Pulse/Bounce/Sparkle): Wave stays procedural, the rest are authored `List<Set<Int>>` frame sequences (built generatively in `buildMatrixFrames`, addressed `row*cols+col`), crossfaded between frames. New `thinkingMatrixPattern` pref + a Matrix-only "Pattern" selector in Chat settings. Width widened twice on request (column pitch now 9dp).
- [x] **Per-indicator color** *(impl 2026-06-28)* — `ThinkingMatrixColor` (Auto + brand accents relay/cyan/green/amber/purple/pink) resolved against `LocalBrand` via `toColor()`, so accents re-theme per app theme. New `thinkingMatrixColor` pref + a Matrix-only swatch row in Chat settings; Auto follows the bubble text color. Possible later add-on: a freeform custom-color picker.
- **OS-level reduce-motion / TalkBack** — currently gates only on the app's `animationEnabled` pref. Also honor OS reduce-motion + touch-exploration like `CleanChatMode` does (`rememberCleanMotionState().osAnimations`).
- **Optional: promote to a full avatar style** — the alternative scope (a `DotMatrixAvatar` `AgentAvatar` shown everywhere via `LocalAvailableAvatars`, selected in Appearance). Deferred in favor of the narrower in-bubble indicator.
- **OS-level reduce-motion / TalkBack** — currently gates only on the app's `animationEnabled` pref. Also honor OS reduce-motion + touch-exploration like `CleanChatMode` does (`rememberCleanMotionState().osAnimations`).
- **Optional: promote to a full avatar style** — the alternative scope (a `DotMatrixAvatar` `AgentAvatar` shown everywhere via `LocalAvailableAvatars`, selected in Appearance). Deferred in favor of the narrower in-bubble indicator.
## Demo mode (2026-06-27) — deferred polish
Shipped offline Demo / Explore mode (see DEVLOG 2026-06-27). Core is in; these are non-blocking polish items, none required for the Play "App access" fix:
- **On-device verify (Studio).** Confirm: "Try the demo" on the onboarding Connect page and the standalone Connect screen lands on Chat showing the canned transcript (Markdown, tool-progress card, weather card, code block); the persistent banner shows and its Connect exits demo into the real wizard; demo runs in airplane mode with no network; Manage/Voice show the demo empty state; Bridge/Terminal show their pair-gate; backing out of demo Chat clears the flag so a real connection still works.
- **Demo composer is a silent no-op.** `ChatViewModel.sendMessage()` early-returns with no API client, so typing + Send in demo does nothing. Polish: intercept sends while `isDemoMode` to append a canned "This is a demo — connect your Hermes server to chat for real" assistant bubble (or disable the composer with a hint), so it doesn't read as broken.
- **Demo composer is a silent no-op — DONE 2026-07-08.** `sendMessage` now intercepts while `isDemoMode`: echoes the user bubble and appends `DemoContent.composerReply` ("offline demo, can't answer for real — tap Connect in the banner"), both clientOnly so demo-exit's `clearMessages()` wipes them. Wired via `setDemoModeWiring` (unconditional in RelayApp — the client-gated chat init never runs in demo, so ChatViewModel's own handler is null there). On-device check rides the existing demo verify item above.
- **Live voice mode in demo.** The voice-mode overlay (mic) launched from Chat isn't demo-gated — a tap would attempt a transcribe (fails gracefully, no crash). Add a demo notice / disable the mic in demo. (Voice settings screen already shows the demo empty state.)
- **Light typewriter/stream simulation.** The transcript is statically populated; an optional per-token reveal on first entry would better convey the "streaming" feel. Acceptable as static for v1.
- **Optional richer demo.** Could add a second tool type or an image attachment to the transcript to showcase more surfaces; kept minimal/one-file for now.
@@ -485,10 +989,10 @@ Things to look into:
- **Update discovery (shipped 2026-06-30 — CLI + dashboard + app).** `hermes relay update-check`, a dashboard "Plugin version" card, and an app **About → "Relay"** row all compare the installed plugin against the latest `plugin-v*` release and surface the right update command (`hermes plugins update hermes-relay` vs `hermes-relay-update`). The app polls the relay's `GET /relay/update-check` (`:8767`, bearer) on each `auth.ok`; the relay is the single source of truth (the app never hits GitHub). Possible polish (deferred): a more prominent dismissible "relay is behind" banner outside About (today it's capability-first + the About row), and showing the app's own version alongside the relay's in the same readout (the app-Version row already exists separately just above it).
- **Per-profile enablement (shipped 2026-06-30).** `hermes relay profiles list|enable [--all|NAME]` + `plugin/profiles.py` resolve the install-once/enable-per-profile papercut; docs now cover the pair-once/one-relay model. Possible follow-up: an `install.sh` / `hermes plugins install` prompt offering "enable for all existing profiles" so new installs don't need the manual `profiles enable --all`.
- `**hermes-relay-self-setup` SKILL.md as a precedent** — we just shipped a self-installing skill that an LLM can fetch from a raw GitHub URL and execute. Does this pattern generalize? Could it become a recommended way for any third-party Hermes project to ship setup automation?
- **Bootstrap injection** — `hermes_relay_bootstrap/` monkey-patches `aiohttp.web.Application` to inject endpoints into vanilla upstream. This is intentional but feels like a hack. Upstream PR #8556 (`feat/session-api`) will eventually let us delete it — verified 2026-04-15 that its scope covers the full bootstrap surface (sessions, memory, skills, config, available-models). Track that PR's status periodically.
- **Gateway slash-command preprocessor — upstream Stage 1 PR.** Sibling follow-up to #8556. Intercepts known gateway commands on `/v1/runs` + `/v1/chat/completions`, dispatches the stateless ones (`/help`, `/commands`) via `gateway_help_lines()`, returns a deterministic "use a channel with session state" notice for the stateful majority. Currently being prepared in `C:/Users/Bailey/Desktop/Open-Projects/hermes-agent-pr-prep/` on branch `feat/api-server-gateway-commands`; awaiting subagent's code + draft PR body before pushing. See `docs/upstream-contributions.md` §5.
- **Bootstrap injection** — `hermes_relay_bootstrap/` monkey-patches `aiohttp.web.Application` to inject endpoints into vanilla/partial upstream. This is intentional but feels like a hack. The original broad PR #8556 was **closed as superseded**; native upstream now covers sessions/chat/fork via [#33134](https://github.com/NousResearch/hermes-agent/pull/33134) and skill/toolset discovery via `/v1/skills` + `/v1/toolsets` (#33016). **Done (2026-07-08, HRUI-002):** the bootstrap's sessions CRUD/messages/fork handlers and the legacy `GET /api/skills` list were retired outright — no pre-#33134 fallback remains; old core builds degrade via the client capability probe. **Still gapped (bootstrap remains for these):** config, memory, legacy `/api/skills/{name}` detail + `PUT /api/skills/toggle` (501 stub), available-models, `/api/sessions/search`, and the slash-command middleware — each retires individually when a native replacement lands or the dependent UX is removed. Track upstream per surface.
- **Gateway slash-command preprocessor — upstream Stage 1 PR.** Sibling follow-up to the native session-control baseline (#33134). Intercepts known gateway commands on `/v1/runs` + `/v1/chat/completions`, dispatches the stateless ones (`/help`, `/commands`) via `gateway_help_lines()`, returns a deterministic "use a channel with session state" notice for the stateful majority. Currently being prepared in `C:/Users/Bailey/Desktop/Open-Projects/hermes-agent-pr-prep/` on branch `feat/api-server-gateway-commands`; awaiting subagent's code + draft PR body before pushing. See `docs/upstream-contributions.md` §5.
- **Gateway slash-command preprocessor — bootstrap middleware (Stage 1 equivalent).** Sibling shim in `hermes_relay_bootstrap/_command_middleware.py` that mirrors the upstream Stage 1 PR as an aiohttp middleware injected at bootstrap time. Ships the hallucination fix to vanilla-upstream installs before the upstream PR lands. Planned for v0.4.1, after the current bridge feature branch wraps. See `ROADMAP.md` v0.4.1 entry.
- **Stage 2 — stateful slash-command dispatch on `/api/sessions/{id}/chat/stream`.** Blocked on PR #8556 merging. Once session primitives ship upstream, add a preprocessor scoped to the session chat stream endpoint only, using `session_id` as the persistence handle. Separate upstream PR + matching bootstrap middleware. See `docs/upstream-contributions.md` §5 ("Stage 2").
- **Stage 2 — stateful slash-command dispatch on `/api/sessions/{id}/chat/stream`.** Unblocked now that session primitives shipped upstream (#33134 / `f7527b0`). Add a preprocessor scoped to the session chat stream endpoint only, using `session_id` as the persistence handle. Separate upstream PR + matching bootstrap middleware. See `docs/upstream-contributions.md` §5 ("Stage 2").
When the answer becomes clearer, this section becomes either an ADR in `docs/decisions.md` or a Plan under `Plans/`.
@@ -545,7 +1049,7 @@ Follow-ups:
## Voice overhaul (shipped 2026-06-18 — `docs/plans/2026-06-18-voice-overhaul.md`)
- **Per-profile voice on Standard (upstream PR).** Upstream `/api/profiles/*` has no voice field and `/api/audio/*` is host-global. Long-term: PR a voice section to the profile config + make `/api/audio/*` honor the active/`?profile=` profile. The relay path already carries per-profile voice; ship that first.
- **Wire connectionId for per-profile voice namespacing.** `VoicePreferencesRepository` is scope-aware (`base_connId_profile`), but `RelayApp` passes only the profile *name* to `onProfileChanged`, so `connectionId` is null and keys namespace by profile-only. Wire `setVoicePrefsConnection` to `ConnectionViewModel.activeConnectionId` (in `RelayApp`) so two connections with same-named profiles don't share voice settings.
- ~~**Wire connectionId for per-profile voice namespacing.**~~ **Already shipped — stale entry (verified 2026-07-08).** The wiring landed in `0aa1b38` (2026-06-21, the same batch this list belongs to): `RelayApp` has a `LaunchedEffect(activeConnectionId, selectedProfile?.name)` calling `voiceViewModel.setVoicePrefsConnection(activeConnectionId)` *before* `onProfileChanged(...)`, and `applyVoicePrefsScope` pushes `(connectionId, profile)` into `VoicePreferencesRepository.setActiveScope`. Two connections with same-named profiles namespace separately.
- **Realtime-PCM waveform output gating.** The basic-TTS output waveform is now Visualizer-accurate (gated on real playback amplitude), but the realtime path gates `outputAudioActive` on `audioSeen` (first decoded PCM bytes) in `VoiceViewModel.handleRealtimeVoiceEvent`, which can still lead audible output by the `RealtimePcmPlayer` start prebuffer. Gate realtime on actual playback-start (head moved) to match the basic-TTS path.
## Chat clean-mode + pets (shipped 2026-06-18 — `docs/plans/2026-06-18-chat-clean-mode-and-pets.md`)
+12 -3
View File
@@ -182,7 +182,13 @@ android {
// [POC] Roborazzi runs without its Gradle plugin (the plugin needs AGP's
// removed TestedExtension). Force record mode via the test-JVM system
// property the plugin would otherwise inject, so captureRoboImage writes.
unitTests.all { it.systemProperty("roborazzi.test.record", "true") }
// Heap: the Roborazzi store renders (1080×2160 native graphics) share a
// worker JVM with the Robolectric suites; Gradle's 512m default OOMs
// once both are in the same run.
unitTests.all {
it.systemProperty("roborazzi.test.record", "true")
it.maxHeapSize = "2g"
}
}
}
@@ -290,6 +296,9 @@ dependencies {
// Security
implementation(libs.security.crypto)
// Force a Tink newer than security-crypto's transitive one — older Tink's
// HybridConfig removeFirst()/removeLast() trips the Android-15 crash lint.
implementation(libs.tink.android)
// DataStore
implementation(libs.datastore.preferences)
@@ -316,8 +325,8 @@ dependencies {
// [POC] Roborazzi host-side screenshot rendering (src/test, Robolectric).
// Renders real composables on the JVM at an exact canvas — no device, no
// status bar, no clipping. See StoreScreenshotTest.
testImplementation("io.github.takahirom.roborazzi:roborazzi:1.43.1")
testImplementation("io.github.takahirom.roborazzi:roborazzi-compose:1.43.1")
testImplementation("io.github.takahirom.roborazzi:roborazzi:1.66.0")
testImplementation("io.github.takahirom.roborazzi:roborazzi-compose:1.66.0")
testImplementation(libs.compose.ui.test.junit4)
testImplementation(libs.compose.ui.test.manifest)
testImplementation("androidx.test.ext:junit:1.3.0")
@@ -1,4 +1,6 @@
v1.2.6 — Tidier chats & calmer status.
v1.4.0 — Realtime voice that finishes the job.
• Chats no longer get stuck on "Untitled" — your first message stands in as the title until the chat is named, plus a new refresh button in the session list. Renaming a chat now sticks on non-default profiles.
• Connection status now slides in as a thin banner at the top instead of a card floating over your chat; the floating alert is kept for persistent errors.
• Long voice tasks can queue, keep running while you ask quick follow-ups, and deliver answers in the selected realtime voice.
• Voice sessions recover more reliably after background or route changes and clear stale task states.
• Refresh model catalogs on demand; add opt-in notification rules and multi-device Bridge targeting.
• Safer startup, server-address handling, long chat turns, and credential media access.
+305 -230
View File
@@ -1,238 +1,313 @@
{
"versions": [
"versions": [
{
"version": "1.4.0",
"title": "Realtime voice that finishes the job",
"date": "2026-07-09",
"sections": [
{
"version": "1.2.6",
"title": "Tidier chats & calmer status",
"date": "2026-06-27",
"sections": [
{
"header": "Tidier chats",
"bullets": [
"Chats no longer get stuck showing \"Untitled\" — your first message stands in as the title until the chat is named, titles refresh once a turn settles, and a new refresh button in the session drawer pulls the latest on demand. Renaming a chat now sticks when you're on a non-default agent profile."
]
},
{
"header": "Calmer status",
"bullets": [
"Connection status — reconnecting, checking, network handoffs — now shows as a thin banner at the top that gently slides the screen down, instead of a card floating over your chat; the floating alert is kept for persistent errors. Quick confirmations (copied, profiles updated, profile/personality switches) land in the same calm banner instead of a pop-up at the bottom."
]
}
]
"header": "Voice that keeps going",
"bullets": [
"Quick follow-ups can be answered while a long Hermes task runs, another long request can wait in a bounded queue, and the finished answer can stay in the selected realtime voice.",
"Voice route recovery now waits for relay confirmation, replays unacknowledged input without starting a second Hermes run, and rejects stale sockets or sessions before they can overwrite a healthy connection.",
"Listening, thinking, reconnecting, and cancellation states now settle cleanly after Stop, exit, route loss, or terminal retry failure."
]
},
{
"version": "1.2.5",
"title": "Stability + Try the demo",
"date": "2026-06-27",
"sections": [
{
"header": "Stability",
"bullets": [
"Fixed a crash that could close the app when a non-URL value — a UI label, or a line copied from the docs — was entered in the API server or Dashboard URL field. The setup fields now reject anything that isn't a valid host or http(s) URL with an inline error, and the dashboard and voice request paths treat a bad address as unreachable instead of crashing."
]
},
{
"header": "Try the demo",
"bullets": [
"A new \"Try the demo\" option on the setup screen — and on the empty chat screen if you skip setup — opens an offline preview of the real chat experience: a sample conversation with Markdown, a tool-progress card, and a rich card, with no server, account, or network. A banner shows it's a demo, with a one-tap Connect to set up for real."
]
}
]
"header": "Models and phone automation",
"bullets": [
"Realtime Agent model and voice choices apply to the next session, persist per connection/profile, and survive restart.",
"Chat and Manage can refresh dynamic provider model catalogs on demand.",
"Opt-in notification rules can offer a local Ask Hermes action, and Bridge tools can target a specific paired Android device."
]
},
{
"version": "1.2.4",
"title": "Stability + connection security",
"date": "2026-06-25",
"sections": [
{
"header": "Stability",
"bullets": [
"Fixed a crash that could close the app when the dashboard connection check hit a transient network failure — a pooled connection aborting or timing out over Tailscale. The check now reports the failure cleanly and the connection probe degrades gracefully instead of force-closing."
]
},
{
"header": "See if you're secure",
"bullets": [
"The chat status chip, connection card, and route picker now show at a glance whether your connection is encrypted — Encrypted · TLS, Encrypted · Tailscale (both secure), Mixed routes, or Not encrypted — and tapping it opens a per-transport breakdown (chat, API, relay tools). A Tailscale or WireGuard route is now correctly shown as encrypted rather than implied insecure."
]
}
]
},
{
"version": "1.2.3",
"title": "Connection crash fix",
"date": "2026-06-23",
"sections": [
{
"header": "Stability",
"bullets": [
"Fixed a crash that could close the app right after connecting over an encrypted link (Tailscale or HTTPS) — a live secure connection was being torn down on the main thread as it came up. Securing your connection no longer force-closes the app; plain-LAN connections were never affected."
]
}
]
},
{
"version": "1.2.2",
"title": "Multi-profile polish",
"date": "2026-06-22",
"sections": [
{
"header": "Profiles that behave",
"bullets": [
"Deleting a session while a non-default agent profile is active now sticks — it no longer reappears after the list refreshes.",
"On a cold start with a non-default profile selected, the session drawer opens on that profile's chats directly instead of briefly showing the default profile's."
]
},
{
"header": "Clearer diagnostics",
"bullets": [
"Diagnostics is now a full screen led by a top-to-bottom list of subsystem health checks — network, API server, chat transport, pairing, relay, and voice — each with a pass / warning / fail state and the reason when something's wrong; tap a failing check for full detail. The recent-activity log stays below."
]
},
{
"header": "Small touches",
"bullets": [
"The default connection is now simply \"Hermes\" (and the optional power features are labelled \"Relay\"), across setup, the switcher, voice, and permissions.",
"Distraction-free chat mode gives its text a taller, scrollable area."
]
}
]
},
{
"version": "1.2.1",
"title": "Polish & control",
"date": "2026-06-21",
"sections": [
{
"header": "Yours to control",
"bullets": [
"Lock the app to a single agent profile (Settings → Profile lock) and hide the rest from the pickers."
]
},
{
"header": "Find your way back",
"bullets": [
"A new \"What's New\" entry in Settings shows current and past release notes any time — not just after an update."
]
},
{
"header": "When something breaks",
"bullets": [
"Diagnostics show clean error titles — tap any entry for a detail view with Copy, Share, and a one-tap GitHub issue.",
"A tasteful in-app banner tells you when a newer version is live (Play or sideload) — dismissable, and it never nags."
]
},
{
"header": "Voice fixes",
"bullets": [
"Stop now halts realtime speech instantly, hold-to-talk is steadier, the voice overlay is easier to read, and a chosen voice applies in Auto mode.",
"Realtime turns that reach back to Hermes no longer drop with a session error."
]
}
]
},
{
"version": "1.2.0",
"title": "Make it yours",
"date": "2026-06-20",
"sections": [
{
"header": "Personalize",
"bullets": [
"Eight app themes in Settings → Appearance — the Hermes Relay brand plus ports of the Nous Hermes looks (Teal, Nous Blue, Midnight, Ember, Mono, Cyberpunk, Rosé), with light/dark.",
"Swap the agent orb for an animated pet that reacts to what the agent is doing — add, preview, and tune pets right in the app, or generate one from sprite art with the AI authoring kit.",
"Reskin the sphere, and give each agent profile its own icon."
]
},
{
"header": "See what's happening",
"bullets": [
"The chat status strip names the actual streaming path (Gateway, Sessions, Completions, Runs), with a basic→best tier ladder in Chat Settings.",
"Tap the context meter for a \"What the agent sees\" sheet — the exact extra context prepended to your next turn.",
"Voice and Realtime turns are badged in the scrollback."
]
},
{
"header": "Privacy",
"bullets": [
"When paired to the relay, the agent can mark private media and the phone blurs it per your setting — sensitivity stays model-emitted."
]
},
{
"header": "Faster & more reliable",
"bullets": [
"Cold start is about 3× faster, and model/personality/approvals load honestly instead of showing a maybe-wrong value.",
"In-app crash reporting offers a one-tap, pre-filled bug report.",
"QR pairing no longer force-closes on unusual cameras (foldables); fixed crashes opening server images and PDFs; in-chat model picks now apply."
]
},
{
"header": "Voice & terminal",
"bullets": [
"Enhanced voice control for Gemini and xAI providers.",
"Leaner terminal with TUI-correct input and an isolated, tuned tmux."
]
}
]
},
{
"version": "1.1.0",
"title": "Release plumbing & polish",
"date": "2026-06-16",
"sections": [
{
"header": "New",
"bullets": [
"Automated Play Console upload when a release tag ships (a human still starts the rollout).",
"/relay slash commands — status, devices, and pair from any platform — plus a relay-status badge in the dashboard header.",
"The relay plugin prompts for its optional voice-provider keys on install, and a tools-only native install path."
]
},
{
"header": "Improved",
"bullets": [
"Settings overhaul: status pills are now exception-only, Power tools shows a single Plugin active/required/offline badge, and Connections moved to the top.",
"Release names and notes are now split per surface (Android, plugin, CLI)."
]
},
{
"header": "Fixed",
"bullets": [
"No more force-close on connect when the stored credential keyset was corrupt — it now heals in place.",
"The installer works on uv-managed Hermes hosts, and the dashboard relay panel buttons are readable again."
]
}
]
},
{
"version": "1.0.0",
"title": "Stable launch",
"date": "2026-06-14",
"sections": [
{
"header": "Gateway chat with live thinking",
"bullets": [
"Chat can ride the upstream dashboard gateway — the only vanilla-upstream path that streams reasoning live, so the Thinking block and sphere light up during generation. \"Auto\" prefers it and falls back to the SSE endpoints per turn.",
"Desktop parity: native image/PDF/file attachments, mid-turn steering, edit & resend, approval/clarify/sudo/secret cards, live subagent lanes, a context-window meter, server slash commands, and turn-complete notifications.",
"Warm-start and an opt-in Keep connected in background toggle so long-backgrounded conversations resume instantly."
]
},
{
"header": "Agents, Manage & media",
"bullets": [
"Switch agent profiles per conversation — model, SOUL, personality, and skills — with the selection bound to the session, never changing the server default for other clients.",
"Manage parity with the desktop dashboard: change models, manage provider keys, edit profiles and SOUL.md, and browse/install skills.",
"Open and save chat images and attachments — full-screen viewer with pinch-zoom, plus an Open/Share/Save menu."
]
},
{
"header": "Standard path is first-class",
"bullets": [
"Chat, Manage, and voice all work against an unmodified upstream Hermes agent; the relay plugin is now purely additive.",
"Seamless connection UX — LAN↔Tailscale handoffs and reconnects no longer reload the chat, and status shows as in-theme slide-down toasts.",
"Persistent Realtime Agent voice that keeps one session across turns, with long runs promoted to tracked background tasks."
]
}
]
"header": "Reliability and safety",
"bullets": [
"Long chat turns avoid premature transport fallback, and supported voice, card, and attachment context now reaches upstream Hermes through channels it consumes.",
"Malformed server addresses fail through normal connection errors, older Android versions avoid newer collection APIs, and relay media blocks credential and token paths.",
"Model management keeps unconfigured providers visible with key-setup guidance, and session cleanup gains export, prune preview/apply, archive, and restore plumbing."
]
}
]
]
},
{
"version": "1.3.0",
"title": "Voice that multitasks & sturdier chats",
"date": "2026-07-06",
"sections": [
{
"header": "Voice, hands-free",
"bullets": [
"Ask for something big and keep talking — long tasks hand off to the background with a live chip showing the current step, steps done, and a running timer, with a tap-to-cancel. The answer is spoken when it's ready, even after a brief disconnect — and if the voice session is gone, it arrives as a notification (the full answer is always in the chat).",
"Leaving voice mode (or tapping stop to interrupt speech) no longer cancels a running background task — the chip's ✕ is the one deliberate kill switch, and a delivered answer keeps its text instead of flipping to \"Cancelled.\"",
"Quieter and quicker: the agent speaks at milestones instead of narrating every step, clearly long tasks hand off to the background right away, and the first turn starts faster — the session warms up when you open voice mode."
]
},
{
"header": "Chats that keep their answers",
"bullets": [
"An answer is no longer lost when the connection drops mid-reply on a long turn (slow local models, delegating skills) — the app quietly re-checks the conversation and completes the turn when the server finishes, with the usual done-notification if you've switched away.",
"Markdown reads like chat: headings are proportionate instead of billboard-sized, lists and paragraphs share one size, links are clearly styled, and timestamps show once per message group."
]
},
{
"header": "Your agent can reach out",
"bullets": [
"Proactive messages: your Hermes agent can message your phone first (off by default, opt-in on both server and phone), and you can reply straight from the notification or the new Hermes inbox — the conversation continues like any other chat."
]
},
{
"header": "Make it yours",
"bullets": [
"Pick your app font — Inter (new default), Nunito, or your system font — applied instantly, everywhere.",
"The in-bubble working indicator can be a small animated dot-matrix (Wave, Pulse, Bounce, Sparkle) with a color of your choice.",
"Quick Controls at the top of Settings puts Persistent connection and Turn-complete alerts one tap away."
]
},
{
"header": "Setup & housekeeping",
"bullets": [
"Onboarding slides now scroll on small screens and large font sizes, so no setup guidance is cut off.",
"Reporting a diagnostic files the right kind of issue: informational entries ask what you expected and file as a question, and every report carries your actual connection mode.",
"Connections is a scannable list with a tabbed detail screen (Overview, Routes, Advanced, Security), and voice settings can now read and edit your server's voice engine (provider, voice, model) over the dashboard."
]
}
]
},
{
"version": "1.2.6",
"title": "Tidier chats & calmer status",
"date": "2026-06-27",
"sections": [
{
"header": "Tidier chats",
"bullets": [
"Chats no longer get stuck showing \"Untitled\" — your first message stands in as the title until the chat is named, titles refresh once a turn settles, and a new refresh button in the session drawer pulls the latest on demand. Renaming a chat now sticks when you're on a non-default agent profile."
]
},
{
"header": "Calmer status",
"bullets": [
"Connection status — reconnecting, checking, network handoffs — now shows as a thin banner at the top that gently slides the screen down, instead of a card floating over your chat; the floating alert is kept for persistent errors. Quick confirmations (copied, profiles updated, profile/personality switches) land in the same calm banner instead of a pop-up at the bottom."
]
}
]
},
{
"version": "1.2.5",
"title": "Stability + Try the demo",
"date": "2026-06-27",
"sections": [
{
"header": "Stability",
"bullets": [
"Fixed a crash that could close the app when a non-URL value — a UI label, or a line copied from the docs — was entered in the API server or Dashboard URL field. The setup fields now reject anything that isn't a valid host or http(s) URL with an inline error, and the dashboard and voice request paths treat a bad address as unreachable instead of crashing."
]
},
{
"header": "Try the demo",
"bullets": [
"A new \"Try the demo\" option on the setup screen — and on the empty chat screen if you skip setup — opens an offline preview of the real chat experience: a sample conversation with Markdown, a tool-progress card, and a rich card, with no server, account, or network. A banner shows it's a demo, with a one-tap Connect to set up for real."
]
}
]
},
{
"version": "1.2.4",
"title": "Stability + connection security",
"date": "2026-06-25",
"sections": [
{
"header": "Stability",
"bullets": [
"Fixed a crash that could close the app when the dashboard connection check hit a transient network failure — a pooled connection aborting or timing out over Tailscale. The check now reports the failure cleanly and the connection probe degrades gracefully instead of force-closing."
]
},
{
"header": "See if you're secure",
"bullets": [
"The chat status chip, connection card, and route picker now show at a glance whether your connection is encrypted — Encrypted · TLS, Encrypted · Tailscale (both secure), Mixed routes, or Not encrypted — and tapping it opens a per-transport breakdown (chat, API, relay tools). A Tailscale or WireGuard route is now correctly shown as encrypted rather than implied insecure."
]
}
]
},
{
"version": "1.2.3",
"title": "Connection crash fix",
"date": "2026-06-23",
"sections": [
{
"header": "Stability",
"bullets": [
"Fixed a crash that could close the app right after connecting over an encrypted link (Tailscale or HTTPS) — a live secure connection was being torn down on the main thread as it came up. Securing your connection no longer force-closes the app; plain-LAN connections were never affected."
]
}
]
},
{
"version": "1.2.2",
"title": "Multi-profile polish",
"date": "2026-06-22",
"sections": [
{
"header": "Profiles that behave",
"bullets": [
"Deleting a session while a non-default agent profile is active now sticks — it no longer reappears after the list refreshes.",
"On a cold start with a non-default profile selected, the session drawer opens on that profile's chats directly instead of briefly showing the default profile's."
]
},
{
"header": "Clearer diagnostics",
"bullets": [
"Diagnostics is now a full screen led by a top-to-bottom list of subsystem health checks — network, API server, chat transport, pairing, relay, and voice — each with a pass / warning / fail state and the reason when something's wrong; tap a failing check for full detail. The recent-activity log stays below."
]
},
{
"header": "Small touches",
"bullets": [
"The default connection is now simply \"Hermes\" (and the optional power features are labelled \"Relay\"), across setup, the switcher, voice, and permissions.",
"Distraction-free chat mode gives its text a taller, scrollable area."
]
}
]
},
{
"version": "1.2.1",
"title": "Polish & control",
"date": "2026-06-21",
"sections": [
{
"header": "Yours to control",
"bullets": [
"Lock the app to a single agent profile (Settings → Profile lock) and hide the rest from the pickers."
]
},
{
"header": "Find your way back",
"bullets": [
"A new \"What's New\" entry in Settings shows current and past release notes any time — not just after an update."
]
},
{
"header": "When something breaks",
"bullets": [
"Diagnostics show clean error titles — tap any entry for a detail view with Copy, Share, and a one-tap GitHub issue.",
"A tasteful in-app banner tells you when a newer version is live (Play or sideload) — dismissable, and it never nags."
]
},
{
"header": "Voice fixes",
"bullets": [
"Stop now halts realtime speech instantly, hold-to-talk is steadier, the voice overlay is easier to read, and a chosen voice applies in Auto mode.",
"Realtime turns that reach back to Hermes no longer drop with a session error."
]
}
]
},
{
"version": "1.2.0",
"title": "Make it yours",
"date": "2026-06-20",
"sections": [
{
"header": "Personalize",
"bullets": [
"Eight app themes in Settings → Appearance — the Hermes Relay brand plus ports of the Nous Hermes looks (Teal, Nous Blue, Midnight, Ember, Mono, Cyberpunk, Rosé), with light/dark.",
"Swap the agent orb for an animated pet that reacts to what the agent is doing — add, preview, and tune pets right in the app, or generate one from sprite art with the AI authoring kit.",
"Reskin the sphere, and give each agent profile its own icon."
]
},
{
"header": "See what's happening",
"bullets": [
"The chat status strip names the actual streaming path (Gateway, Sessions, Completions, Runs), with a basic→best tier ladder in Chat Settings.",
"Tap the context meter for a \"What the agent sees\" sheet — the exact extra context prepended to your next turn.",
"Voice and Realtime turns are badged in the scrollback."
]
},
{
"header": "Privacy",
"bullets": [
"When paired to the relay, the agent can mark private media and the phone blurs it per your setting — sensitivity stays model-emitted."
]
},
{
"header": "Faster & more reliable",
"bullets": [
"Cold start is about 3× faster, and model/personality/approvals load honestly instead of showing a maybe-wrong value.",
"In-app crash reporting offers a one-tap, pre-filled bug report.",
"QR pairing no longer force-closes on unusual cameras (foldables); fixed crashes opening server images and PDFs; in-chat model picks now apply."
]
},
{
"header": "Voice & terminal",
"bullets": [
"Enhanced voice control for Gemini and xAI providers.",
"Leaner terminal with TUI-correct input and an isolated, tuned tmux."
]
}
]
},
{
"version": "1.1.0",
"title": "Release plumbing & polish",
"date": "2026-06-16",
"sections": [
{
"header": "New",
"bullets": [
"Automated Play Console upload when a release tag ships (a human still starts the rollout).",
"/relay slash commands — status, devices, and pair from any platform — plus a relay-status badge in the dashboard header.",
"The relay plugin prompts for its optional voice-provider keys on install, and a tools-only native install path."
]
},
{
"header": "Improved",
"bullets": [
"Settings overhaul: status pills are now exception-only, Power tools shows a single Plugin active/required/offline badge, and Connections moved to the top.",
"Release names and notes are now split per surface (Android, plugin, CLI)."
]
},
{
"header": "Fixed",
"bullets": [
"No more force-close on connect when the stored credential keyset was corrupt — it now heals in place.",
"The installer works on uv-managed Hermes hosts, and the dashboard relay panel buttons are readable again."
]
}
]
},
{
"version": "1.0.0",
"title": "Stable launch",
"date": "2026-06-14",
"sections": [
{
"header": "Gateway chat with live thinking",
"bullets": [
"Chat can ride the upstream dashboard gateway — the only vanilla-upstream path that streams reasoning live, so the Thinking block and sphere light up during generation. \"Auto\" prefers it and falls back to the SSE endpoints per turn.",
"Desktop parity: native image/PDF/file attachments, mid-turn steering, edit & resend, approval/clarify/sudo/secret cards, live subagent lanes, a context-window meter, server slash commands, and turn-complete notifications.",
"Warm-start and an opt-in Keep connected in background toggle so long-backgrounded conversations resume instantly."
]
},
{
"header": "Agents, Manage & media",
"bullets": [
"Switch agent profiles per conversation — model, SOUL, personality, and skills — with the selection bound to the session, never changing the server default for other clients.",
"Manage parity with the desktop dashboard: change models, manage provider keys, edit profiles and SOUL.md, and browse/install skills.",
"Open and save chat images and attachments — full-screen viewer with pinch-zoom, plus an Open/Share/Save menu."
]
},
{
"header": "Standard path is first-class",
"bullets": [
"Chat, Manage, and voice all work against an unmodified upstream Hermes agent; the relay plugin is now purely additive.",
"Seamless connection UX — LAN↔Tailscale handoffs and reconnects no longer reload the chat, and status shows as in-theme slide-down toasts.",
"Persistent Realtime Agent voice that keeps one session across turns, with long runs promoted to tracked background tasks."
]
}
]
}
]
}
+20 -11
View File
@@ -1,13 +1,22 @@
v1.2.6 - Tidier chats & calmer status
v1.4.0 - Realtime voice that finishes the job
Chats
* Chats no longer get stuck on "Untitled" — your first message stands
in as the title until the chat is named, titles refresh once a turn
settles, and a new refresh button in the session drawer pulls the
latest. Renaming a chat now sticks on non-default agent profiles.
Voice
* Keep talking while long work runs: quick follow-ups can be
answered, another long request can queue, and the finished
answer stays in your selected realtime voice.
* Background and route changes recover more reliably. Stale
listening, thinking, reconnecting, and cancel states clear
instead of trapping the voice screen.
* Pick a Realtime Agent model and voice per connection/profile;
the next session uses it and the choice survives restart.
Calmer status
* Connection status (reconnecting, checking, network handoffs) now
shows as a thin banner at the top that gently slides the screen
down, instead of a card floating over your chat — the floating alert
is kept for persistent errors. Quick confirmations land there too.
More control
* Refresh provider model catalogs from Chat or Manage.
* Opt-in notification rules can offer a local "Ask Hermes?"
action, and Bridge tools can target a specific Android device.
Reliability
* Long chats avoid premature transport fallback, phone context
reaches upstream Hermes on supported paths, malformed server
addresses fail safely, and credential files cannot be served
through relay media.
@@ -15,6 +15,7 @@ import android.os.HandlerThread
import android.util.DisplayMetrics
import android.util.Log
import android.view.WindowManager
import kotlinx.coroutines.delay
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.sync.withLock
import kotlinx.coroutines.withContext
@@ -114,8 +115,27 @@ class ScreenCapture(
*/
private const val MAX_IMAGES = 2
/** Capture timeout — if no frame arrives in this window, fail loudly. */
private const val CAPTURE_TIMEOUT_MS = 2_500L
/**
* Capture timeout — if no frame arrives in this window, fail loudly.
*
* BOOX / e-ink devices can take several seconds before a
* VirtualDisplay-backed ImageReader emits its first frame, especially
* after a fresh MediaProjection grant or when the display is idle. Keep
* the default generous enough for those devices while still bounded so
* a dead capture pipeline reports a clear error.
*/
private const val DEFAULT_CAPTURE_TIMEOUT_MS = 10_000L
/** Optional JVM/system-property override for local QA and OEM tuning. */
private const val CAPTURE_TIMEOUT_PROPERTY =
"hermes.relay.screen_capture_timeout_ms"
private const val MIN_CAPTURE_TIMEOUT_MS = 2_500L
private const val MAX_CAPTURE_TIMEOUT_MS = 30_000L
/** One retry covers stale VirtualDisplay/ImageReader pipelines. */
private const val MAX_CAPTURE_ATTEMPTS = 2
private const val CAPTURE_RETRY_DELAY_MS = 350L
}
// === PHASE3-bridge-ui-followup: MediaProjection reuse fix ===
@@ -211,7 +231,27 @@ class ScreenCapture(
// mutex keeps us honest if anything ever parallelizes.
val pngBytes = try {
captureMutex.withLock {
captureFrame(projection)
var lastTimeout: CaptureTimeoutException? = null
for (attempt in 1..MAX_CAPTURE_ATTEMPTS) {
try {
return@withLock captureFrame(projection)
} catch (e: CaptureTimeoutException) {
lastTimeout = e
Log.w(
TAG,
"screen capture timed out on attempt " +
"$attempt/$MAX_CAPTURE_ATTEMPTS: ${e.message}"
)
if (attempt < MAX_CAPTURE_ATTEMPTS) {
// A timeout can leave an OEM VirtualDisplay path
// wedged without invalidating the MediaProjection
// grant. Rebuild our pipeline once before giving up.
releaseCache()
delay(CAPTURE_RETRY_DELAY_MS)
}
}
}
throw lastTimeout ?: IOException("screen capture timed out")
}
} catch (e: Exception) {
Log.w(TAG, "captureFrame failed: ${e.message}")
@@ -286,16 +326,28 @@ class ScreenCapture(
}
return try {
kotlinx.coroutines.withTimeout(CAPTURE_TIMEOUT_MS) { deferred.await() }
val timeoutMs = captureTimeoutMs()
kotlinx.coroutines.withTimeout(timeoutMs) { deferred.await() }
} catch (e: kotlinx.coroutines.TimeoutCancellationException) {
pendingCaptureRef.compareAndSet(deferred, null)
throw IOException("screen capture timed out")
throw CaptureTimeoutException(
"screen capture timed out after ${captureTimeoutMs()}ms"
)
} catch (t: Throwable) {
pendingCaptureRef.compareAndSet(deferred, null)
throw t
}
}
private fun captureTimeoutMs(): Long {
val configured = System.getProperty(CAPTURE_TIMEOUT_PROPERTY)
?.toLongOrNull()
?.coerceIn(MIN_CAPTURE_TIMEOUT_MS, MAX_CAPTURE_TIMEOUT_MS)
return configured ?: DEFAULT_CAPTURE_TIMEOUT_MS
}
private class CaptureTimeoutException(message: String) : IOException(message)
/**
* Build (or reuse) the cached VirtualDisplay + ImageReader + HandlerThread
* for this projection. Rebuilds when:
@@ -489,7 +541,7 @@ class ScreenCapture(
fastClient.newCall(request).execute().use { response ->
when (response.code) {
200 -> {
val raw = response.body?.string().orEmpty()
val raw = response.body.string()
val token = extractToken(raw)
if (token.isNullOrBlank()) {
Result.failure(
@@ -225,7 +225,7 @@ class RealtimePcmPlayer(context: Context? = null) {
// is the chunk's end frame. The cursor reaches this amplitude once
// playbackHeadPosition passes the previous end frame.
playbackAmpQueue.addLast(FrameAmp(endFrame = totalFramesWritten, rms = rms))
while (playbackAmpQueue.size > MAX_AMP_QUEUE) playbackAmpQueue.removeFirst()
while (playbackAmpQueue.size > MAX_AMP_QUEUE) playbackAmpQueue.removeAt(0)
}
/**
@@ -240,7 +240,7 @@ class RealtimePcmPlayer(context: Context? = null) {
val head = readHeadFrames(track).toLong()
// Drop fully-played chunks so the head of the queue is the one playing now.
while (playbackAmpQueue.size > 1 && playbackAmpQueue.first().endFrame <= head) {
playbackAmpQueue.removeFirst()
playbackAmpQueue.removeAt(0)
}
amplitudeAtHead(playbackAmpQueue, head)
}
@@ -141,6 +141,9 @@ class VoiceRecorder(
fun stopRecording(): File {
val file = currentOutputFile
?: throw IllegalStateException("stopRecording called with no active recording")
// Claim the capture exactly once. A stale UI stop must not repackage
// the previous PCM as a second voice turn.
currentOutputFile = null
val record = audioRecord
stopRequested.set(true)
@@ -207,8 +210,15 @@ class VoiceRecorder(
}
}
updateAmplitude(buffer, read)
} else if (read < 0) {
Log.w(TAG, "AudioRecord.read ended with error code $read")
break
}
}
// Android can terminate capture while the app is backgrounded without
// stopRecording() running. Reflect that loss in isRecording() so the
// foreground UI can recover instead of remaining stuck on Listening.
stopRequested.set(true)
}
private fun updateAmplitude(buffer: ByteArray, read: Int) {
@@ -108,9 +108,36 @@ object DemoContent {
),
)
/**
* Assistant reply appended when the user sends a message INSIDE demo
* mode. The composer must not be a silent no-op (it reads as broken —
* see the demo-polish TODO), but there is no server to answer, so the
* "reply" is an honest notice pointing at the exit path. Same content
* contract as the transcript: clientOnly, terminal, zero network.
*
* @param id unique message id supplied by the caller (UUID-based; two
* rapid sends must not collide on LazyColumn keys).
* @param nowMs wall-clock timestamp for the bubble.
*/
fun composerReply(id: String, nowMs: Long): ChatMessage = ChatMessage(
id = id,
role = MessageRole.ASSISTANT,
content = COMPOSER_REPLY,
timestamp = nowMs,
agentName = DEMO_AGENT_NAME,
badges = listOf("Demo"),
clientOnly = true,
)
// --- Message bodies (Markdown). Kept as constants so the content is easy
// to scan and the [transcript] builder stays readable. ---
private val COMPOSER_REPLY: String = """
This is the offline demo, so I can't answer for real — nothing here talks to a server.
Connect your own Hermes server to chat live: tap **Connect** in the demo banner above.
""".trimIndent()
private val ASSISTANT_TOUR: String = """
I'm **Hermes**, the agent running on *your* server. Here's a quick tour of what this app surfaces:
@@ -44,6 +44,9 @@ data class VoiceSettings(
* docs/plans/2026-05-24-realtime-persistent-session.md.
*/
val realtimePersistentSession: Boolean = true,
/** Per-profile Realtime Agent session overrides; blank uses relay config. */
val realtimeModel: String = "",
val realtimeVoice: String = "",
/**
* Enhanced-voice overrides for the relay TTS path, mapped onto the active
* provider (Gemini / xAI). Empty string / false means "use the server's
@@ -154,9 +157,9 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
// over the hard default — see [scopedName] / [resolveString].
//
// Why these are per-profile: engine mode, audio route, and the
// enhanced-voice overrides describe *which voice the agent speaks
// with*, which is a property of the profile (the relay already
// persists `voice_output:`/`realtime_voice:` per profile and
// enhanced-voice and realtime-session overrides describe *which voice
// the agent speaks with*, which is a property of the profile (the relay
// already persists `voice_output:`/`realtime_voice:` per profile and
// `RelayVoiceClient` already sends `?profile=`). Keeping them global
// leaked one profile's voice onto every other profile.
private const val KEY_ENGINE_MODE = "voice_engine_mode"
@@ -166,6 +169,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
private const val KEY_ENH_AUDIO_TAGS = "voice_enh_audio_tags"
private const val KEY_ENH_PERSONA = "voice_enh_persona"
private const val KEY_ENH_LANGUAGE = "voice_enh_language"
private const val KEY_REALTIME_MODEL = "voice_realtime_model"
private const val KEY_REALTIME_VOICE = "voice_realtime_voice"
// --- Global keys (shared across profiles; never namespaced) ----------
// Why these stay global: interaction-mode and silence-threshold are
@@ -214,8 +219,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
/**
* Point the repository at a (connection, profile) scope. Per-profile reads
* and writes (engine/route/enhanced) re-target the namespaced keys for that
* profile; global prefs are unaffected. Passing a null/blank profile name
* and writes (engine/route/enhanced/realtime) re-target the namespaced keys
* for that profile; global prefs are unaffected. Passing a null/blank profile name
* reverts per-profile reads/writes to the global base layer (the default
* profile). Idempotent — a no-op when the normalized scope is unchanged.
*/
@@ -248,6 +253,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
enhancedAudioTags = resolveBoolean(prefs, KEY_ENH_AUDIO_TAGS, scope, false),
enhancedPersona = resolveString(prefs, KEY_ENH_PERSONA, scope, ""),
enhancedLanguage = resolveString(prefs, KEY_ENH_LANGUAGE, scope, ""),
realtimeModel = resolveString(prefs, KEY_REALTIME_MODEL, scope, ""),
realtimeVoice = resolveString(prefs, KEY_REALTIME_VOICE, scope, ""),
// --- global (shared across profiles) ---
interactionMode = prefs[KEY_INTERACTION_MODE] ?: DEFAULT_INTERACTION_MODE,
silenceThresholdMs = prefs[KEY_SILENCE_THRESHOLD_MS] ?: DEFAULT_SILENCE_THRESHOLD_MS,
@@ -327,6 +334,29 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
dataStore.edit { it[key] = language.trim() }
}
/** "" clears the override so new sessions use the relay's saved model. */
suspend fun setRealtimeModel(model: String) {
val key = stringPreferencesKey(scopedName(KEY_REALTIME_MODEL, _scope.value))
dataStore.edit { it[key] = model.trim() }
}
/** "" clears the override so new sessions use the relay's saved voice. */
suspend fun setRealtimeVoice(voice: String) {
val key = stringPreferencesKey(scopedName(KEY_REALTIME_VOICE, _scope.value))
dataStore.edit { it[key] = voice.trim() }
}
/** Persist a compatible model/voice pair without exposing a half-updated snapshot. */
suspend fun setRealtimeSelection(model: String, voice: String) {
val scope = _scope.value
val modelKey = stringPreferencesKey(scopedName(KEY_REALTIME_MODEL, scope))
val voiceKey = stringPreferencesKey(scopedName(KEY_REALTIME_VOICE, scope))
dataStore.edit {
it[modelKey] = model.trim()
it[voiceKey] = voice.trim()
}
}
// --- global setters (always the un-namespaced key) -----------------------
suspend fun setInteractionMode(mode: String) {
@@ -177,6 +177,14 @@ object DiagnosticsLog {
return noUserInfo.take(MAX_TEXT_LENGTH)
}
/**
* Public secret redaction for user-composed report text (e.g. the "what
* were you expecting?" answer embedded in a GitHub issue body). Same
* redaction + cap as the stored stacktraces — entry fields are already
* sanitized at record time; this covers text added after the fact.
*/
fun redactReportText(value: String?): String? = redactTrace(value)
private fun clean(value: String?): String? {
val trimmed = value?.trim()?.takeIf { it.isNotBlank() } ?: return null
return redact(trimmed).take(MAX_TEXT_LENGTH)
@@ -169,7 +169,7 @@ object EventStore {
)
if (buffer.size >= MAX_ENTRIES) {
buffer.removeFirst()
buffer.removeAt(0)
}
buffer.addLast(entry)
}
@@ -42,6 +42,20 @@ enum class ConnectionState {
Reconnecting
}
/**
* Build an OkHttp request for a relay socket URL, or `null` if the URL is
* malformed. OkHttp's [Request.Builder.url] throws [IllegalArgumentException]
* on an invalid host; the relay connect runs on a background coroutine, so an
* uncaught throw crashes the app (the #131 "Invalid URL host" class). Callers
* treat `null` as a connection failure instead of letting it propagate.
*/
internal fun buildRelayRequestOrNull(url: String): Request? =
try {
Request.Builder().url(url).build()
} catch (e: IllegalArgumentException) {
null
}
class ConnectionManager(
private val multiplexer: ChannelMultiplexer,
/**
@@ -828,9 +842,30 @@ class ConnectionManager(
authenticated = false
client = buildClient()
val request = Request.Builder()
.url(url)
.build()
val request = buildRelayRequestOrNull(url)
if (request == null) {
// A malformed relay URL (an invalid/empty host from a corrupt or
// hand-edited pairing payload) can't be built into a request. This
// runs on a background coroutine, so letting OkHttp's url() throw
// would crash the app — the #131 "Invalid URL host" class, relay-
// socket half. Route it through the same path onFailure uses.
Log.e(TAG, "doConnect: malformed relay URL '$url' — not connecting")
DiagnosticsLog.record(
category = DiagnosticCategory.Relay,
severity = DiagnosticSeverity.Error,
title = "Invalid relay URL",
detail = "The relay address could not be parsed; re-pair to refresh it.",
url = url,
)
authenticated = false
_connectionState.value = ConnectionState.Disconnected
previousSocketToClose?.let { stale ->
runCatching { stale.close(1000, replaceReason) }
stale.cancel()
}
scheduleReconnect()
return
}
Log.i(TAG, "doConnect: opening WSS to $url")
val newSocket = client.newWebSocket(request, object : WebSocketListener() {
@@ -13,6 +13,7 @@ import kotlinx.serialization.builtins.ListSerializer
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import okhttp3.HttpUrl.Companion.toHttpUrl
import okhttp3.HttpUrl.Companion.toHttpUrlOrNull
import okhttp3.MediaType.Companion.toMediaType
import okhttp3.OkHttpClient
import okhttp3.Request
@@ -153,7 +154,10 @@ class RelayHttpClient(
.replace(Regex("^ws://", RegexOption.IGNORE_CASE), "http://")
.trimEnd('/')
val url = "$httpBase/media/$token"
val url = "$httpBase/media/$token".toHttpUrlOrNull()
?: return@withContext Result.failure(
IllegalArgumentException("Invalid relay URL: $httpBase")
)
val request = Request.Builder()
.url(url)
@@ -597,7 +601,10 @@ class RelayHttpClient(
.replace(Regex("^ws://", RegexOption.IGNORE_CASE), "http://")
.trimEnd('/')
val url = "$httpBase/sessions"
val url = "$httpBase/sessions".toHttpUrlOrNull()
?: return@withContext Result.failure(
IllegalArgumentException("Invalid relay URL: $httpBase")
)
val request = Request.Builder()
.url(url)
.get()
File diff suppressed because it is too large Load Diff
@@ -15,6 +15,7 @@ import com.hermesandroid.relay.network.upstream.GatewaySubagentEvent
import com.hermesandroid.relay.network.upstream.models.MessageItem
import com.hermesandroid.relay.network.upstream.models.RelayStreamEventEnvelope
import com.hermesandroid.relay.network.upstream.models.SessionItem
import com.hermesandroid.relay.voice.RealtimeTurnSyncBuilder
import kotlinx.coroutines.flow.MutableStateFlow
import kotlinx.coroutines.flow.StateFlow
import kotlinx.coroutines.flow.asStateFlow
@@ -220,6 +221,22 @@ class ChatHandler {
private val _isStreaming = MutableStateFlow(false)
val isStreaming: StateFlow<Boolean> = _isStreaming.asStateFlow()
/**
* Silently drop the global streaming flag + turn-status caption without
* touching the message list, error, or per-message `isStreaming` flags.
* For abandoning an in-flight answer recovery (issue #166) on a path that
* clears or reloads the transcript itself: there is no placeholder to
* finalize and nothing went wrong, so [onStreamComplete] (which reconciles
* a specific message) and [onStreamError] (which raises an error banner)
* are both the wrong tool. Leaves per-message streaming flags intact so a
* caller that still needs to find the placeholder afterwards (e.g.
* cancelStream's Stopped-badge pass) can.
*/
fun clearStreamingStatus() {
_isStreaming.value = false
_turnStatus.value = null
}
private val _sessions = MutableStateFlow<List<ChatSession>>(emptyList())
val sessions: StateFlow<List<ChatSession>> = _sessions.asStateFlow()
@@ -460,6 +477,11 @@ class ChatHandler {
}
}
/** Remove a provisional client-side message that never became a real turn. */
fun removeMessage(messageId: String) {
_messages.update { messages -> messages.filterNot { it.id == messageId } }
}
/**
* Append a local-only voice-intent trace to the chat scroll. Used by
* the sideload voice intent flow (`RealVoiceBridgeIntentHandler`) so
@@ -1011,6 +1033,11 @@ class ChatHandler {
}
}
// Trimmed assistant texts of synced provider-answered realtime turns
// found in this reload — used below to drop their superseded local
// clientOnly bubbles (same exchange, pre-sync copy).
val syncedRealtimeTurnContents = mutableSetOf<String>()
val loaded = items.mapNotNull { item ->
val role = when (item.role) {
"user" -> MessageRole.USER
@@ -1060,7 +1087,7 @@ class ChatHandler {
// straight onto the reconstructed ChatMessage and strip their
// lines from the displayed content in the same pass. No
// post-assignment dispatch needed.
val (cleanedContent, extractedCards) = if (
val (cardCleanedContent, extractedCards) = if (
role == MessageRole.ASSISTANT && afterMedia.isNotEmpty()
) {
extractCardsFromContent(afterMedia)
@@ -1068,6 +1095,23 @@ class ChatHandler {
afterMedia to emptyList()
}
// A provider-answered realtime voice turn synced into the session
// (RealtimeTurnSyncBuilder) carries a trailing provenance marker —
// "[Realtime Agent provider-native voice turn: provider=…]" — in
// its assistant text. Render it as the quiet "Realtime Agent"
// badge (same chip live turns get) instead of raw bracket noise,
// and remember the stripped text so the superseded local
// clientOnly bubble can be dropped below instead of duplicating
// the exchange.
val strippedRealtimeContent = if (role == MessageRole.ASSISTANT) {
RealtimeTurnSyncBuilder.stripProvenanceMarker(cardCleanedContent)
} else {
null
}
val isSyncedRealtimeTurn = strippedRealtimeContent != null
val cleanedContent = strippedRealtimeContent ?: cardCleanedContent
if (isSyncedRealtimeTurn) syncedRealtimeTurnContents.add(cleanedContent.trim())
val prior = priorById[messageId]
// Outbound attachments: prefer an id-match (covers any future
// user-message id reconciliation), else fall back to the
@@ -1119,6 +1163,11 @@ class ChatHandler {
} else {
""
},
badges = if (isSyncedRealtimeTurn && "Realtime Agent" !in prior.badges) {
prior.badges + "Realtime Agent"
} else {
prior.badges
},
)
} else {
// INSERT — a server message with no local row yet. Built from
@@ -1137,6 +1186,7 @@ class ChatHandler {
// Server persists per-message reasoning — restore it so the
// Thought-process block survives returning to the chat.
thinkingContent = if (role == MessageRole.ASSISTANT) serverThinking ?: "" else "",
badges = if (isSyncedRealtimeTurn) listOf("Realtime Agent") else emptyList(),
)
}
}
@@ -1164,7 +1214,19 @@ class ChatHandler {
// but IS in the transcript, so it reconciles normally; only clientOnly +
// absent-from-transcript marks a preservable orphan.
val loadedIds = loaded.mapTo(HashSet()) { it.id }
val preservedLocal = _messages.value.filter { it.clientOnly && it.id !in loadedIds }
val preservedLocal = _messages.value.filter { msg ->
if (!msg.clientOnly || msg.id in loadedIds) return@filter false
// Drop a provider-answered realtime bubble whose SYNCED copy just
// loaded from the server transcript (matched on the synced
// assistant text) — keeping both would render the exchange twice.
// Unsynced traces are always preserved: they are still the only
// record of the turn.
val trace = msg.realtimeTurn
!(
trace != null && trace.syncedToServer &&
trace.assistantText.trim() in syncedRealtimeTurnContents
)
}
val merged = if (preservedLocal.isEmpty()) {
loaded
} else {
@@ -2644,6 +2706,9 @@ class ChatHandler {
fun onStreamError(message: String) {
_isStreaming.value = false
// The turn is over — a stale lifecycle/recovery caption must not
// outlive it (onStreamComplete clears the same way).
_turnStatus.value = null
_error.value = message
// Clear streaming flag on any actively streaming message
_messages.update { messages ->
@@ -7,6 +7,9 @@ import com.hermesandroid.relay.network.upstream.models.MessageItem
import com.hermesandroid.relay.network.upstream.models.MessageListResponse
import com.hermesandroid.relay.network.upstream.models.SessionItem
import com.hermesandroid.relay.network.upstream.models.SessionListResponse
import com.hermesandroid.relay.network.upstream.models.SessionPruneFilters
import com.hermesandroid.relay.network.upstream.models.SessionPrunePreview
import com.hermesandroid.relay.network.upstream.models.SessionPruneResult
import com.hermesandroid.relay.auth.SecureStoreCache
import com.hermesandroid.relay.auth.SessionTokenStore
import com.hermesandroid.relay.auth.buildRawTokenStore
@@ -270,8 +273,27 @@ class DashboardApiClient(
getJson("/api/audio/elevenlabs/voices").mapCatching { parseElevenLabsVoices(it) }
}
/** Full provider/model universe — REST twin of the TUI's `model.options` RPC. */
suspend fun getModelOptions(): Result<JsonObject> = getJsonObject("/api/model/options")
/**
* Full provider/model universe — REST twin of the TUI's `model.options` RPC.
*
* Always opts into `include_unconfigured=1`: newer upstream defaults this
* route to configured-providers-only, which would silently drop the
* unauthenticated skeleton rows Manage renders as its Keys-setup
* affordance. Older upstream returned the full universe by default and
* ignores the extra param, so both generations serve the same catalog.
*
* [refresh] maps to upstream's explicit `refresh=1` path, which refreshes
* dynamic/custom-provider catalogs on demand without probing every
* provider during normal picker opens.
*/
suspend fun getModelOptions(refresh: Boolean = false): Result<JsonObject> =
getJsonObject(
if (refresh) {
"/api/model/options?refresh=1&include_unconfigured=1"
} else {
"/api/model/options?include_unconfigured=1"
},
)
/**
* Assign the main model in `~/.hermes/config.yaml` (new sessions only).
@@ -508,7 +530,11 @@ class DashboardApiClient(
* ordering where the host honors it. Android still sorts by decoded
* `last_active` locally because older hosts return started-time order.
*/
suspend fun listSessions(profile: String? = null, limit: Int = 200): Result<List<SessionItem>> =
suspend fun listSessions(
profile: String? = null,
limit: Int = 200,
archived: String? = null,
): Result<List<SessionItem>> =
withContext(Dispatchers.IO) {
val query = buildList {
add("limit=${limit.coerceIn(1, 200)}")
@@ -516,6 +542,10 @@ class DashboardApiClient(
add("min_messages=1")
val name = profile?.trim().orEmpty()
if (name.isNotBlank()) add("profile=${pathSegment(name)}")
// Upstream `archived` filter: exclude (default) | only | include.
// Omitted unless requested so older hosts see an unchanged request.
val archivedMode = archived?.trim().orEmpty()
if (archivedMode.isNotBlank()) add("archived=${pathSegment(archivedMode)}")
}.joinToString(prefix = "?", separator = "&")
getJson("/api/sessions$query").mapCatching { root ->
val parsed = json.decodeFromJsonElement(SessionListResponse.serializer(), root)
@@ -555,19 +585,98 @@ class DashboardApiClient(
suspend fun deleteSession(sessionId: String, profile: String? = null): Result<JsonObject> =
deleteJsonObject("/api/sessions/${pathSegment(sessionId)}${profileQuery(profile)}")
/**
* Export one session as server-owned JSON metadata + messages. This is the
* safe "archive a copy before cleanup" primitive for clients that want to
* offer download/share before a destructive delete or prune. Profile scoping
* matches [deleteSession].
*/
suspend fun exportSession(sessionId: String, profile: String? = null): Result<JsonObject> =
getJsonObject("/api/sessions/${pathSegment(sessionId)}/export${profileQuery(profile)}")
/**
* Rename a session scoped to a profile via the dashboard
* `PATCH /api/sessions/{id}?profile=` surface — the write twin of
* [deleteSession]. A non-default profile's sessions live in that profile's
* own `state.db`, so the unscoped api_server rename would patch the wrong
* DB and the new title would never appear in the profile-scoped list.
* `PATCH /api/sessions/{id}` surface — the write twin of [deleteSession].
* A non-default profile's sessions live in that profile's own `state.db`,
* so the unscoped api_server rename would patch the wrong DB and the new
* title would never appear in the profile-scoped list. Current upstream
* reads `profile` from the PATCH body (`SessionRename`); the query param
* rides along for builds that scoped by query.
*/
suspend fun renameSession(sessionId: String, title: String, profile: String? = null): Result<JsonObject> =
patchJsonObject(
"/api/sessions/${pathSegment(sessionId)}${profileQuery(profile)}",
buildJsonObject { put("title", title) },
buildJsonObject {
put("title", title)
profile?.trim()?.takeIf { it.isNotBlank() }?.let { put("profile", it) }
},
)
/**
* Soft-archive or restore a session via the same dashboard
* `PATCH /api/sessions/{id}` surface (`{archived: true|false}`). Archived
* sessions drop out of the default list and are excluded from a prune
* unless [SessionPruneFilters.includeArchived] is set; list them back with
* [listSessions] `archived = "only"`. Profile scoping matches
* [renameSession]: body for current upstream, query for older builds.
*/
suspend fun setSessionArchived(
sessionId: String,
archived: Boolean,
profile: String? = null,
): Result<JsonObject> =
patchJsonObject(
"/api/sessions/${pathSegment(sessionId)}${profileQuery(profile)}",
buildJsonObject {
put("archived", archived)
profile?.trim()?.takeIf { it.isNotBlank() }?.let { put("profile", it) }
},
)
/**
* Dry-run a server-backed bulk session cleanup via the dashboard
* `POST /api/sessions/prune` (`dry_run: true`). Returns what WOULD be
* deleted — matched count, started-at span, and the candidate rows —
* without deleting anything. This is the required first step of the
* prune flow: show the preview, then pass it to [pruneSessions].
*/
suspend fun previewSessionPrune(filters: SessionPruneFilters): Result<SessionPrunePreview> =
postJsonObject("/api/sessions/prune", filters.toPrunePayload(dryRun = true))
.mapCatching { root ->
json.decodeFromJsonElement(SessionPrunePreview.serializer(), root)
}
/**
* Apply a server-backed bulk session cleanup (`POST /api/sessions/prune`,
* `dry_run: false`). Destructive — [confirmedPreview] is required so no
* caller can reach this without first running [previewSessionPrune] with
* the same [filters] and showing the user its count/span. A preview that
* matched nothing short-circuits without touching the server: sessions
* that aged into the filter after the preview are not covered by what the
* user confirmed.
*/
suspend fun pruneSessions(
filters: SessionPruneFilters,
confirmedPreview: SessionPrunePreview,
): Result<SessionPruneResult> {
if (confirmedPreview.matched <= 0) {
return Result.success(SessionPruneResult(ok = true, removed = 0))
}
return postJsonObject("/api/sessions/prune", filters.toPrunePayload(dryRun = false))
.mapCatching { root ->
json.decodeFromJsonElement(SessionPruneResult.serializer(), root)
}
}
private fun SessionPruneFilters.toPrunePayload(dryRun: Boolean): JsonObject =
buildJsonObject {
olderThanDays?.let { put("older_than_days", it) }
source?.trim()?.takeIf { it.isNotBlank() }?.let { put("source", it) }
profile?.trim()?.takeIf { it.isNotBlank() }?.let { put("profile", it) }
if (includeArchived) put("include_archived", true)
put("dry_run", dryRun)
}
private fun parseProfiles(root: JsonObject): List<Profile> {
fun decode(element: JsonElement, nameOverride: String?): Profile? = runCatching {
val obj = element as? JsonObject ?: return null
@@ -71,11 +71,23 @@ class GatewayChatClient(
private val scope: CoroutineScope = CoroutineScope(SupervisorJob() + Dispatchers.IO),
/** Max wall-clock a single mid-turn reconnect keeps retrying before failing the turn. */
private val midTurnRejoinWindowMs: Long = MAX_MIDTURN_REJOIN_MS,
/** Test seam — generic RPC ack timeout. Production keeps [RPC_TIMEOUT_MS]. */
private val rpcTimeoutMs: Long = RPC_TIMEOUT_MS,
/** Test seam — `prompt.submit` ack ceiling. Production keeps [PROMPT_SUBMIT_REQUEST_TIMEOUT_MS]. */
private val promptSubmitTimeoutMs: Long = PROMPT_SUBMIT_REQUEST_TIMEOUT_MS,
/** Test seam — idle-progress watchdog base. Production keeps [TURN_TIMEOUT_MS]. */
private val turnIdleTimeoutMs: Long = TURN_TIMEOUT_MS,
) {
companion object {
private const val TAG = "GatewayChatClient"
/** Mirrors the desktop CLI's turn timeout — reset on every received event. */
/**
* Idle-progress turn watchdog — reset on EVERY received gateway event
* (deltas, tool events, status lines), so it only fires after this
* long with no events at all. It is NOT a hard turn cap: a turn that
* keeps streaming lives indefinitely, and a slow `prompt.submit` ack
* is bounded separately by [PROMPT_SUBMIT_REQUEST_TIMEOUT_MS].
*/
private const val TURN_TIMEOUT_MS = 180_000L
/**
@@ -88,14 +100,22 @@ class GatewayChatClient(
private const val ASK_SUDO_TIMEOUT_MS = 150_000L
private const val ASK_UNBOUNDED_TIMEOUT_MS = 600_000L
private fun watchdogTimeoutFor(eventType: String): Long = when (eventType) {
"clarify.request", "secret.request" -> ASK_CLARIFY_SECRET_TIMEOUT_MS
"sudo.request" -> ASK_SUDO_TIMEOUT_MS
"approval.request", "terminal.read.request" -> ASK_UNBOUNDED_TIMEOUT_MS
else -> TURN_TIMEOUT_MS
}
private const val RPC_TIMEOUT_MS = 15_000L
/**
* `prompt.submit` ack ceiling — mirrors upstream desktop's
* PROMPT_SUBMIT_REQUEST_TIMEOUT_MS (apps/desktop/src/hermes.ts,
* upstream commit 164144183). The submit is effectively
* fire-and-forget: turn completion is signaled by stream events
* (`message.complete`), NOT by the RPC return, and MoA/deep-reasoning/
* tool-heavy turns can legitimately take minutes to ack. Bounding the
* ack by [RPC_TIMEOUT_MS] false-failed a running turn into the SSE
* preflight fallback — which resubmits the same prompt → duplicate
* turn. Matches the backend's own agent-turn ceiling
* (agent.gateway_timeout = 1800s), so this only fires when the turn
* would have been abandoned server-side anyway.
*/
private const val PROMPT_SUBMIT_REQUEST_TIMEOUT_MS = 1_800_000L
private const val CONNECT_TIMEOUT_MS = 20_000L
/**
@@ -417,8 +437,26 @@ class GatewayChatClient(
put("text", text)
truncateBeforeUserOrdinal?.let { put("truncate_before_user_ordinal", it) }
},
// Long-running RPC, not a generic 15s ack — see the
// constant's doc. The idle watchdog (armed above, reset by
// every event) owns liveness while this await is pending.
timeoutMs = promptSubmitTimeoutMs,
)
if (submitted.isFailure) {
// Once this turn's own events are flowing (or it already
// finished), the prompt provably reached the server — a
// slow, lost, or socket-severed ack must NOT preflight-fail
// into the SSE fallback, which would resubmit the same
// prompt as a duplicate turn. Recovery belongs to the
// stream: the watchdog and mid-turn rejoin own it.
if (turn.started || turn.ended) {
Log.w(
TAG,
"prompt.submit ack failed after turn start " +
"(${submitted.exceptionOrNull()?.message}) — no SSE fallback",
)
return@launch
}
activeTurn = null
turn.disarmWatchdog()
throw GatewayPreflightException(
@@ -723,7 +761,7 @@ class GatewayChatClient(
* generic alias. Connects on demand if needed. Switching a model is then a
* `/model <model> --provider <slug>` slash dispatch.
*/
suspend fun modelOptions(): Result<GatewayModelOptions> {
suspend fun modelOptions(refresh: Boolean = false): Result<GatewayModelOptions> {
if (webSocket == null || readySignal?.isCompleted != true) {
try {
connectMutex.withLock { ensureConnected() }
@@ -731,7 +769,10 @@ class GatewayChatClient(
return Result.failure(e)
}
}
val params = buildJsonObject { liveSessionId?.let { put("session_id", it) } }
val params = buildJsonObject {
liveSessionId?.let { put("session_id", it) }
if (refresh) put("refresh", true)
}
return rpc("model.options", params).map { result ->
val providers = (result["providers"] as? JsonArray).orEmpty().mapNotNull { el ->
val obj = el as? JsonObject ?: return@mapNotNull null
@@ -1294,7 +1335,7 @@ class GatewayChatClient(
retargetedThisTurn = false
POST_RETARGET_SETTLE_MS
} else {
TURN_TIMEOUT_MS
turnIdleTimeoutMs
},
)
return
@@ -1336,7 +1377,7 @@ class GatewayChatClient(
private suspend fun rpc(
method: String,
params: JsonObject,
timeoutMs: Long = RPC_TIMEOUT_MS,
timeoutMs: Long = rpcTimeoutMs,
): Result<JsonObject> {
val socket = webSocket ?: return Result.failure(GatewayRpcException("not connected"))
val id = rpcId.getAndIncrement()
@@ -1448,6 +1489,14 @@ class GatewayChatClient(
// Turn handle
// ------------------------------------------------------------------
/** Per-event idle-watchdog duration — asks block server-side with no events, so they arm longer. */
private fun watchdogTimeoutFor(eventType: String): Long = when (eventType) {
"clarify.request", "secret.request" -> ASK_CLARIFY_SECRET_TIMEOUT_MS
"sudo.request" -> ASK_SUDO_TIMEOUT_MS
"approval.request", "terminal.read.request" -> ASK_UNBOUNDED_TIMEOUT_MS
else -> turnIdleTimeoutMs
}
private inner class GatewayTurn(
val callbacks: GatewayTurnCallbacks,
) : ActiveTurnHandle {
@@ -1470,7 +1519,19 @@ class GatewayChatClient(
val ended: Boolean get() = mapper.turnEnded || cancelled
/**
* True once any turn-scoped event has arrived — proof the server
* received the submit and is running the turn. `session.info` doesn't
* count: it's connection-level (resume/config echoes) and can arrive
* independent of this turn, so it must not suppress a legitimate
* preflight fallback.
*/
@Volatile
var started = false
private set
fun onEvent(type: String, payload: JsonObject?) {
if (type != "session.info") started = true
tracer.mark("ttfe")
if (type == "message.delta" || type == "reasoning.delta" || type == "thinking.delta") {
tracer.mark("ttft")
@@ -1486,7 +1547,7 @@ class GatewayChatClient(
}
}
fun armWatchdog(timeoutMs: Long = TURN_TIMEOUT_MS) {
fun armWatchdog(timeoutMs: Long = turnIdleTimeoutMs) {
watchdog?.cancel()
watchdog = scope.launch {
delay(timeoutMs)
@@ -28,6 +28,7 @@ import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.booleanOrNull
import kotlinx.serialization.json.contentOrNull
import kotlinx.serialization.json.decodeFromJsonElement
import okhttp3.HttpUrl.Companion.toHttpUrlOrNull
import okhttp3.MediaType.Companion.toMediaType
import okhttp3.OkHttpClient
import okhttp3.Request
@@ -178,6 +179,32 @@ class HermesApiClient(
companion object {
private const val TAG = "HermesApiClient"
private val JSON_MEDIA = "application/json".toMediaType()
/**
* Prefix stamped by [streamFailureMessage] on stream failures raised
* by the transport layer (the IOException family: socket reset/close,
* DNS, TLS, timeouts) as opposed to a server-reported error. The
* dropped-stream answer recovery (issue #166) keys on it via
* [isTransportStreamError].
*/
const val TRANSPORT_ERROR_PREFIX = "Connection failed"
/**
* True when a stream `onError` message came from a transport-layer
* failure (see [TRANSPORT_ERROR_PREFIX]) — the class of error where
* the server may still be running (and persisting) the turn.
*/
fun isTransportStreamError(errorMsg: String): Boolean =
errorMsg.startsWith(TRANSPORT_ERROR_PREFIX)
/** Shared human-readable message for an SSE [EventSourceListener.onFailure]. */
private fun streamFailureMessage(t: Throwable?, response: Response?): String = when {
response != null && !response.isSuccessful ->
"API error ${response.code}: ${response.message}"
t is IOException -> "$TRANSPORT_ERROR_PREFIX: ${t.message}"
t != null -> "Stream error: ${t.message}"
else -> "Unknown stream error"
}
}
private val mainHandler = Handler(Looper.getMainLooper())
@@ -522,6 +549,9 @@ class HermesApiClient(
* blank the `model` field is omitted entirely and the server falls
* back to its session default. Used by the agent-profile picker so
* an explicit user choice wins over implicit session/server defaults.
* Best-effort hint: current native upstream does not parse `model`
* on this route (legacy fork builds honor it) — see the contract
* notes in `HermesChatPayloads.kt`.
*/
fun sendChatStream(
sessionId: String,
@@ -529,23 +559,25 @@ class HermesApiClient(
systemMessage: String? = null,
attachments: List<com.hermesandroid.relay.data.Attachment>? = null,
/**
* Pre-built OpenAI-format synthetic messages to splice into the
* payload alongside the live `message`. Produced by
* Pre-built OpenAI-format synthetic messages carrying phone-local
* context (voice intents, card dispatches, realtime voice turns).
* Produced by
* [com.hermesandroid.relay.voice.VoiceIntentSyncBuilder.buildSyntheticMessages]
* for the v0.4.1 voice-intent → server session sync feature.
* and its twin builders; the param name is historical — it accepts
* any synthetic-message array.
*
* When non-empty, the request body grows a top-level `messages`
* array containing the synthetic `assistant` (with `tool_calls`)
* + `tool` (with `tool_call_id`) pairs. The server-side session
* absorbs them into its conversation history so the LLM sees
* prior phone-local voice actions in its session memory.
* Upstream's session-chat handler consumes only `message` and
* `system_message` — a top-level `messages` array is NOT parsed
* (verified in `gateway/platforms/api_server.py`,
* `_handle_session_chat_stream`), so these can't ride the request
* as real history entries. Instead [buildSessionChatStreamPayload]
* renders them as a plain-text digest folded into this turn's
* ephemeral `system_message`. The model sees the context for THIS
* turn only; it is not persisted server-side. See the mapping notes
* in `HermesChatPayloads.kt`.
*
* Null / empty on every send that has no unsynced voice intents
* to communicate, which is the common case after the first sync.
* The Hermes API server treats unrecognised body fields
* permissively (matches OpenAI Chat Completions semantics), so
* this stays a safe additive change against any conformant
* upstream.
* Null / empty on every send that has no unsynced traces to
* communicate, which is the common case after the first sync.
*/
voiceIntentMessages: JsonArray? = null,
onSessionId: (String) -> Unit,
@@ -568,7 +600,7 @@ class HermesApiClient(
AgentDisplay.profileRequestName(profileName)?.let {
Log.d(TAG, "sendChatStream: profile=$it")
}
val requestPayload = buildSessionChatStreamPayload(
val built = buildSessionChatStreamPayload(
message = message,
systemMessage = systemMessage,
attachments = attachments,
@@ -576,12 +608,19 @@ class HermesApiClient(
modelOverride = modelOverride,
profileName = profileName,
)
val requestBody = json.encodeToString(JsonObject.serializer(), requestPayload)
logDroppedAttachments("sessions chat/stream", built.droppedAttachments)
val requestBody = json.encodeToString(JsonObject.serializer(), built.payload)
val request = authRequest("$baseUrl/api/sessions/$sessionId/chat/stream")
.header("Accept", "text/event-stream")
.post(requestBody.toRequestBody(JSON_MEDIA))
.build()
val request = authRequestOrNull("$baseUrl/api/sessions/$sessionId/chat/stream")
?.header("Accept", "text/event-stream")
?.post(requestBody.toRequestBody(JSON_MEDIA))
?.build()
?: run {
// #131: malformed base URL — fail the turn through the normal
// error channel instead of throwing out of the ViewModel.
mainHandler.post { onError(invalidBaseUrlMessage()) }
return failedEventSource()
}
val completeCalled = AtomicBoolean(false)
// Comparable to the gateway's turn[gateway] line — see TurnLatencyTracer.
@@ -762,13 +801,7 @@ class HermesApiClient(
) {
tracer.done("error")
if (completeCalled.compareAndSet(false, true)) {
val msg = when {
response != null && !response.isSuccessful ->
"API error ${response.code}: ${response.message}"
t is IOException -> "Connection failed: ${t.message}"
t != null -> "Stream error: ${t.message}"
else -> "Unknown stream error"
}
val msg = streamFailureMessage(t, response)
mainHandler.post { onError(msg) }
}
}
@@ -819,7 +852,7 @@ class HermesApiClient(
AgentDisplay.profileRequestName(profileName)?.let {
Log.d(TAG, "sendChatCompletionsStream: profile=$it")
}
val requestPayload = buildChatCompletionsStreamPayload(
val built = buildChatCompletionsStreamPayload(
message = message,
model = model,
systemMessage = systemMessage,
@@ -828,12 +861,18 @@ class HermesApiClient(
modelOverride = modelOverride,
profileName = profileName,
)
val requestBody = json.encodeToString(JsonObject.serializer(), requestPayload)
logDroppedAttachments("chat completions", built.droppedAttachments)
val requestBody = json.encodeToString(JsonObject.serializer(), built.payload)
val request = authRequest("$baseUrl/v1/chat/completions")
.header("Accept", "text/event-stream")
.post(requestBody.toRequestBody(JSON_MEDIA))
.build()
val request = authRequestOrNull("$baseUrl/v1/chat/completions")
?.header("Accept", "text/event-stream")
?.post(requestBody.toRequestBody(JSON_MEDIA))
?.build()
?: run {
// #131: malformed base URL — see sendChatStream.
mainHandler.post { onError(invalidBaseUrlMessage()) }
return failedEventSource()
}
val completeCalled = AtomicBoolean(false)
val messageStarted = AtomicBoolean(false)
@@ -904,13 +943,7 @@ class HermesApiClient(
) {
tracer.done("error")
if (completeCalled.compareAndSet(false, true)) {
val msg = when {
response != null && !response.isSuccessful ->
"API error ${response.code}: ${response.message}"
t is IOException -> "Connection failed: ${t.message}"
t != null -> "Stream error: ${t.message}"
else -> "Unknown stream error"
}
val msg = streamFailureMessage(t, response)
mainHandler.post { onError(msg) }
}
}
@@ -986,7 +1019,13 @@ class HermesApiClient(
model: String? = null,
systemMessage: String? = null,
attachments: List<com.hermesandroid.relay.data.Attachment>? = null,
/** See [sendChatStream]'s `voiceIntentMessages` doc — same semantics. */
/**
* See [sendChatStream]'s `voiceIntentMessages` doc. On the runs
* path the mapping differs slightly: plain user/assistant text
* turns ride the upstream-parsed `conversation_history` field,
* while tool-call pairs fold into the `instructions` digest —
* see [buildRunStreamPayload].
*/
voiceIntentMessages: JsonArray? = null,
onSessionId: (String) -> Unit,
onMessageStarted: (String) -> Unit,
@@ -1008,7 +1047,7 @@ class HermesApiClient(
AgentDisplay.profileRequestName(profileName)?.let {
Log.d(TAG, "sendRunStream: profile=$it")
}
val requestPayload = buildRunStreamPayload(
val built = buildRunStreamPayload(
message = message,
model = model,
systemMessage = systemMessage,
@@ -1017,12 +1056,18 @@ class HermesApiClient(
modelOverride = modelOverride,
profileName = profileName,
)
val requestBody = json.encodeToString(JsonObject.serializer(), requestPayload)
logDroppedAttachments("runs", built.droppedAttachments)
val requestBody = json.encodeToString(JsonObject.serializer(), built.payload)
val request = authRequest("$baseUrl/v1/runs")
.header("Accept", "text/event-stream")
.post(requestBody.toRequestBody(JSON_MEDIA))
.build()
val request = authRequestOrNull("$baseUrl/v1/runs")
?.header("Accept", "text/event-stream")
?.post(requestBody.toRequestBody(JSON_MEDIA))
?.build()
?: run {
// #131: malformed base URL — see sendChatStream.
mainHandler.post { onError(invalidBaseUrlMessage()) }
return failedEventSource()
}
val completeCalled = AtomicBoolean(false)
// Comparable to the gateway's turn[gateway] line — see TurnLatencyTracer.
@@ -1203,13 +1248,7 @@ class HermesApiClient(
) {
tracer.done("error")
if (completeCalled.compareAndSet(false, true)) {
val msg = when {
response != null && !response.isSuccessful ->
"API error ${response.code}: ${response.message}"
t is IOException -> "Connection failed: ${t.message}"
t != null -> "Stream error: ${t.message}"
else -> "Unknown stream error"
}
val msg = streamFailureMessage(t, response)
mainHandler.post { onError(msg) }
}
}
@@ -1364,6 +1403,65 @@ class HermesApiClient(
return builder
}
/**
* Non-throwing twin of [authRequest] for the streaming entry points
* (#131 crash class). The three send*Stream methods build their Request
* BEFORE any try/catch or EventSource listener exists, so a malformed
* [baseUrl] (hand-edited connection, corrupt settings import) made
* `Request.Builder.url(String)` throw `IllegalArgumentException`
* synchronously up through the ViewModel. Returns null on a bad URL so
* the caller can route the failure through its normal `onError` channel
* instead. Non-streaming methods keep [authRequest] — their existing
* try/catch already contains the throw.
*/
private fun authRequestOrNull(url: String): Request.Builder? {
val builder = buildApiRequestOrNull(url) ?: return null
if (apiKey.isNotBlank()) {
builder.header("Authorization", "Bearer $apiKey")
}
return builder
}
/**
* Inert [EventSource] returned by the streaming methods when the request
* couldn't even be built (bad base URL). The turn already failed via
* `onError`; this just satisfies the return type so callers' cancel()
* handling stays uniform.
*/
private fun failedEventSource(): EventSource = object : EventSource {
// Guaranteed-parseable placeholder; never dispatched.
private val placeholder = Request.Builder().url("http://invalid.invalid/").build()
override fun request(): Request = placeholder
override fun cancel() {}
}
/** Human message for a base URL that fails to parse (#131). */
private fun invalidBaseUrlMessage(): String =
"Invalid server address ($baseUrl) — edit the connection's API URL or re-pair."
/**
* Make attachment drops on the SSE fallback transports explicit
* (HRUI-001): the payload builders return attachments that have no
* upstream-supported channel on the target endpoint instead of
* silently omitting them. The user-visible notice lives in
* ChatViewModel (`warnIfAttachmentsDropped`) — this log line is the
* network-layer audit trail that the bytes never left the device.
*/
private fun logDroppedAttachments(
endpoint: String,
dropped: List<com.hermesandroid.relay.data.Attachment>,
) {
if (dropped.isEmpty()) return
val names = dropped.joinToString(", ") {
it.fileName ?: if (it.isImage) "image" else "file"
}
Log.w(
TAG,
"Dropped ${dropped.size} attachment(s) with no supported channel " +
"on the $endpoint endpoint (not sent): $names",
)
}
private fun apiFailure(response: Response, operation: String): IOException {
val detail = response.message.takeIf { it.isNotBlank() }?.let { ": $it" }.orEmpty()
val message = when (response.code) {
@@ -1380,3 +1478,13 @@ class HermesApiClient(
private fun firstNonBlank(vararg values: String?): String =
values.firstOrNull { !it.isNullOrBlank() }.orEmpty()
}
/**
* #131 guard, api_server half: parse-or-null Request builder for a URL string.
* `Request.Builder.url(String)` throws `IllegalArgumentException` on a
* malformed host; the streaming send paths must fail through `onError`
* instead. Top-level (like `buildRelayRequestOrNull` in ConnectionManager)
* so the guard is unit-testable without instantiating the client.
*/
internal fun buildApiRequestOrNull(url: String): Request.Builder? =
url.toHttpUrlOrNull()?.let { Request.Builder().url(it) }
@@ -4,14 +4,233 @@ import com.hermesandroid.relay.data.AgentDisplay
import com.hermesandroid.relay.data.Attachment
import kotlinx.serialization.json.JsonArray
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.add
import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.addJsonObject
import kotlinx.serialization.json.buildJsonArray
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.contentOrNull
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.putJsonObject
/*
* === Upstream request contract (HRUI-001) ===
*
* Verified against hermes-agent `gateway/platforms/api_server.py`. These
* builders send ONLY fields the target handler consumes (plus a small,
* documented set of legacy hint fields — see below). Fields upstream
* ignores are never emitted: a dead field on the wire misrepresents
* capability and masks data loss.
*
* Per-endpoint parsing truth (current upstream main):
*
* - `POST /api/sessions/{id}/chat/stream` (`_handle_session_chat_stream`)
* consumes `message` (or `input`) and `system_message` (or
* `instructions`, string only). `message` accepts either a plain string
* or OpenAI-style content parts (text + `image_url`) via
* `_normalize_multimodal_content`. Top-level `messages`, `attachments`,
* `model`, and `profile` are NOT parsed.
*
* - `POST /v1/runs` (`_handle_runs`) consumes `input` (string or message
* array), `instructions`, `conversation_history` (array of
* `{role, content}` objects, string-coerced), `previous_response_id`,
* `session_id`, and `model`. It does NOT parse `system_message`,
* `stream`, `messages`, `attachments`, or `profile` — and always
* answers `202 {"run_id": ...}` JSON (no SSE on POST).
*
* - `POST /v1/chat/completions` (`_handle_chat_completions`) consumes
* `messages`, `stream`, and `model`. Within `messages`: `system` roles
* fold into the ephemeral system prompt; `user`/`assistant` entries are
* kept as history with multimodal content normalization; `tool`-role
* entries are silently skipped and `tool_calls` fields are stripped.
* Top-level `attachments` and `profile` are NOT parsed.
*
* Legacy hint fields we deliberately keep sending although current native
* upstream ignores them: `model` + `profile` on the sessions path,
* `profile` on runs/completions, and `stream` on runs. They are
* configuration hints (never user content, so they cannot mask data
* loss) honored by legacy fork builds — the runs path in particular only
* activates against servers that explicitly advertise SSE-on-POST, which
* vanilla upstream never does. See `ServerCapabilities`.
*
* === Synthetic-history mapping ===
*
* Phone-local synthetic turns (voice-intent traces, card dispatches,
* provider-answered realtime voice turns — see `VoiceIntentSyncBuilder`,
* `CardDispatchSyncBuilder`, `RealtimeTurnSyncBuilder`) arrive here as one
* OpenAI-format array. Historically they were sent as a top-level
* `messages` field on sessions/runs, which upstream never consumed —
* silent data loss. They now map onto channels each endpoint actually
* supports:
*
* - Tool-call pairs (`assistant` + `tool` with `tool_call_id`) have no
* surviving wire shape on ANY fallback endpoint, so they render as a
* plain-text digest ([renderSyntheticHistoryDigest]) folded into the
* per-turn ephemeral system prompt: `system_message` on sessions,
* `instructions` on runs, the `system` message on completions.
* - Plain `user`/`assistant` text turns ride a real history channel
* where one exists: spliced into `messages` on completions, sent as
* `conversation_history` on runs. The sessions endpoint has no
* client-provided history channel, so there they join the digest.
*
* This mapping is ephemeral where the digest is used: the model sees the
* context for THIS turn only; it is not persisted into the server-side
* session transcript. That is strictly better than the previous behavior
* (context arrived never) and matches the existing voice-turn pattern of
* per-turn non-persisted instructions.
*
* === Attachments ===
*
* Only the completions endpoint has an upstream-supported attachment
* channel on this surface: inline `image_url` content parts (images
* only). Sessions/runs payloads carry no attachments at all. Anything
* that cannot be delivered is returned in
* [ChatPayloadResult.droppedAttachments] so callers can surface the drop
* (HermesApiClient logs it; ChatViewModel shows a user-visible notice) —
* never a silent discard. Note: current upstream's sessions `message`
* field does accept inline `image_url` content parts, so image delivery
* on the sessions path is a possible future improvement; it is not wired
* yet because the caller's attachment warning and this builder must move
* together.
*/
/**
* Result of building a fallback-transport chat payload.
*
* @property payload The JSON request body — contains only fields the
* target endpoint consumes (plus documented legacy hint fields).
* @property droppedAttachments Attachments that have NO supported channel
* on the target endpoint and were therefore not encoded into [payload].
* Callers must surface these (log + user notice), never ignore them.
*/
internal data class ChatPayloadResult(
val payload: JsonObject,
val droppedAttachments: List<Attachment>,
)
/**
* Header line for the synthetic phone-context digest. Tells the model the
* listed activity already happened on-device so it treats the lines as
* history, not instructions to act on.
*/
internal const val SYNTHETIC_DIGEST_HEADER =
"Phone-side activity since the previous server turn " +
"(already completed on-device; context only — do not re-execute):"
private fun JsonObject.roleOrNull(): String? =
(this["role"] as? JsonPrimitive)?.contentOrNull
private fun JsonObject.contentStringOrNull(): String? =
(this["content"] as? JsonPrimitive)?.contentOrNull
/**
* True for a synthetic entry deliverable as a REAL conversation turn on
* endpoints with a client-history channel: plain `user`/`assistant` role,
* string content, no `tool_calls`. Matches the shape emitted by
* `RealtimeTurnSyncBuilder`; tool-call pairs from the voice-intent and
* card-dispatch builders fail this check and go through the digest.
*/
internal fun isPlainSyntheticTurn(entry: JsonObject): Boolean {
val role = entry.roleOrNull()
if (role != "user" && role != "assistant") return false
if (entry.containsKey("tool_calls")) return false
return !entry.contentStringOrNull().isNullOrBlank()
}
/**
* Render the synthetic sync stream as a compact plain-text digest for the
* per-turn ephemeral system prompt.
*
* Tool-call pairs (`assistant.tool_calls` + matching `tool` result keyed
* by `tool_call_id`) always render, one line per call:
* `- called <name> with <arguments> -> <result>`. Plain text turns render
* as `- user: ...` / `- assistant: ...` lines only when
* [includePlainTurns] is true (sessions path — no real history channel);
* endpoints that deliver plain turns natively pass false so the same turn
* is never delivered twice.
*
* @return null when nothing renders (no synthetic messages, or only plain
* turns while [includePlainTurns] is false).
*/
internal fun renderSyntheticHistoryDigest(
syntheticMessages: JsonArray?,
includePlainTurns: Boolean,
): String? {
if (syntheticMessages.isNullOrEmpty()) return null
// Pair tool results with their originating call.
val resultsByCallId = HashMap<String, String>()
for (element in syntheticMessages) {
val obj = element as? JsonObject ?: continue
if (obj.roleOrNull() != "tool") continue
val callId = (obj["tool_call_id"] as? JsonPrimitive)?.contentOrNull ?: continue
resultsByCallId[callId] = obj.contentStringOrNull().orEmpty()
}
val lines = mutableListOf<String>()
for (element in syntheticMessages) {
val obj = element as? JsonObject ?: continue
when (obj.roleOrNull()) {
"assistant" -> {
val toolCalls = obj["tool_calls"] as? JsonArray
if (toolCalls != null) {
for (call in toolCalls) {
val callObj = call as? JsonObject ?: continue
val function = callObj["function"] as? JsonObject
val name = (function?.get("name") as? JsonPrimitive)
?.contentOrNull ?: "unknown_tool"
val args = (function?.get("arguments") as? JsonPrimitive)
?.contentOrNull ?: "{}"
val callId = (callObj["id"] as? JsonPrimitive)?.contentOrNull
val result = callId?.let(resultsByCallId::get)
lines += if (result.isNullOrBlank()) {
"- called $name with $args"
} else {
"- called $name with $args -> $result"
}
}
} else if (includePlainTurns) {
obj.contentStringOrNull()?.takeIf { it.isNotBlank() }
?.let { lines += "- assistant: $it" }
}
}
"user" -> if (includePlainTurns) {
obj.contentStringOrNull()?.takeIf { it.isNotBlank() }
?.let { lines += "- user: $it" }
}
// "tool" entries fold into their assistant line via resultsByCallId.
}
}
if (lines.isEmpty()) return null
return SYNTHETIC_DIGEST_HEADER + "\n" + lines.joinToString("\n")
}
/**
* Merge the caller's per-turn system message with the synthetic-history
* digest into one ephemeral prompt string. Either side may be absent.
*/
internal fun mergeEphemeralContext(systemMessage: String?, digest: String?): String? = when {
digest.isNullOrBlank() -> systemMessage?.takeIf { it.isNotBlank() }
systemMessage.isNullOrBlank() -> digest
else -> systemMessage + "\n\n" + digest
}
/** Synthetic entries deliverable as real history turns (see [isPlainSyntheticTurn]). */
private fun plainSyntheticTurns(syntheticMessages: JsonArray?): List<JsonObject> =
(syntheticMessages ?: emptyList())
.mapNotNull { it as? JsonObject }
.filter(::isPlainSyntheticTurn)
/**
* Body for `POST /api/sessions/{id}/chat/stream`.
*
* Emits `message` + `system_message` (upstream-consumed) and `model` +
* `profile` (legacy hints — current native upstream ignores both on this
* route; legacy fork builds honor them; see the file header). ALL
* synthetic history folds into `system_message` via the digest: the
* endpoint has no client-provided history channel. Attachments have no
* supported channel here and are returned as dropped.
*/
internal fun buildSessionChatStreamPayload(
message: String,
systemMessage: String? = null,
@@ -19,30 +238,33 @@ internal fun buildSessionChatStreamPayload(
voiceIntentMessages: JsonArray? = null,
modelOverride: String? = null,
profileName: String? = null,
): JsonObject = buildJsonObject {
put("message", message)
if (!systemMessage.isNullOrBlank()) {
put("system_message", systemMessage)
}
if (!modelOverride.isNullOrBlank()) {
put("model", modelOverride)
}
AgentDisplay.profileRequestName(profileName)?.let { put("profile", it) }
if (!attachments.isNullOrEmpty()) {
putJsonArray("attachments") {
attachments.forEach { att ->
addJsonObject {
put("contentType", att.contentType)
put("content", att.content)
}
}
): ChatPayloadResult {
val digest = renderSyntheticHistoryDigest(voiceIntentMessages, includePlainTurns = true)
val effectiveSystem = mergeEphemeralContext(systemMessage, digest)
val payload = buildJsonObject {
put("message", message)
if (!effectiveSystem.isNullOrBlank()) {
put("system_message", effectiveSystem)
}
if (!modelOverride.isNullOrBlank()) {
put("model", modelOverride)
}
AgentDisplay.profileRequestName(profileName)?.let { put("profile", it) }
}
if (voiceIntentMessages != null && voiceIntentMessages.isNotEmpty()) {
put("messages", voiceIntentMessages)
}
return ChatPayloadResult(payload, droppedAttachments = attachments.orEmpty())
}
/**
* Body for `POST /v1/runs`.
*
* Emits `input`, `model`, and `instructions` (upstream-consumed; note the
* runs handler reads `instructions`, NOT `system_message` — the latter was
* a silent drop before HRUI-001), plus `stream` + `profile` legacy hints.
* Synthetic history: plain text turns ride `conversation_history` (a real
* upstream channel — entries are `{role, content}` objects); tool-call
* pairs fold into the `instructions` digest. Attachments have no
* supported channel here and are returned as dropped.
*/
internal fun buildRunStreamPayload(
message: String,
model: String? = null,
@@ -51,36 +273,46 @@ internal fun buildRunStreamPayload(
voiceIntentMessages: JsonArray? = null,
modelOverride: String? = null,
profileName: String? = null,
): JsonObject {
): ChatPayloadResult {
val resolvedModel = when {
!modelOverride.isNullOrBlank() -> modelOverride
!model.isNullOrBlank() -> model
else -> "default"
}
return buildJsonObject {
val digest = renderSyntheticHistoryDigest(voiceIntentMessages, includePlainTurns = false)
val effectiveInstructions = mergeEphemeralContext(systemMessage, digest)
val plainTurns = plainSyntheticTurns(voiceIntentMessages)
val payload = buildJsonObject {
put("model", resolvedModel)
put("input", message)
put("stream", true)
if (!systemMessage.isNullOrBlank()) {
put("system_message", systemMessage)
if (!effectiveInstructions.isNullOrBlank()) {
put("instructions", effectiveInstructions)
}
AgentDisplay.profileRequestName(profileName)?.let { put("profile", it) }
if (!attachments.isNullOrEmpty()) {
putJsonArray("attachments") {
attachments.forEach { att ->
addJsonObject {
put("contentType", att.contentType)
put("content", att.content)
}
}
if (plainTurns.isNotEmpty()) {
putJsonArray("conversation_history") {
plainTurns.forEach { add(it) }
}
}
if (voiceIntentMessages != null && voiceIntentMessages.isNotEmpty()) {
put("messages", voiceIntentMessages)
}
AgentDisplay.profileRequestName(profileName)?.let { put("profile", it) }
}
return ChatPayloadResult(payload, droppedAttachments = attachments.orEmpty())
}
/**
* Body for `POST /v1/chat/completions`.
*
* Emits `model`, `stream`, and `messages` (all upstream-consumed) plus
* the `profile` legacy hint. Synthetic history: plain text turns splice
* into `messages` before the live user message (upstream keeps
* `user`/`assistant` history entries verbatim); tool-call pairs fold into
* the system message digest, because upstream SKIPS `tool`-role messages
* and STRIPS `tool_calls` — splicing them produced junk empty-content
* assistant entries and lost the results entirely. Image attachments ride
* inline `image_url` content parts on the user message (upstream vision
* format); non-image attachments have no channel and are returned as
* dropped.
*/
internal fun buildChatCompletionsStreamPayload(
message: String,
model: String? = null,
@@ -89,35 +321,37 @@ internal fun buildChatCompletionsStreamPayload(
voiceIntentMessages: JsonArray? = null,
modelOverride: String? = null,
profileName: String? = null,
): JsonObject {
): ChatPayloadResult {
val resolvedModel = when {
!modelOverride.isNullOrBlank() -> modelOverride
!model.isNullOrBlank() -> model
else -> "default"
}
return buildJsonObject {
val digest = renderSyntheticHistoryDigest(voiceIntentMessages, includePlainTurns = false)
val effectiveSystem = mergeEphemeralContext(systemMessage, digest)
val plainTurns = plainSyntheticTurns(voiceIntentMessages)
val imageAttachments = attachments.orEmpty().filter { it.isImage }
val payload = buildJsonObject {
put("model", resolvedModel)
put("stream", true)
AgentDisplay.profileRequestName(profileName)?.let { put("profile", it) }
putJsonArray("messages") {
if (!systemMessage.isNullOrBlank()) {
if (!effectiveSystem.isNullOrBlank()) {
addJsonObject {
put("role", "system")
put("content", systemMessage)
put("content", effectiveSystem)
}
}
if (voiceIntentMessages != null && voiceIntentMessages.isNotEmpty()) {
voiceIntentMessages.forEach { add(it) }
}
plainTurns.forEach { add(it) }
addJsonObject {
put("role", "user")
if (!attachments.isNullOrEmpty() && attachments.any { it.isImage }) {
if (imageAttachments.isNotEmpty()) {
put("content", buildJsonArray {
addJsonObject {
put("type", "text")
put("text", message)
}
attachments.filter { it.isImage }.forEach { att ->
imageAttachments.forEach { att ->
addJsonObject {
put("type", "image_url")
putJsonObject("image_url") {
@@ -131,15 +365,9 @@ internal fun buildChatCompletionsStreamPayload(
}
}
}
if (!attachments.isNullOrEmpty() && attachments.any { !it.isImage }) {
putJsonArray("attachments") {
attachments.filter { !it.isImage }.forEach { att ->
addJsonObject {
put("contentType", att.contentType)
put("content", att.content)
}
}
}
}
}
return ChatPayloadResult(
payload = payload,
droppedAttachments = attachments.orEmpty().filter { !it.isImage },
)
}
@@ -179,6 +179,58 @@ data class RenameSessionRequest(
val title: String
)
// --- Server-backed bulk cleanup (dashboard POST /api/sessions/prune) ---
/**
* Client-side subset of upstream's `SessionPrune` body. Nulls are omitted from
* the request; a fully-bare filter set is a "bare prune", where upstream
* applies its own implicit ended-more-than-90-days-ago cutoff.
*/
data class SessionPruneFilters(
val olderThanDays: Double? = null,
val source: String? = null,
val profile: String? = null,
val includeArchived: Boolean = false,
)
/** One row of the dry-run preview (`sessions` in the prune response). */
@Serializable
data class SessionPruneCandidate(
@Serializable(with = FlexibleIdNonNullSerializer::class)
val id: String = "",
val source: String? = null,
val title: String? = null,
val model: String? = null,
@SerialName("started_at")
@Serializable(with = FlexibleTimestampSerializer::class)
val startedAt: Double? = null,
@SerialName("message_count") val messageCount: Int? = null,
)
/**
* Dry-run response: what a prune WOULD delete — count, started-at span, and
* the candidate rows — without deleting anything. Upstream orders candidates
* oldest-first.
*/
@Serializable
data class SessionPrunePreview(
val matched: Int = 0,
@SerialName("oldest_started_at")
@Serializable(with = FlexibleTimestampSerializer::class)
val oldestStartedAt: Double? = null,
@SerialName("newest_started_at")
@Serializable(with = FlexibleTimestampSerializer::class)
val newestStartedAt: Double? = null,
val sessions: List<SessionPruneCandidate> = emptyList(),
)
/** Apply response — how many sessions the server actually removed. */
@Serializable
data class SessionPruneResult(
val ok: Boolean = true,
val removed: Int = 0,
)
// --- Messages ---
@Serializable
@@ -8,6 +8,11 @@ import android.service.notification.StatusBarNotification
import android.util.Log
import com.hermesandroid.relay.network.relay.ChannelMultiplexer
import com.hermesandroid.relay.network.relay.models.Envelope
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.cancel
import kotlinx.coroutines.launch
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
@@ -42,6 +47,11 @@ import java.util.concurrent.ConcurrentLinkedQueue
*/
class HermesNotificationCompanion : NotificationListenerService() {
private val serviceScope = CoroutineScope(SupervisorJob() + Dispatchers.IO)
private val triggerStore by lazy {
NotificationTriggerStore(applicationContext.notificationTriggerDataStore)
}
/**
* Buffer for entries that arrive before [multiplexer] has been
* wired up (e.g. notifications during app cold-start). Bounded so
@@ -68,6 +78,7 @@ class HermesNotificationCompanion : NotificationListenerService() {
if (active === this) {
active = null
}
serviceScope.cancel()
super.onDestroy()
}
@@ -75,6 +86,13 @@ class HermesNotificationCompanion : NotificationListenerService() {
if (sbn == null) return
val entry = sbn.toEntry() ?: return
// The trigger MVP posts its own local prompt notifications. Never feed
// Hermes-Relay's notifications back into the rule engine, or a broad
// rule could prompt on its own prompt. Still forward them to the relay
// cache to preserve existing notification-companion semantics.
if (entry.packageName != packageName) {
evaluateNotificationTriggers(entry)
}
val envelope = entry.toEnvelope()
// Drain any backlog first so order is preserved.
@@ -141,6 +159,31 @@ class HermesNotificationCompanion : NotificationListenerService() {
)
}
private fun evaluateNotificationTriggers(entry: NotificationEntry) {
serviceScope.launch {
val match = triggerStore.firstMatchingRule(entry) ?: return@launch
val result = when (match.rule.action) {
NotificationTriggerAction.AskMe -> NotificationTriggerPromptNotifier.notifyAskMe(
context = applicationContext,
rule = match.rule,
entry = entry,
)
}
triggerStore.appendActivity(
NotificationTriggerActivityEntry(
ruleId = match.rule.id,
ruleLabel = match.rule.label,
action = match.rule.action,
packageName = entry.packageName,
title = entry.title,
textPreview = entry.text?.take(160) ?: entry.subText?.take(160),
matchedAt = System.currentTimeMillis(),
result = result,
)
)
}
}
private fun NotificationEntry.toEnvelope(): Envelope {
val payload = JSON.encodeToJsonElement(NotificationEntry.serializer(), this) as JsonObject
return Envelope(
@@ -0,0 +1,292 @@
package com.hermesandroid.relay.notifications
import android.Manifest
import android.annotation.SuppressLint
import android.app.NotificationChannel
import android.app.NotificationManager
import android.app.PendingIntent
import android.content.Context
import android.content.Intent
import android.content.pm.PackageManager
import android.os.Build
import android.util.Log
import androidx.core.app.NotificationCompat
import androidx.core.app.NotificationManagerCompat
import androidx.core.content.ContextCompat
import androidx.datastore.core.DataStore
import androidx.datastore.preferences.core.Preferences
import androidx.datastore.preferences.core.booleanPreferencesKey
import androidx.datastore.preferences.core.edit
import androidx.datastore.preferences.core.stringPreferencesKey
import androidx.datastore.preferences.preferencesDataStore
import com.hermesandroid.relay.MainActivity
import com.hermesandroid.relay.R
import kotlinx.coroutines.flow.Flow
import kotlinx.coroutines.flow.first
import kotlinx.coroutines.flow.map
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.decodeFromString
import kotlinx.serialization.encodeToString
import kotlinx.serialization.json.Json
import java.util.UUID
/**
* Minimal notification-trigger MVP schema and persistence.
*
* Storage location: Android DataStore preferences file `notification_triggers`
* under the app-private data directory. Rules and the visible activity log are
* JSON strings so schema evolution remains additive and lenient.
*/
@Serializable
data class NotificationTriggerRule(
val id: String = UUID.randomUUID().toString(),
val label: String = "Ask me about matching notifications",
val enabled: Boolean = true,
@SerialName("app_package")
val appPackage: String? = null,
@SerialName("title_contains")
val titleContains: String? = null,
@SerialName("text_contains")
val textContains: String? = null,
val action: NotificationTriggerAction = NotificationTriggerAction.AskMe,
@SerialName("require_confirmation")
val requireConfirmation: Boolean = false,
)
@Serializable
enum class NotificationTriggerAction {
@SerialName("ask_me")
AskMe,
}
@Serializable
data class NotificationTriggerActivityEntry(
val id: String = UUID.randomUUID().toString(),
@SerialName("rule_id")
val ruleId: String,
@SerialName("rule_label")
val ruleLabel: String,
val action: NotificationTriggerAction,
@SerialName("package_name")
val packageName: String,
val title: String? = null,
@SerialName("text_preview")
val textPreview: String? = null,
@SerialName("matched_at")
val matchedAt: Long,
val result: String,
)
@Serializable
data class NotificationTriggerSettings(
@SerialName("master_enabled")
val masterEnabled: Boolean = false,
@SerialName("kill_switch")
val killSwitch: Boolean = false,
val rules: List<NotificationTriggerRule> = emptyList(),
@SerialName("activity_log")
val activityLog: List<NotificationTriggerActivityEntry> = emptyList(),
)
data class NotificationTriggerMatch(
val rule: NotificationTriggerRule,
val entry: NotificationEntry,
)
internal val Context.notificationTriggerDataStore: DataStore<Preferences> by
preferencesDataStore(name = "notification_triggers")
class NotificationTriggerStore(
private val dataStore: DataStore<Preferences>,
) {
private val json = Json {
ignoreUnknownKeys = true
encodeDefaults = true
}
val settings: Flow<NotificationTriggerSettings> = dataStore.data.map { prefs ->
NotificationTriggerSettings(
masterEnabled = prefs[KEY_MASTER_ENABLED] ?: false,
killSwitch = prefs[KEY_KILL_SWITCH] ?: false,
rules = decodeList<NotificationTriggerRule>(prefs[KEY_RULES_JSON]),
activityLog = decodeList<NotificationTriggerActivityEntry>(prefs[KEY_ACTIVITY_LOG_JSON]),
)
}
suspend fun setMasterEnabled(enabled: Boolean) {
dataStore.edit { prefs -> prefs[KEY_MASTER_ENABLED] = enabled }
}
suspend fun setKillSwitch(enabled: Boolean) {
dataStore.edit { prefs -> prefs[KEY_KILL_SWITCH] = enabled }
}
suspend fun saveSingleRule(rule: NotificationTriggerRule) {
dataStore.edit { prefs ->
prefs[KEY_RULES_JSON] = json.encodeToString(listOf(rule.normalized()))
}
}
suspend fun clearActivityLog() {
dataStore.edit { prefs -> prefs.remove(KEY_ACTIVITY_LOG_JSON) }
}
suspend fun firstMatchingRule(entry: NotificationEntry): NotificationTriggerMatch? {
val snapshot = settings.first()
if (!snapshot.masterEnabled || snapshot.killSwitch) return null
val rule = snapshot.rules.firstOrNull { it.matches(entry) } ?: return null
return NotificationTriggerMatch(rule = rule, entry = entry)
}
suspend fun appendActivity(entry: NotificationTriggerActivityEntry) {
dataStore.edit { prefs ->
val current = decodeList<NotificationTriggerActivityEntry>(prefs[KEY_ACTIVITY_LOG_JSON])
prefs[KEY_ACTIVITY_LOG_JSON] = json.encodeToString(
(listOf(entry) + current).take(MAX_ACTIVITY_LOG_ENTRIES),
)
}
}
private inline fun <reified T> decodeList(raw: String?): List<T> {
if (raw.isNullOrBlank()) return emptyList()
return runCatching { json.decodeFromString<List<T>>(raw) }.getOrDefault(emptyList())
}
private fun NotificationTriggerRule.normalized(): NotificationTriggerRule = copy(
label = label.trim().ifBlank { "Ask me about matching notifications" },
appPackage = appPackage.cleanBlank(),
titleContains = titleContains.cleanBlank(),
textContains = textContains.cleanBlank(),
)
companion object {
private val KEY_MASTER_ENABLED = booleanPreferencesKey("notification_triggers_enabled")
private val KEY_KILL_SWITCH = booleanPreferencesKey("notification_triggers_kill_switch")
private val KEY_RULES_JSON = stringPreferencesKey("notification_trigger_rules_json")
private val KEY_ACTIVITY_LOG_JSON = stringPreferencesKey("notification_trigger_activity_log_json")
const val MAX_ACTIVITY_LOG_ENTRIES = 25
fun defaultRule(): NotificationTriggerRule = NotificationTriggerRule()
}
}
fun NotificationTriggerRule.matches(entry: NotificationEntry): Boolean {
if (!enabled) return false
val app = appPackage.cleanBlank()
val titleNeedle = titleContains.cleanBlank()
val textNeedle = textContains.cleanBlank()
// Avoid accidental "match every notification on the phone" rules. The UI
// requires at least one filter too, but this keeps imported/future schema
// data safe.
if (app == null && titleNeedle == null && textNeedle == null) return false
if (app != null && !entry.packageName.equals(app, ignoreCase = true)) return false
if (titleNeedle != null && !entry.title.orEmpty().contains(titleNeedle, ignoreCase = true)) {
return false
}
if (textNeedle != null) {
val haystack = listOfNotNull(entry.text, entry.subText).joinToString("\n")
if (!haystack.contains(textNeedle, ignoreCase = true)) return false
}
return true
}
fun NotificationTriggerRule.summary(): String {
val parts = buildList {
appPackage.cleanBlank()?.let { add("app $it") }
titleContains.cleanBlank()?.let { add("title contains “$it”") }
textContains.cleanBlank()?.let { add("text contains “$it”") }
}
return if (parts.isEmpty()) "No filters set" else parts.joinToString(" · ")
}
private fun String?.cleanBlank(): String? = this?.trim()?.takeIf { it.isNotBlank() }
object NotificationTriggerPromptNotifier {
private const val TAG = "NotifTriggerPrompt"
private const val CHANNEL_ID = "notification_triggers"
private const val CHANNEL_NAME = "Notification triggers"
private const val NOTIFICATION_ID_BASE = 4300
private const val CHAT_ROUTE = "chat"
/**
* Safe automatic action: post a local prompt that asks the user whether to
* involve Hermes. It does not send an LLM request, reply, tap, text, route,
* or otherwise act on another app without the user tapping first.
*/
@SuppressLint("MissingPermission", "NotificationPermission")
fun notifyAskMe(
context: Context,
rule: NotificationTriggerRule,
entry: NotificationEntry,
): String {
ensureChannel(context)
if (!hasPostNotificationsPermission(context)) {
Log.i(TAG, "POST_NOTIFICATIONS not granted — logging trigger without prompt")
return "skipped: post-notifications permission missing"
}
val tapIntent = Intent(context, MainActivity::class.java).apply {
flags = Intent.FLAG_ACTIVITY_NEW_TASK or Intent.FLAG_ACTIVITY_CLEAR_TOP
putExtra(MainActivity.EXTRA_NAV_ROUTE, CHAT_ROUTE)
}
val pendingFlags = PendingIntent.FLAG_UPDATE_CURRENT or PendingIntent.FLAG_IMMUTABLE
val tapPending = PendingIntent.getActivity(context, notificationId(entry), tapIntent, pendingFlags)
val title = "Ask Hermes about this?"
val source = entry.title?.takeIf { it.isNotBlank() } ?: entry.packageName
val body = entry.text?.takeIf { it.isNotBlank() }
?: "Rule matched: ${rule.summary()}"
val expanded = "Matched ${rule.summary()}\n\n$source\n$body"
val notification = NotificationCompat.Builder(context, CHANNEL_ID)
.setSmallIcon(R.mipmap.ic_launcher)
.setContentTitle(title)
.setContentText("$source — ${body.take(96)}")
.setStyle(NotificationCompat.BigTextStyle().bigText(expanded.take(700)))
.setContentIntent(tapPending)
.setAutoCancel(true)
.setOnlyAlertOnce(false)
.setPriority(NotificationCompat.PRIORITY_DEFAULT)
.setCategory(NotificationCompat.CATEGORY_REMINDER)
.build()
return runCatching {
NotificationManagerCompat.from(context).notify(notificationId(entry), notification)
"prompt posted"
}.getOrElse { exc ->
Log.w(TAG, "notifyAskMe: notify failed", exc)
"skipped: prompt failed (${exc.javaClass.simpleName})"
}
}
private fun notificationId(entry: NotificationEntry): Int {
val suffix = (entry.key.hashCode() and 0x0fff)
return NOTIFICATION_ID_BASE + suffix
}
private fun ensureChannel(context: Context) {
if (Build.VERSION.SDK_INT < Build.VERSION_CODES.O) return
val nm = context.getSystemService(NotificationManager::class.java) ?: return
if (nm.getNotificationChannel(CHANNEL_ID) != null) return
val channel = NotificationChannel(
CHANNEL_ID,
CHANNEL_NAME,
NotificationManager.IMPORTANCE_DEFAULT,
).apply {
description = "Prompts shown when an explicitly enabled notification trigger matches."
setShowBadge(true)
}
nm.createNotificationChannel(channel)
}
private fun hasPostNotificationsPermission(context: Context): Boolean {
if (Build.VERSION.SDK_INT < Build.VERSION_CODES.TIRAMISU) return true
return ContextCompat.checkSelfPermission(
context,
Manifest.permission.POST_NOTIFICATIONS,
) == PackageManager.PERMISSION_GRANTED
}
}
@@ -421,6 +421,7 @@ fun RelayApp() {
System.currentTimeMillis() - lastPausedAtMs.value
}
connectionViewModel.revalidateOnResume(awayMs)
voiceViewModel.onAppResumed()
}
else -> {}
}
@@ -791,6 +792,23 @@ fun RelayApp() {
chatViewModel.notifyOnTurnComplete = notifyTurnComplete
}
// Demo-mode composer wiring: unconditional — a demo session has no API
// client, so the client-gated chat init effect above never runs and
// ChatViewModel's own handler stays null. Lambdas read live state on
// every send.
LaunchedEffect(Unit) {
chatViewModel.setDemoModeWiring(
isDemo = { connectionViewModel.isDemoMode.value },
handler = { connectionViewModel.chatHandler },
)
// Voice → chat breadcrumbs (e.g. "background task still running" when
// voice mode exits with a detached run) land as system notices in the
// shared chat transcript.
voiceViewModel.chatNoticeSink = { notice ->
connectionViewModel.chatHandler.addSystemNotice(notice)
}
}
// Sync tool annotation parsing toggle to ChatHandler
val parseAnnotations by connectionViewModel.parseToolAnnotations.collectAsState()
LaunchedEffect(parseAnnotations) {
@@ -21,11 +21,15 @@ import androidx.compose.foundation.verticalScroll
import androidx.compose.material3.Button
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.OutlinedButton
import androidx.compose.material3.OutlinedTextField
import androidx.compose.material3.Surface
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.runtime.Composable
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.platform.LocalContext
@@ -38,6 +42,7 @@ import androidx.compose.ui.window.DialogProperties
import com.hermesandroid.relay.BuildConfig
import com.hermesandroid.relay.diagnostics.DiagnosticLogEntry
import com.hermesandroid.relay.diagnostics.DiagnosticSeverity
import com.hermesandroid.relay.util.DiagnosticIssuePrefill
import com.hermesandroid.relay.util.IssueReport
/**
@@ -57,6 +62,13 @@ fun DiagnosticDetailDialog(entry: DiagnosticLogEntry, onDismiss: () -> Unit) {
val plainText = remember(entry) { entry.toPlainText() }
val severityName = entry.severity.name
// Info-severity pre-flight: routine log lines only become GitHub issues once
// the reporter says what they expected instead (that answer replaces the
// boilerplate "What happened" line). Error entries keep the direct flow.
val needsExpectation = entry.severity == DiagnosticSeverity.Info
var expectationVisible by remember(entry) { mutableStateOf(false) }
var expectation by remember(entry) { mutableStateOf("") }
Dialog(
onDismissRequest = onDismiss,
properties = DialogProperties(usePlatformDefaultWidth = false),
@@ -127,6 +139,20 @@ fun DiagnosticDetailDialog(entry: DiagnosticLogEntry, onDismiss: () -> Unit) {
)
}
if (expectationVisible) {
Spacer(Modifier.height(14.dp))
OutlinedTextField(
value = expectation,
onValueChange = { expectation = it },
label = { Text("What were you expecting to happen?") },
supportingText = {
Text("This is a routine log entry — telling us what looked wrong turns it into an answerable report.")
},
minLines = 2,
modifier = Modifier.fillMaxWidth(),
)
}
Spacer(Modifier.height(18.dp))
FlowRow(
modifier = Modifier.fillMaxWidth(),
@@ -155,16 +181,24 @@ fun DiagnosticDetailDialog(entry: DiagnosticLogEntry, onDismiss: () -> Unit) {
},
) { Text("Export") }
Button(
enabled = !expectationVisible || expectation.isNotBlank(),
onClick = {
if (needsExpectation && !expectationVisible) {
expectationVisible = true
return@Button
}
// Copy full text first; the GitHub URL only carries the
// head of long traces, so the user can paste the rest.
IssueReport.copyToClipboard(context, plainText)
val opened = IssueReport.openUrl(
context,
IssueReport.buildGithubIssueUrl(
title = "[Bug]: ${entry.title}",
bodyMarkdown = entry.toIssueBody(),
labels = "bug",
title = DiagnosticIssuePrefill.issueTitle(entry),
bodyMarkdown = DiagnosticIssuePrefill.issueBody(
entry,
expectation = expectation.takeIf { expectationVisible },
),
labels = DiagnosticIssuePrefill.issueLabels(entry),
),
)
toast(
@@ -245,54 +279,3 @@ private fun DiagnosticLogEntry.toPlainText(): String = buildString {
append(it)
}
}
/**
* Markdown issue body mirroring the crash-report issue format: environment block
* + the captured entry. Trace is capped so the prefilled GitHub URL stays within
* browser limits (full text is on the clipboard).
*/
private const val MAX_TRACE_FOR_URL = 3000
private fun DiagnosticLogEntry.toIssueBody(): String {
val trace = (stacktrace ?: detail).orEmpty().let {
if (it.length > MAX_TRACE_FOR_URL) {
it.take(MAX_TRACE_FOR_URL) + "\n… (truncated — full diagnostic copied to your clipboard)"
} else {
it
}
}
val surface = if (BuildConfig.FLAVOR.equals("sideload", ignoreCase = true)) "sideload APK" else "Google Play"
return buildString {
appendLine(
"> ⚠️ Before submitting: remove any secrets, tokens, real hostnames/IPs, " +
"or personal data from the detail below.",
)
appendLine()
appendLine("### Affected area")
appendLine("Android app")
appendLine()
appendLine("### What happened?")
appendLine("Captured diagnostic from the in-app activity log.")
appendLine()
appendLine("### Environment")
appendLine("- Hermes-Relay version/tag: ${BuildConfig.VERSION_NAME} (code ${BuildConfig.VERSION_CODE})")
appendLine("- Install surface: $surface")
appendLine("- Connection mode: LAN / Tailscale / public TLS / other")
appendLine()
appendLine("### Diagnostic")
appendLine("- Title: $title")
appendLine("- Category: ${category.label}")
appendLine("- Severity: ${severity.name}")
endpointRole?.let { appendLine("- Route: $it") }
url?.let { appendLine("- URL: $it") }
elapsedMs?.let { appendLine("- Elapsed: ${it}ms") }
if (trace.isNotBlank()) {
appendLine()
appendLine("```")
appendLine(trace)
appendLine("```")
}
appendLine()
append("<sub>Captured by the Hermes-Relay in-app diagnostics log</sub>")
}
}
@@ -121,6 +121,14 @@ fun MessageBubble(
* conversation). Null hides the entry.
*/
onEditMessage: ((ChatMessage) -> Unit)? = null,
/**
* True while the ViewModel is recovering a dropped stream's answer by
* polling the session transcript (issue #166) — the streaming
* placeholder's slow-turn label reads "Reconnecting to your answer…"
* instead of "Still working…" so the wait is honest about what's
* happening.
*/
recoveringAnswer: Boolean = false,
) {
val isUser = message.role == MessageRole.USER
val isSystem = message.role == MessageRole.SYSTEM
@@ -499,10 +507,17 @@ fun MessageBubble(
modifier = Modifier.padding(top = 4.dp),
)
}
if (showStillWorking && awaitingFirstToken) {
// During dropped-stream answer recovery the label shows
// immediately (the 4s escalation is for a slow first
// token; a recovery is already known to be slow).
if ((showStillWorking || recoveringAnswer) && awaitingFirstToken) {
Spacer(modifier = Modifier.width(8.dp))
Text(
text = "Still working…",
text = if (recoveringAnswer) {
"Reconnecting to your answer…"
} else {
"Still working…"
},
style = MaterialTheme.typography.labelSmall,
color = textColor.copy(alpha = 0.6f),
modifier = Modifier.padding(top = 4.dp),
@@ -17,7 +17,9 @@ import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.filled.Check
import androidx.compose.material.icons.filled.Close
import androidx.compose.material.icons.filled.Refresh
import androidx.compose.material.icons.filled.Search
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.ExperimentalMaterial3Api
import androidx.compose.material3.HorizontalDivider
import androidx.compose.material3.Icon
@@ -26,6 +28,7 @@ import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.ModalBottomSheet
import androidx.compose.material3.OutlinedTextField
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.material3.rememberModalBottomSheetState
import androidx.compose.runtime.Composable
import androidx.compose.runtime.getValue
@@ -54,6 +57,8 @@ import androidx.compose.ui.unit.dp
@Composable
fun ModelPickerSheet(
options: List<ChatInputPickerOption>,
refreshing: Boolean = false,
onRefresh: (() -> Unit)? = null,
onSelect: (ChatInputPickerOption) -> Unit,
onDismiss: () -> Unit,
) {
@@ -98,12 +103,35 @@ fun ModelPickerSheet(
horizontalArrangement = Arrangement.SpaceBetween,
verticalAlignment = Alignment.CenterVertically,
) {
Text(text = "Model", style = MaterialTheme.typography.titleMedium)
Text(
text = "${modelOptions.size} models",
style = MaterialTheme.typography.labelMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Column(modifier = Modifier.weight(1f)) {
Text(text = "Model", style = MaterialTheme.typography.titleMedium)
Text(
text = "${modelOptions.size} models",
style = MaterialTheme.typography.labelMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
if (onRefresh != null) {
TextButton(
onClick = onRefresh,
enabled = !refreshing,
) {
if (refreshing) {
CircularProgressIndicator(
modifier = Modifier.size(16.dp),
strokeWidth = 2.dp,
)
} else {
Icon(
imageVector = Icons.Filled.Refresh,
contentDescription = null,
modifier = Modifier.size(18.dp),
)
}
Spacer(modifier = Modifier.size(6.dp))
Text(if (refreshing) "Refreshing" else "Refresh")
}
}
}
Spacer(modifier = Modifier.height(8.dp))
@@ -4,6 +4,7 @@ import android.app.Activity
import android.content.Context
import android.content.ContextWrapper
import android.content.pm.ActivityInfo
import android.view.WindowManager
import androidx.compose.runtime.Composable
import androidx.compose.runtime.DisposableEffect
import androidx.compose.ui.platform.LocalContext
@@ -40,3 +41,30 @@ private tailrec fun Context.findActivity(): Activity? = when (this) {
is ContextWrapper -> baseContext.findActivity()
else -> null
}
/**
* Holds `Window.FLAG_KEEP_SCREEN_ON` while [enabled] is true, clearing it the
* moment it flips false or this composable leaves the composition. This is
* the Android-recommended mechanism for "keep the screen on while this UI is
* active" — see [com.hermesandroid.relay.power.WakeLockManager]'s doc comment
* for why a `PowerManager` wake lock is the wrong tool for a visible surface.
*
* Single call site by design: the flag is a plain bit on the window, not
* ref-counted, so two independent callers toggling it independently could
* stomp each other (one disposing clears a flag the other still wants held).
* Callers that need to OR multiple conditions (e.g. "streaming a reply" OR
* "voice mode is open") should combine them into one boolean and pass that.
*/
@Composable
fun KeepScreenOnWhile(enabled: Boolean) {
val context = LocalContext.current
DisposableEffect(enabled) {
val window = context.findActivity()?.window
if (enabled) {
window?.addFlags(WindowManager.LayoutParams.FLAG_KEEP_SCREEN_ON)
}
onDispose {
window?.clearFlags(WindowManager.LayoutParams.FLAG_KEEP_SCREEN_ON)
}
}
}
@@ -155,6 +155,7 @@ fun VoiceModeOverlay(
// Cancels the promoted/durable background Hermes run from the chip's ✕.
// Default no-op so existing call sites/previews keep compiling.
onBackgroundRunCancel: () -> Unit = {},
onBackgroundRunTap: () -> Unit = {},
onHermesConfirmationAnswer: (String) -> Unit = {},
// === END v0.4.1 ===
) {
@@ -174,14 +175,12 @@ fun VoiceModeOverlay(
}
}
// Pipe classified voice errors to the app-wide snackbar host. The inline
// error banner stays as a belt-and-suspenders for longer-lived messages.
val snackbarHost = LocalSnackbarHost.current
LaunchedEffect(errorEvents) {
errorEvents?.collect { err ->
snackbarHost.showHumanError(err)
}
}
// Voice errors surface ONLY on the overlay's own inline top banner
// (uiState.error) while the overlay is up — we deliberately do NOT also pipe
// them to the app-wide bottom snackbar. Doing both duplicated the message
// and left a retry-only, un-dismissable toast at the bottom during long /
// timed-out background runs. (errorEvents is still used by the chat + voice
// settings surfaces when the overlay isn't the active surface.)
Box(
modifier = modifier
@@ -351,6 +350,7 @@ fun VoiceModeOverlay(
BackgroundRunChip(
run = uiState.backgroundRun,
onCancel = onBackgroundRunCancel,
onTap = onBackgroundRunTap,
modifier = Modifier
.fillMaxWidth()
.padding(horizontal = 24.dp, vertical = 4.dp),
@@ -447,6 +447,25 @@ fun VoiceModeOverlay(
}
}
// Compact mode: the background-run chip must survive outside focus
// mode too — a running task with no visible presence reads as lost
// (the chip previously existed ONLY in the focus layout).
AnimatedVisibility(
visible = !focusMode && uiState.backgroundRun != null,
enter = fadeIn(tween(140)),
exit = fadeOut(tween(180)),
modifier = Modifier
.align(Alignment.BottomCenter)
.padding(bottom = 120.dp, start = 24.dp, end = 24.dp),
) {
BackgroundRunChip(
run = uiState.backgroundRun,
onCancel = onBackgroundRunCancel,
onTap = onBackgroundRunTap,
modifier = Modifier.fillMaxWidth(),
)
}
// Error banner
AnimatedVisibility(
visible = uiState.error != null,
@@ -470,8 +489,14 @@ fun VoiceModeOverlay(
text = uiState.error.orEmpty(),
color = MaterialTheme.colorScheme.onErrorContainer,
style = MaterialTheme.typography.bodySmall,
modifier = Modifier.weight(1f, fill = false),
modifier = Modifier.weight(1f),
)
// Dismiss clears the error and returns to Idle without
// retrying, so a failed/timed-out turn never traps the user
// on a retry-only banner.
TextButton(onClick = { onClearError() }) {
Text("Dismiss")
}
TextButton(
onClick = {
onClearError()
@@ -1239,6 +1264,18 @@ private fun CompactTranscriptRow(
)
}
val hasText = message.content.isNotBlank()
// Tools run before the agent's reply, so render the tool rows above the
// reply text — the bubble then reads in chronological order (tool calls
// ran → answer) instead of showing the answer above the tools that
// preceded it.
if (message.toolCalls.isNotEmpty()) {
Column(verticalArrangement = Arrangement.spacedBy(6.dp)) {
message.toolCalls.forEach { toolCall ->
VoiceToolStatusRow(toolCall)
}
}
if (hasText) Spacer(Modifier.height(6.dp))
}
when {
isVoiceActionBubble && hasText -> MarkdownContent(
content = message.content,
@@ -1259,14 +1296,6 @@ private fun CompactTranscriptRow(
overflow = TextOverflow.Ellipsis,
)
}
if (message.toolCalls.isNotEmpty()) {
Spacer(Modifier.height(6.dp))
Column(verticalArrangement = Arrangement.spacedBy(6.dp)) {
message.toolCalls.forEach { toolCall ->
VoiceToolStatusRow(toolCall)
}
}
}
}
}
@@ -1480,6 +1509,7 @@ private fun DestructiveCountdownRow(
private fun BackgroundRunChip(
run: BackgroundRunState?,
onCancel: () -> Unit,
onTap: () -> Unit = {},
modifier: Modifier = Modifier,
) {
var latest by remember { mutableStateOf<BackgroundRunState?>(null) }
@@ -1509,14 +1539,19 @@ private fun BackgroundRunChip(
animationSpec = infiniteRepeatable(tween(900), RepeatMode.Reverse),
label = "bgRunPulseAlpha",
)
val done = display.phase == BackgroundRunPhase.DONE
val dotColor = when (display.phase) {
BackgroundRunPhase.RECONNECTING -> MaterialTheme.colorScheme.tertiary
else -> MaterialTheme.colorScheme.primary
}
// A settled (DONE) chip reads as an outcome, not activity: solid dot,
// no pulse, no live ticker.
val dotAlpha = if (done) 1f else pulse
val title = when (display.phase) {
BackgroundRunPhase.RECONNECTING -> "Reconnecting — your task is still running"
BackgroundRunPhase.DELIVERING -> display.message
BackgroundRunPhase.RUNNING -> display.message
BackgroundRunPhase.DONE -> display.message
}
val detail = buildList {
display.statusLine
@@ -1528,12 +1563,18 @@ private fun BackgroundRunChip(
if (display.completedToolCount == 1) "" else "s"
)
}
add(elapsedLabel)
if (display.queuedCount > 0) {
add("+${display.queuedCount} queued")
}
if (!done) add(elapsedLabel)
}.joinToString(" · ")
Surface(
shape = RoundedCornerShape(16.dp),
color = MaterialTheme.colorScheme.secondaryContainer,
// Tap on a settled chip = respeak the delivered answer (the VM
// no-ops the tap for live phases).
onClick = onTap,
) {
Row(
verticalAlignment = Alignment.CenterVertically,
@@ -1543,7 +1584,7 @@ private fun BackgroundRunChip(
modifier = Modifier
.size(8.dp)
.clip(CircleShape)
.background(dotColor.copy(alpha = pulse)),
.background(dotColor.copy(alpha = dotAlpha)),
)
Spacer(Modifier.width(10.dp))
Column(modifier = Modifier.weight(1f)) {
@@ -1570,7 +1611,9 @@ private fun BackgroundRunChip(
) {
Icon(
imageVector = Icons.Filled.Close,
contentDescription = "Cancel background task",
// The VM treats ✕ on a DONE chip as a local dismiss,
// never a cancel — label it accordingly for TalkBack.
contentDescription = if (done) "Dismiss" else "Cancel background task",
tint = MaterialTheme.colorScheme.onSecondaryContainer,
modifier = Modifier.size(16.dp),
)
@@ -167,7 +167,7 @@ object PetImporter {
val queue = ArrayDeque<File>()
queue.add(root)
while (queue.isNotEmpty()) {
val dir = queue.removeFirst()
val dir = queue.removeAt(0)
val manifest = File(dir, "pet.json")
if (manifest.isFile) return manifest
dir.listFiles()?.forEach { if (it.isDirectory) queue.add(it) }
@@ -5,6 +5,7 @@ import com.hermesandroid.relay.ui.theme.LocalBrand
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.BoxScope
import androidx.compose.foundation.layout.BoxWithConstraints
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.ColumnScope
import androidx.compose.foundation.layout.Spacer
@@ -13,6 +14,8 @@ import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.widthIn
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.verticalScroll
import androidx.compose.foundation.shape.CircleShape
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.material.icons.Icons
@@ -69,66 +72,80 @@ fun OnboardingPage(
}
)
Column(
modifier = modifier
.fillMaxWidth()
.widthIn(max = 560.dp)
.padding(horizontal = 24.dp, vertical = 12.dp),
horizontalAlignment = Alignment.CenterHorizontally,
verticalArrangement = Arrangement.Center
) {
Card(
modifier = Modifier
.fillMaxWidth()
.height(232.dp)
.gradientBorder(shape = heroShape, isDarkTheme = isDarkTheme),
shape = heroShape,
colors = CardDefaults.cardColors(
containerColor = if (transparentHero) Color.Transparent else MaterialTheme.colorScheme.surfaceContainer
)
) {
val heroModifier = Modifier
.fillMaxWidth()
.height(232.dp)
Box(
modifier = if (transparentHero) heroModifier else heroModifier.background(heroBrush),
contentAlignment = Alignment.Center
) {
heroContent()
}
// Short viewports (small phones, large font scale, split screen) shrink or
// drop the hero so the body text fits; the vertical scroll below is the
// safety net when even that isn't enough. The enclosing pager Box centers
// short content, so no Arrangement.Center here — it conflicts with
// verticalScroll when content overflows.
BoxWithConstraints(modifier = modifier.fillMaxWidth()) {
val heroHeight = when {
maxHeight < 480.dp -> 0.dp
maxHeight < 620.dp -> 160.dp
else -> 232.dp
}
Spacer(modifier = Modifier.height(18.dp))
Card(
Column(
modifier = Modifier
.fillMaxWidth()
.gradientBorder(shape = bodyShape, isDarkTheme = isDarkTheme),
shape = bodyShape,
colors = CardDefaults.cardColors(
containerColor = MaterialTheme.colorScheme.surfaceVariant
)
.widthIn(max = 560.dp)
.verticalScroll(rememberScrollState())
.padding(horizontal = 24.dp, vertical = 12.dp),
horizontalAlignment = Alignment.CenterHorizontally
) {
Column(
modifier = Modifier.padding(horizontal = 24.dp, vertical = 22.dp),
horizontalAlignment = Alignment.CenterHorizontally,
verticalArrangement = Arrangement.spacedBy(14.dp)
if (heroHeight > 0.dp) {
Card(
modifier = Modifier
.fillMaxWidth()
.height(heroHeight)
.gradientBorder(shape = heroShape, isDarkTheme = isDarkTheme),
shape = heroShape,
colors = CardDefaults.cardColors(
containerColor = if (transparentHero) Color.Transparent else MaterialTheme.colorScheme.surfaceContainer
)
) {
val heroModifier = Modifier
.fillMaxWidth()
.height(heroHeight)
Box(
modifier = if (transparentHero) heroModifier else heroModifier.background(heroBrush),
contentAlignment = Alignment.Center
) {
heroContent()
}
}
Spacer(modifier = Modifier.height(18.dp))
}
Card(
modifier = Modifier
.fillMaxWidth()
.gradientBorder(shape = bodyShape, isDarkTheme = isDarkTheme),
shape = bodyShape,
colors = CardDefaults.cardColors(
containerColor = MaterialTheme.colorScheme.surfaceVariant
)
) {
Text(
text = title,
style = MaterialTheme.typography.headlineMedium,
textAlign = TextAlign.Center,
color = MaterialTheme.colorScheme.onSurface
)
Column(
modifier = Modifier.padding(horizontal = 24.dp, vertical = 22.dp),
horizontalAlignment = Alignment.CenterHorizontally,
verticalArrangement = Arrangement.spacedBy(14.dp)
) {
Text(
text = title,
style = MaterialTheme.typography.headlineMedium,
textAlign = TextAlign.Center,
color = MaterialTheme.colorScheme.onSurface
)
Text(
text = description,
style = MaterialTheme.typography.bodyLarge,
textAlign = TextAlign.Center,
color = MaterialTheme.colorScheme.onSurfaceVariant
)
Text(
text = description,
style = MaterialTheme.typography.bodyLarge,
textAlign = TextAlign.Center,
color = MaterialTheme.colorScheme.onSurfaceVariant
)
content()
content()
}
}
}
}
@@ -47,6 +47,7 @@ import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.clip
import androidx.compose.ui.platform.LocalConfiguration
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.res.painterResource
import androidx.compose.ui.text.font.FontWeight
@@ -208,11 +209,15 @@ fun OnboardingScreen(
// Bottom navigation only on informational pages — the wizard
// owns its own back/pair affordances.
if (pages[pagerState.currentPage] != OnboardingPage.Connect) {
// Short viewports get a tighter footer so more of the pager
// content stays above the fold; indicator + Back/Next remain
// pinned outside the (scrollable) pager pages either way.
val compactHeight = LocalConfiguration.current.screenHeightDp < 620
Column(
modifier = Modifier
.fillMaxWidth()
.padding(horizontal = 32.dp)
.padding(bottom = 48.dp),
.padding(bottom = if (compactHeight) 16.dp else 48.dp),
horizontalAlignment = Alignment.CenterHorizontally
) {
PageIndicator(
@@ -220,7 +225,7 @@ fun OnboardingScreen(
currentPage = pagerState.currentPage
)
Spacer(modifier = Modifier.height(24.dp))
Spacer(modifier = Modifier.height(if (compactHeight) 12.dp else 24.dp))
Row(
modifier = Modifier.fillMaxWidth(),
@@ -154,6 +154,7 @@ import com.hermesandroid.relay.ui.components.ChatInputPickerControl
import com.hermesandroid.relay.ui.components.ChatInputPickerOption
import com.hermesandroid.relay.ui.components.ChatInputTrailing
import com.hermesandroid.relay.ui.components.CommandPalette
import com.hermesandroid.relay.ui.components.KeepScreenOnWhile
import com.hermesandroid.relay.ui.components.ModelPickerSheet
import com.hermesandroid.relay.ui.components.ConnectionStatusBadge
import com.hermesandroid.relay.ui.components.CommandRow
@@ -461,7 +462,16 @@ fun ChatScreen(
val messages by chatViewModel.messages.collectAsState()
val isStreaming by chatViewModel.isStreaming.collectAsState()
// Keep the screen on for the two "actively engaged, hands-off-keyboard"
// cases: voice mode is a call-like continuous session (mirrors Assistant/
// phone-call UIs, held the whole time the overlay is up), and an
// in-flight chat reply is closer to video playback (held only while
// isStreaming — reading/scrolling an idle transcript uses the OS default,
// matching WhatsApp/Telegram/Signal norms). Single call site: the window
// flag isn't ref-counted, see KeepScreenOnWhile's doc comment.
KeepScreenOnWhile(enabled = voiceUiState.voiceMode || isStreaming)
val turnStatus by chatViewModel.turnStatus.collectAsState()
val recoveringAnswer by chatViewModel.recoveringAnswer.collectAsState()
val voiceStats by voiceViewModel.voiceStats.collectAsState()
var voiceOutputConfig by remember { mutableStateOf<VoiceOutputConfig?>(null) }
var realtimeAgentConfig by remember { mutableStateOf<RealtimeVoiceConfig?>(null) }
@@ -499,6 +509,7 @@ fun ChatScreen(
val serverModelName by chatViewModel.serverModelName.collectAsState()
val availableModels by chatViewModel.availableModels.collectAsState()
val modelProviders by chatViewModel.modelProviders.collectAsState()
val modelOptionsRefreshing by chatViewModel.modelOptionsRefreshing.collectAsState()
val selectedModelOverride by chatViewModel.selectedModelOverride.collectAsState()
val gatewayCurrentModel by chatViewModel.gatewayCurrentModel.collectAsState()
val selectedReasoningEffort by chatViewModel.selectedReasoningEffort.collectAsState()
@@ -702,12 +713,14 @@ fun ChatScreen(
voiceOutputConfig?.default_provider
}
val activeVoiceModel = if (realtimeAgentActive) {
realtimeAgentConfig?.default_model
voiceStats.realtimeModel.takeIf { it.isNotBlank() }
?: realtimeAgentConfig?.default_model
} else {
voiceOutputConfig?.default_model
}
val activeVoiceName = if (realtimeAgentActive) {
realtimeAgentConfig?.default_voice
voiceStats.realtimeVoice.takeIf { it.isNotBlank() }
?: realtimeAgentConfig?.default_voice
} else {
voiceOutputConfig?.default_voice
}
@@ -2068,6 +2081,7 @@ fun ChatScreen(
showThinking = showThinking,
isFirstInGroup = isFirstInGroup,
isLastInGroup = isLastInGroup,
recoveringAnswer = recoveringAnswer,
onAttachmentRetry = { msgId, idx ->
chatViewModel.manualFetchAttachment(msgId, idx)
},
@@ -2708,6 +2722,8 @@ fun ChatScreen(
if (showModelSheet) {
ModelPickerSheet(
options = modelPickerOptions,
refreshing = modelOptionsRefreshing,
onRefresh = { chatViewModel.refreshModelOptions(refresh = true) },
onSelect = { option ->
showModelSheet = false
chatViewModel.selectModel(option.value, option.provider)
@@ -2832,6 +2848,7 @@ fun ChatScreen(
onModeChange = { voiceViewModel.setInteractionMode(it) },
onClearError = { voiceViewModel.clearError() },
onBackgroundRunCancel = { voiceViewModel.cancelBackgroundRun() },
onBackgroundRunTap = { voiceViewModel.respeakBackgroundResult() },
// Agent B's overlay collects this flow and renders classified
// voice errors (mic capture, STT/TTS failures, relay drops).
errorEvents = voiceViewModel.errorEvents,
@@ -54,6 +54,7 @@ import androidx.compose.material3.ButtonDefaults
import androidx.compose.material3.Card
import androidx.compose.material3.CardDefaults
import androidx.compose.material3.Checkbox
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.DropdownMenu
import androidx.compose.material3.DropdownMenuItem
import androidx.compose.material3.ExperimentalMaterial3Api
@@ -2444,24 +2445,29 @@ private data class ExpensiveModelConfirm(
val warning: String,
)
private data class ModelProviderOption(
internal data class ModelProviderOption(
val id: String,
val label: String,
val authenticated: Boolean,
val models: List<String>,
/** Upstream setup hint for unconfigured rows, e.g. "paste OPENAI_API_KEY to activate". */
val setupHint: String? = null,
)
/**
* Tolerant reader for `GET /api/model/options` (the REST twin of the TUI's
* `model.options` RPC). Unauthenticated providers come back as skeleton rows
* — keep them visible but unselectable so the user learns which key to add
* in the Keys section instead of the provider silently missing.
* `model.options` RPC, requested with `include_unconfigured=1`).
* Unauthenticated providers come back as skeleton rows — empty `models` on
* newer upstream — and MUST survive parsing: they render greyed/unselectable
* so the user learns which key to add in the Keys section instead of the
* provider silently missing.
*/
private fun parseModelOptions(root: JsonObject): List<ModelProviderOption> {
internal fun parseModelOptions(root: JsonObject): List<ModelProviderOption> {
val providers = root["providers"] as? JsonArray ?: return emptyList()
return providers.mapNotNull { element ->
val obj = element as? JsonObject ?: return@mapNotNull null
val id = obj.stringField("id")
val id = obj.stringField("slug")
?: obj.stringField("id")
?: obj.stringField("provider")
?: obj.stringField("name")
?: return@mapNotNull null
@@ -2472,15 +2478,17 @@ private fun parseModelOptions(root: JsonObject): List<ModelProviderOption> {
else -> null
}?.trim()?.takeIf { it.isNotBlank() }
}.orEmpty()
if (models.isEmpty()) return@mapNotNull null
ModelProviderOption(
id = id,
label = obj.stringField("label")
?: obj.stringField("display_name")
?: obj.stringField("name")
?: id,
authenticated = obj.booleanField("authenticated") != false,
// Absent hint field: a row with models is assumed usable; an empty
// row can only be an unconfigured skeleton, so grey it.
authenticated = obj.booleanField("authenticated") ?: models.isNotEmpty(),
models = models,
setupHint = obj.stringField("warning"),
)
}.sortedByDescending { it.authenticated }
}
@@ -2494,27 +2502,36 @@ private fun ModelPickerDialog(
onDismiss: () -> Unit,
) {
var loading by remember { mutableStateOf(true) }
var refreshing by remember { mutableStateOf(false) }
var error by remember { mutableStateOf<String?>(null) }
var providers by remember { mutableStateOf<List<ModelProviderOption>>(emptyList()) }
val scope = rememberCoroutineScope()
fun loadOptions(refresh: Boolean = false) {
if (refresh && refreshing) return
if (refresh) refreshing = true else loading = true
error = null
scope.launch {
val result = try {
withDashboardClient(clientFactory) { client -> client.getModelOptions(refresh = refresh) }
} catch (e: Exception) {
Result.failure(e)
}
result.fold(
onSuccess = { root ->
providers = parseModelOptions(root)
if (providers.isEmpty()) {
error = "The dashboard returned no model options."
}
},
onFailure = { err -> error = err.message ?: "Could not load model options" },
)
if (refresh) refreshing = false else loading = false
}
}
LaunchedEffect(target) {
loading = true
error = null
val result = try {
withDashboardClient(clientFactory) { client -> client.getModelOptions() }
} catch (e: Exception) {
Result.failure(e)
}
result.fold(
onSuccess = { root ->
providers = parseModelOptions(root)
if (providers.isEmpty()) {
error = "The dashboard returned no model options."
}
},
onFailure = { err -> error = err.message ?: "Could not load model options" },
)
loading = false
loadOptions()
}
AlertDialog(
@@ -2538,12 +2555,37 @@ private fun ModelPickerDialog(
color = MaterialTheme.colorScheme.error,
)
else -> Column {
Text(
text = "Applies to new sessions. Greyed providers need a key — " +
"add one under Manage → Keys.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Row(
modifier = Modifier.fillMaxWidth(),
horizontalArrangement = Arrangement.SpaceBetween,
verticalAlignment = Alignment.CenterVertically,
) {
Text(
text = "Applies to new sessions. Greyed providers need a key — " +
"add one under Manage → Keys.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
modifier = Modifier.weight(1f),
)
TextButton(
onClick = { loadOptions(refresh = true) },
enabled = !refreshing && !actionInFlight,
) {
if (refreshing) {
CircularProgressIndicator(
modifier = Modifier.size(16.dp),
strokeWidth = 2.dp,
)
} else {
Icon(
imageVector = Icons.Filled.Refresh,
contentDescription = null,
modifier = Modifier.size(18.dp),
)
}
Text(if (refreshing) "Refreshing" else "Refresh")
}
}
LazyColumn(modifier = Modifier.heightIn(max = 400.dp)) {
providers.forEach { provider ->
item(key = "provider-${provider.id}") {
@@ -2559,6 +2601,22 @@ private fun ModelPickerDialog(
modifier = Modifier.padding(top = 12.dp, bottom = 2.dp),
)
}
if (provider.models.isEmpty()) {
// Unconfigured skeleton row (include_unconfigured=1):
// no models until a key lands, so show the server's
// setup hint in place of the model list.
item(key = "setup-${provider.id}") {
Text(
text = provider.setupHint
?: "Add a key under Manage → Keys to unlock models.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
modifier = Modifier
.fillMaxWidth()
.padding(vertical = 6.dp),
)
}
}
items(
items = provider.models,
key = { model -> "model-${provider.id}-$model" },
@@ -19,8 +19,8 @@ import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.verticalScroll
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.automirrored.filled.ArrowBack
import androidx.compose.material.icons.automirrored.filled.OpenInNew
import androidx.compose.material.icons.filled.Notifications
import androidx.compose.material.icons.filled.OpenInNew
import androidx.compose.material.icons.filled.Refresh
import androidx.compose.material3.Button
import androidx.compose.material3.Card
@@ -31,15 +31,21 @@ import androidx.compose.material3.HorizontalDivider
import androidx.compose.material3.Icon
import androidx.compose.material3.IconButton
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.OutlinedTextField
import androidx.compose.material3.Scaffold
import androidx.compose.material3.Switch
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.material3.TopAppBar
import androidx.compose.material3.TopAppBarDefaults
import androidx.compose.runtime.Composable
import androidx.compose.runtime.DisposableEffect
import androidx.compose.runtime.LaunchedEffect
import androidx.compose.runtime.collectAsState
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.rememberCoroutineScope
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
@@ -51,20 +57,20 @@ import androidx.lifecycle.Lifecycle
import androidx.lifecycle.LifecycleEventObserver
import androidx.lifecycle.compose.LocalLifecycleOwner
import com.hermesandroid.relay.notifications.HermesNotificationCompanion
import com.hermesandroid.relay.notifications.NotificationTriggerAction
import com.hermesandroid.relay.notifications.NotificationTriggerRule
import com.hermesandroid.relay.notifications.NotificationTriggerSettings
import com.hermesandroid.relay.notifications.NotificationTriggerStore
import com.hermesandroid.relay.notifications.notificationTriggerDataStore
import com.hermesandroid.relay.notifications.summary
import kotlinx.coroutines.launch
import java.text.DateFormat
import java.util.Date
/**
* Notification companion settings screen — opt-in helper that lets the
* user's Hermes assistant read notifications they've explicitly granted
* access to.
*
* Sections:
* 1. Status — granted or not granted (live, observed via lifecycle)
* 2. Open Android Settings — fires `ACTION_NOTIFICATION_LISTENER_SETTINGS`
* 3. About — explains what the feature does and how to revoke
*
* Mirrors the layout/style of [VoiceSettingsScreen] for consistency.
* access to, plus the event-trigger MVP rule editor/activity log.
*/
@OptIn(ExperimentalMaterial3Api::class)
@Composable
@@ -72,11 +78,18 @@ fun NotificationCompanionSettingsScreen(
onBack: () -> Unit,
) {
val context = LocalContext.current
val appContext = context.applicationContext
val lifecycleOwner = LocalLifecycleOwner.current
val scope = rememberCoroutineScope()
val triggerStore = remember(appContext) {
NotificationTriggerStore(appContext.notificationTriggerDataStore)
}
val triggerSettings by triggerStore.settings.collectAsState(
initial = NotificationTriggerSettings(),
)
// Track grant state. Re-check on every ON_RESUME so the screen
// updates immediately when the user comes back from Android
// Settings after toggling the listener.
// updates immediately when the user comes back from Android Settings.
var granted by remember {
mutableStateOf(HermesNotificationCompanion.isAccessGranted(context))
}
@@ -117,214 +130,440 @@ fun NotificationCompanionSettingsScreen(
.padding(16.dp),
verticalArrangement = Arrangement.spacedBy(16.dp),
) {
NotificationAccessCard(
granted = granted,
onRefresh = { granted = HermesNotificationCompanion.isAccessGranted(context) },
onOpenSettings = {
val intent = Intent(Settings.ACTION_NOTIFICATION_LISTENER_SETTINGS)
intent.addFlags(Intent.FLAG_ACTIVITY_NEW_TASK)
runCatching { context.startActivity(intent) }
},
)
// --- Status --- (action first: the user came to turn it on)
NotifSectionCard(title = "Status") {
Row(
modifier = Modifier.fillMaxWidth(),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.spacedBy(12.dp),
) {
Box(
modifier = Modifier
.size(12.dp)
.clip(CircleShape)
.background(
if (granted) {
Color(0xFF43A047) // green
} else {
MaterialTheme.colorScheme.outline
},
),
NotificationTriggerCard(
settings = triggerSettings,
onMasterChanged = { enabled ->
scope.launch { triggerStore.setMasterEnabled(enabled) }
},
onKillSwitchChanged = { enabled ->
scope.launch { triggerStore.setKillSwitch(enabled) }
},
onSaveRule = { rule ->
scope.launch { triggerStore.saveSingleRule(rule) }
},
)
NotificationActivityLogCard(
settings = triggerSettings,
onClear = { scope.launch { triggerStore.clearActivityLog() } },
)
NotificationAboutCard()
NotificationTestCard(granted = granted)
}
}
}
@Composable
private fun NotificationAccessCard(
granted: Boolean,
onRefresh: () -> Unit,
onOpenSettings: () -> Unit,
) {
NotifSectionCard(title = "Status") {
Row(
modifier = Modifier.fillMaxWidth(),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.spacedBy(12.dp),
) {
Box(
modifier = Modifier
.size(12.dp)
.clip(CircleShape)
.background(
if (granted) Color(0xFF43A047) else MaterialTheme.colorScheme.outline,
),
)
Column(modifier = Modifier.weight(1f)) {
Text(
text = if (granted) "Access granted" else "Access not granted",
style = MaterialTheme.typography.bodyLarge,
)
Text(
text = if (granted) {
"Hermes-Relay can read posted notifications"
} else {
"Tap below to enable in Android Settings"
},
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
IconButton(onClick = onRefresh) {
Icon(
imageVector = Icons.Filled.Refresh,
contentDescription = "Refresh status",
)
}
}
HorizontalDivider(modifier = Modifier.padding(vertical = 12.dp))
Button(
onClick = onOpenSettings,
modifier = Modifier.fillMaxWidth(),
) {
Icon(
imageVector = Icons.AutoMirrored.Filled.OpenInNew,
contentDescription = null,
)
Spacer(Modifier.size(8.dp))
Text(if (granted) "Manage in Android Settings" else "Open Android Settings")
}
}
}
@Composable
private fun NotificationTriggerCard(
settings: NotificationTriggerSettings,
onMasterChanged: (Boolean) -> Unit,
onKillSwitchChanged: (Boolean) -> Unit,
onSaveRule: (NotificationTriggerRule) -> Unit,
) {
val emptyRule = remember { NotificationTriggerStore.defaultRule() }
val savedRule = settings.rules.firstOrNull() ?: emptyRule
var ruleEnabled by remember { mutableStateOf(savedRule.enabled) }
var label by remember { mutableStateOf(savedRule.label) }
var appPackage by remember { mutableStateOf(savedRule.appPackage.orEmpty()) }
var titleContains by remember { mutableStateOf(savedRule.titleContains.orEmpty()) }
var textContains by remember { mutableStateOf(savedRule.textContains.orEmpty()) }
LaunchedEffect(
savedRule.id,
savedRule.enabled,
savedRule.label,
savedRule.appPackage,
savedRule.titleContains,
savedRule.textContains,
) {
ruleEnabled = savedRule.enabled
label = savedRule.label
appPackage = savedRule.appPackage.orEmpty()
titleContains = savedRule.titleContains.orEmpty()
textContains = savedRule.textContains.orEmpty()
}
val hasAnyFilter = appPackage.isNotBlank() || titleContains.isNotBlank() || textContains.isNotBlank()
NotifSectionCard(title = "Event triggers (MVP)") {
Text(
text = "Rules are off by default. When enabled, the first MVP action is safe: post a local “Ask Hermes?” prompt when a matching notification arrives.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(8.dp))
LabeledSwitchRow(
title = "Enable proactive triggers",
subtitle = "Explicit opt-in. Existing notification forwarding still works when this is off.",
checked = settings.masterEnabled,
onCheckedChange = onMasterChanged,
)
LabeledSwitchRow(
title = "Kill switch",
subtitle = "Immediately pauses all trigger actions without deleting rules or the log.",
checked = settings.killSwitch,
onCheckedChange = onKillSwitchChanged,
)
HorizontalDivider(modifier = Modifier.padding(vertical = 12.dp))
Text("Rule", style = MaterialTheme.typography.titleSmall)
Text(
text = "Match by app package and optional title/text contains filters. Example package: com.slack.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(8.dp))
LabeledSwitchRow(
title = "Rule enabled",
subtitle = savedRule.summary(),
checked = ruleEnabled,
onCheckedChange = { ruleEnabled = it },
)
OutlinedTextField(
value = label,
onValueChange = { label = it },
modifier = Modifier.fillMaxWidth(),
singleLine = true,
label = { Text("Rule label") },
)
OutlinedTextField(
value = appPackage,
onValueChange = { appPackage = it },
modifier = Modifier.fillMaxWidth(),
singleLine = true,
label = { Text("App package") },
placeholder = { Text("com.example.app") },
)
OutlinedTextField(
value = titleContains,
onValueChange = { titleContains = it },
modifier = Modifier.fillMaxWidth(),
singleLine = true,
label = { Text("Title contains (optional)") },
)
OutlinedTextField(
value = textContains,
onValueChange = { textContains = it },
modifier = Modifier.fillMaxWidth(),
singleLine = true,
label = { Text("Text contains (optional)") },
)
Button(
onClick = {
onSaveRule(
NotificationTriggerRule(
id = savedRule.id,
label = label,
enabled = ruleEnabled,
appPackage = appPackage,
titleContains = titleContains,
textContains = textContains,
action = NotificationTriggerAction.AskMe,
requireConfirmation = false,
),
)
},
enabled = hasAnyFilter,
modifier = Modifier.fillMaxWidth(),
) {
Text("Save rule")
}
if (!hasAnyFilter) {
Text(
text = "Set at least one filter before saving so triggers do not match every notification on the phone.",
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.error,
)
}
}
}
@Composable
private fun NotificationActivityLogCard(
settings: NotificationTriggerSettings,
onClear: () -> Unit,
) {
NotifSectionCard(title = "Activity log") {
Row(
modifier = Modifier.fillMaxWidth(),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.SpaceBetween,
) {
Text(
text = "Latest trigger matches",
style = MaterialTheme.typography.bodyMedium,
)
TextButton(
onClick = onClear,
enabled = settings.activityLog.isNotEmpty(),
) { Text("Clear") }
}
if (settings.activityLog.isEmpty()) {
Text(
text = "No trigger activity yet.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
return@NotifSectionCard
}
val df = remember { DateFormat.getDateTimeInstance(DateFormat.SHORT, DateFormat.SHORT) }
settings.activityLog.take(10).forEach { entry ->
Column(modifier = Modifier.padding(vertical = 6.dp)) {
Text(
text = entry.ruleLabel,
style = MaterialTheme.typography.bodyMedium,
)
Text(
text = "${entry.packageName} · ${df.format(Date(entry.matchedAt))}",
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
entry.title?.let {
Text(text = it, style = MaterialTheme.typography.bodySmall)
}
entry.textPreview?.let {
Text(
text = it,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Column(modifier = Modifier.weight(1f)) {
}
Text(
text = entry.result,
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.primary,
)
}
HorizontalDivider()
}
}
}
@Composable
private fun NotificationAboutCard() {
NotifSectionCard(title = "About") {
Text(
text = "Lets your Hermes assistant help you triage notifications. When enabled, your phone forwards each notification's app, title, and text to your paired Hermes server through the Relay pairing used by phone tools.",
style = MaterialTheme.typography.bodyMedium,
)
Spacer(Modifier.height(8.dp))
Text(
text = "Requires Android's notification access permission. You can grant or revoke it at any time in Android Settings. This is the same permission Wear OS, Android Auto, and Tasker use.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(8.dp))
Text(
text = "Confirmation policy: local prompts and local log writes can run automatically after opt-in. Anything that sends a message, replies in another app, routes content to a person/channel, uses bridge gestures, or spends/changes data still requires an explicit user confirmation first.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
@Composable
private fun NotificationTestCard(granted: Boolean) {
NotifSectionCard(title = "Test") {
Text(
text = "Tap below to fetch the last few notifications the listener has captured this session. Helpful for verifying the connection is working end-to-end.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(8.dp))
var lastSnapshot by remember {
mutableStateOf<List<TestNotificationLine>>(emptyList())
}
var lastError by remember { mutableStateOf<String?>(null) }
FilledTonalButton(
onClick = {
val service = HermesNotificationCompanion.active
if (service == null) {
lastSnapshot = emptyList()
lastError = if (granted) {
"Listener has not bound yet. Try posting a test notification and re-tap."
} else {
"Notification access is not granted."
}
} else {
val active = service.activeNotifications
lastError = null
lastSnapshot = active.orEmpty()
.sortedByDescending { it.postTime }
.take(5)
.map {
TestNotificationLine(
pkg = it.packageName,
title = it.notification?.extras
?.getCharSequence(android.app.Notification.EXTRA_TITLE)
?.toString(),
text = it.notification?.extras
?.getCharSequence(android.app.Notification.EXTRA_TEXT)
?.toString(),
postedAt = it.postTime,
)
}
}
},
modifier = Modifier.fillMaxWidth(),
) {
Icon(
imageVector = Icons.Filled.Notifications,
contentDescription = null,
)
Spacer(Modifier.size(8.dp))
Text("Fetch recent")
}
lastError?.let { err ->
Spacer(Modifier.height(8.dp))
Text(
text = err,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.error,
)
}
if (lastSnapshot.isNotEmpty()) {
Spacer(Modifier.height(8.dp))
val df = remember { DateFormat.getTimeInstance(DateFormat.SHORT) }
lastSnapshot.forEach { line ->
Column(modifier = Modifier.padding(vertical = 4.dp)) {
Row(
modifier = Modifier.fillMaxWidth(),
horizontalArrangement = Arrangement.SpaceBetween,
) {
Text(
text = if (granted) "Access granted" else "Access not granted",
style = MaterialTheme.typography.bodyLarge,
text = line.title ?: "(no title)",
style = MaterialTheme.typography.bodyMedium,
)
Text(
text = if (granted) {
"Hermes-Relay can read posted notifications"
} else {
"Tap below to enable in Android Settings"
},
style = MaterialTheme.typography.bodySmall,
text = df.format(Date(line.postedAt)),
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
IconButton(onClick = {
granted = HermesNotificationCompanion.isAccessGranted(context)
}) {
Icon(
imageVector = Icons.Filled.Refresh,
contentDescription = "Refresh status",
Text(
text = line.pkg,
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
line.text?.let {
Text(
text = it,
style = MaterialTheme.typography.bodySmall,
)
}
}
HorizontalDivider(modifier = Modifier.padding(vertical = 12.dp))
Button(
onClick = {
val intent = Intent(Settings.ACTION_NOTIFICATION_LISTENER_SETTINGS)
intent.addFlags(Intent.FLAG_ACTIVITY_NEW_TASK)
runCatching { context.startActivity(intent) }
},
modifier = Modifier.fillMaxWidth(),
) {
Icon(
imageVector = Icons.Filled.OpenInNew,
contentDescription = null,
)
Spacer(Modifier.size(8.dp))
Text(
if (granted) "Manage in Android Settings" else "Open Android Settings",
)
}
}
// --- About ---
NotifSectionCard(title = "About") {
Text(
text = (
"Lets your Hermes assistant help you triage " +
"notifications. When enabled, your phone " +
"forwards each notification's app, title, " +
"and text to your paired Hermes server " +
"through the Relay pairing used by phone tools."
),
style = MaterialTheme.typography.bodyMedium,
)
Spacer(Modifier.height(8.dp))
Text(
text = (
"Requires Android's notification access " +
"permission. You can grant or revoke it " +
"at any time in Android Settings. This is " +
"the same permission Wear OS, Android " +
"Auto, and Tasker use."
),
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
// --- Test (last received) ---
NotifSectionCard(title = "Test") {
Text(
text = (
"Tap below to fetch the last few notifications " +
"the listener has captured this session. " +
"Helpful for verifying the connection is " +
"working end-to-end."
),
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(8.dp))
var lastSnapshot by remember {
mutableStateOf<List<TestNotificationLine>>(emptyList())
}
var lastError by remember { mutableStateOf<String?>(null) }
FilledTonalButton(
onClick = {
// Pull from the live service if it's bound. We
// don't have a server-round-trip helper here on
// purpose — that would require sending a relay
// round-trip and we want this screen to work
// even when the relay is unreachable.
val service = HermesNotificationCompanion.active
if (service == null) {
lastSnapshot = emptyList()
lastError = if (granted) {
"Listener has not bound yet. Try " +
"posting a test notification " +
"(any message) and re-tap."
} else {
"Notification access is not granted."
}
} else {
val active = service.activeNotifications
lastError = null
lastSnapshot = active.orEmpty()
.sortedByDescending { it.postTime }
.take(5)
.map {
TestNotificationLine(
pkg = it.packageName,
title = it.notification?.extras
?.getCharSequence(android.app.Notification.EXTRA_TITLE)
?.toString(),
text = it.notification?.extras
?.getCharSequence(android.app.Notification.EXTRA_TEXT)
?.toString(),
postedAt = it.postTime,
)
}
}
},
modifier = Modifier.fillMaxWidth(),
) {
Icon(
imageVector = Icons.Filled.Notifications,
contentDescription = null,
)
Spacer(Modifier.size(8.dp))
Text("Fetch recent")
}
lastError?.let { err ->
Spacer(Modifier.height(8.dp))
Text(
text = err,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.error,
)
}
if (lastSnapshot.isNotEmpty()) {
Spacer(Modifier.height(8.dp))
val df = remember { DateFormat.getTimeInstance(DateFormat.SHORT) }
lastSnapshot.forEach { line ->
Column(modifier = Modifier.padding(vertical = 4.dp)) {
Row(
modifier = Modifier.fillMaxWidth(),
horizontalArrangement = Arrangement.SpaceBetween,
) {
Text(
text = line.title ?: "(no title)",
style = MaterialTheme.typography.bodyMedium,
)
Text(
text = df.format(Date(line.postedAt)),
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
Text(
text = line.pkg,
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
line.text?.let {
Text(
text = it,
style = MaterialTheme.typography.bodySmall,
)
}
}
HorizontalDivider()
}
}
HorizontalDivider()
}
}
}
}
// Local tiny model for the screen's "Test" preview list — keeps the
// composable scope contained without adding to the wire models file.
@Composable
private fun LabeledSwitchRow(
title: String,
subtitle: String,
checked: Boolean,
onCheckedChange: (Boolean) -> Unit,
) {
Row(
modifier = Modifier
.fillMaxWidth()
.padding(vertical = 4.dp),
horizontalArrangement = Arrangement.spacedBy(12.dp),
verticalAlignment = Alignment.CenterVertically,
) {
Column(modifier = Modifier.weight(1f)) {
Text(text = title, style = MaterialTheme.typography.bodyMedium)
Text(
text = subtitle,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
Switch(checked = checked, onCheckedChange = onCheckedChange)
}
}
private data class TestNotificationLine(
val pkg: String,
val title: String?,
@@ -360,4 +599,3 @@ private fun NotifSectionCard(
}
}
}
@@ -18,8 +18,10 @@ import androidx.compose.foundation.text.KeyboardOptions
import androidx.compose.foundation.verticalScroll
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.automirrored.filled.ArrowBack
import androidx.compose.material.icons.filled.Info
import androidx.compose.material.icons.filled.PlayArrow
import androidx.compose.material.icons.filled.Warning
import androidx.compose.material3.AlertDialog
import androidx.compose.material3.Card
import androidx.compose.material3.CardDefaults
import androidx.compose.material3.DropdownMenuItem
@@ -251,6 +253,8 @@ fun VoiceSettingsScreen(
currentEngine = currentEngine,
output = configState.voiceOutputConfig,
realtime = configState.realtimeConfig,
realtimeModel = voiceSettings.realtimeModel,
realtimeVoice = voiceSettings.realtimeVoice,
)
// --- Voice scope (single home; WP-V3 dedupe + WP-V4 honest label) ---
@@ -1435,22 +1439,29 @@ private fun RealtimeAgentCard(
var realtimeSampleRate by remember { mutableStateOf("24000") }
var realtimeSaving by remember { mutableStateOf(false) }
var realtimeManualOpen by remember { mutableStateOf(false) }
var showDeliveryInfo by remember { mutableStateOf(false) }
val config = configState.realtimeConfig
LaunchedEffect(
config?.enabled,
config?.default_provider,
config?.default_model,
config?.default_voice,
config?.sample_rate,
) {
val c = config ?: return@LaunchedEffect
realtimeEnabled = c.enabled
realtimeProvider = c.default_provider.orEmpty()
realtimeModel = c.default_model.orEmpty()
realtimeVoice = c.default_voice.orEmpty()
realtimeSampleRate = c.sample_rate.toString()
}
LaunchedEffect(
config?.default_model,
config?.default_voice,
voiceSettings.realtimeModel,
voiceSettings.realtimeVoice,
) {
val c = config ?: return@LaunchedEffect
realtimeModel = voiceSettings.realtimeModel.ifBlank { c.default_model.orEmpty() }
realtimeVoice = voiceSettings.realtimeVoice.ifBlank { c.default_voice.orEmpty() }
}
fun refreshRealtimeProviderOptions(providerId: String, applyDefaults: Boolean) {
settingsViewModel.refreshRealtimeProviderOptions(
@@ -1535,10 +1546,10 @@ private fun RealtimeAgentCard(
value = config?.default_provider
?: (configState.realtimeConfigError?.let { "unavailable" } ?: "loading..."),
)
config?.default_model?.let { model ->
realtimeModel.takeIf { it.isNotBlank() }?.let { model ->
ProviderRow(label = "Model", value = model)
}
config?.default_voice?.let { voice ->
realtimeVoice.takeIf { it.isNotBlank() }?.let { voice ->
ProviderRow(label = "Voice", value = voice)
}
config?.let { c ->
@@ -1610,17 +1621,34 @@ private fun RealtimeAgentCard(
)
}
Spacer(Modifier.height(8.dp))
Text(
"When the answer is ready",
style = MaterialTheme.typography.labelMedium,
)
Row(
modifier = Modifier.fillMaxWidth(),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.SpaceBetween,
) {
Text(
"When the answer is ready",
style = MaterialTheme.typography.labelMedium,
)
IconButton(
onClick = { showDeliveryInfo = true },
modifier = Modifier.size(36.dp),
) {
Icon(
imageVector = Icons.Filled.Info,
contentDescription = "Delivery mode details",
tint = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
Spacer(Modifier.height(4.dp))
val deliveryOptions = listOf(
"speak_verbatim",
"speak_when_idle",
"notify_then_speak",
"visual_only",
)
val deliveryLabels = listOf("Speak", "Notify", "Show only")
val deliveryLabels = listOf("Exact", "Summary", "Notify", "Show")
SingleChoiceSegmentedButtonRow(modifier = Modifier.fillMaxWidth()) {
deliveryOptions.forEachIndexed { index, option ->
SegmentedButton(
@@ -1649,6 +1677,10 @@ private fun RealtimeAgentCard(
}
}
if (showDeliveryInfo) {
DeliveryModeInfoDialog(onDismiss = { showDeliveryInfo = false })
}
HorizontalDivider(modifier = Modifier.padding(vertical = 8.dp))
Row(
@@ -1724,6 +1756,9 @@ private fun RealtimeAgentCard(
selectedRealtimeProvider?.let { provider ->
realtimeVoice = voiceForModel(provider, model, realtimeVoice)
}
scope.launch {
prefsRepo.setRealtimeSelection(realtimeModel, realtimeVoice)
}
},
enabled = voiceClient != null,
)
@@ -1735,7 +1770,10 @@ private fun RealtimeAgentCard(
realtimeVoice,
realtimeModel,
),
onValueChange = { realtimeVoice = it },
onValueChange = { voice ->
realtimeVoice = voice
scope.launch { prefsRepo.setRealtimeVoice(voice) }
},
enabled = voiceClient != null,
)
compatibilityNotice(
@@ -1774,14 +1812,20 @@ private fun RealtimeAgentCard(
)
OutlinedTextField(
value = realtimeModel,
onValueChange = { realtimeModel = it },
onValueChange = { model ->
realtimeModel = model
scope.launch { prefsRepo.setRealtimeModel(model) }
},
modifier = Modifier.fillMaxWidth(),
singleLine = true,
label = { Text("Model ID") },
)
OutlinedTextField(
value = realtimeVoice,
onValueChange = { realtimeVoice = it },
onValueChange = { voice ->
realtimeVoice = voice
scope.launch { prefsRepo.setRealtimeVoice(voice) }
},
modifier = Modifier.fillMaxWidth(),
singleLine = true,
label = { Text("Voice ID") },
@@ -1839,6 +1883,7 @@ private fun RealtimeAgentCard(
)
realtimeSaving = false
if (result.isSuccess) {
prefsRepo.setRealtimeSelection(realtimeModel, realtimeVoice)
settingsViewModel.setRealtimeConfig(result.getOrNull())
} else {
val human = classifyError(
@@ -1866,6 +1911,51 @@ private fun RealtimeAgentCard(
}
}
@Composable
private fun DeliveryModeInfoDialog(onDismiss: () -> Unit) {
AlertDialog(
onDismissRequest = onDismiss,
title = { Text("Answer delivery") },
text = {
Column(verticalArrangement = Arrangement.spacedBy(10.dp)) {
DeliveryModeInfoRow(
label = "Exact",
body = "Recommended. The realtime voice reads the Hermes answer word for word, falling back to standard TTS only if it goes off-script.",
)
DeliveryModeInfoRow(
label = "Summary",
body = "The realtime voice rephrases the result in its own words. More conversational, less faithful to the exact answer.",
)
DeliveryModeInfoRow(
label = "Notify",
body = "Shows that the answer is ready first, then speaks when you re-engage.",
)
DeliveryModeInfoRow(
label = "Show",
body = "Keeps the completed answer visual only.",
)
}
},
confirmButton = {
TextButton(onClick = onDismiss) {
Text("Got it")
}
},
)
}
@Composable
private fun DeliveryModeInfoRow(label: String, body: String) {
Column(verticalArrangement = Arrangement.spacedBy(2.dp)) {
Text(label, style = MaterialTheme.typography.titleSmall)
Text(
text = body,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
// ---------------------------------------------------------------------------
// Global voice controls — interaction mode + silence threshold.
// ---------------------------------------------------------------------------
@@ -2844,6 +2934,8 @@ private fun voiceOutputSummary(
currentEngine: VoiceEngineMode,
output: VoiceOutputConfig?,
realtime: RealtimeVoiceConfig?,
realtimeModel: String = "",
realtimeVoice: String = "",
): Pair<String, String> {
val profileLabel = profile?.description?.takeIf { it.isNotBlank() }
?: profile?.name?.takeIf { it.isNotBlank() }
@@ -2856,8 +2948,12 @@ private fun voiceOutputSummary(
} ?: "loading voice output..."
val realtimeLabel = realtime?.let { config ->
val provider = config.default_provider?.takeIf { it.isNotBlank() } ?: "realtime ..."
val voice = config.default_voice?.takeIf { it.isNotBlank() } ?: "voice ..."
"$provider / $voice"
val model = realtimeModel.takeIf { it.isNotBlank() }
?: config.default_model?.takeIf { it.isNotBlank() }
val voice = realtimeVoice.takeIf { it.isNotBlank() }
?: config.default_voice?.takeIf { it.isNotBlank() }
?: "voice ..."
if (model == null) "$provider / $voice" else "$provider / $model / $voice"
} ?: "realtime loading..."
return profileLabel to when (currentEngine) {
VoiceEngineMode.HermesVoiceOutput -> "Hermes Chat + Voice Output - $outputLabel"
@@ -2871,12 +2967,16 @@ private fun VoiceProfileSummaryCard(
currentEngine: VoiceEngineMode,
output: VoiceOutputConfig?,
realtime: RealtimeVoiceConfig?,
realtimeModel: String,
realtimeVoice: String,
) {
val (title, subtitle) = voiceOutputSummary(
profile = selectedProfile,
currentEngine = currentEngine,
output = output,
realtime = realtime,
realtimeModel = realtimeModel,
realtimeVoice = realtimeVoice,
)
Card(
modifier = Modifier.fillMaxWidth(),
@@ -0,0 +1,112 @@
package com.hermesandroid.relay.util
import com.hermesandroid.relay.BuildConfig
import com.hermesandroid.relay.data.Connection
import com.hermesandroid.relay.diagnostics.DiagnosticLogEntry
import com.hermesandroid.relay.diagnostics.DiagnosticSeverity
import com.hermesandroid.relay.diagnostics.DiagnosticsLog
/**
* Pure builder for the GitHub "new issue" prefill derived from a diagnostics
* entry — title, labels, and markdown body. Extracted from the detail dialog so
* the prefill contract is unit-testable without Compose.
*
* Only [DiagnosticSeverity.Error] entries are bug reports. Routine Info/Warning
* log lines ("Testing API connection", probe results, …) were landing on the
* tracker as `[Bug]:` issues with an empty boilerplate body, so non-Error
* entries prefill as `[Diagnostic]:` questions instead; for Info entries the
* caller additionally collects the reporter's expectation before offering the
* link (see [com.hermesandroid.relay.ui.components.DiagnosticDetailDialog]).
*/
object DiagnosticIssuePrefill {
/**
* Cap for the trace embedded in the prefilled GitHub URL so it stays within
* browser limits (the full text is on the clipboard).
*/
private const val MAX_TRACE_FOR_URL = 3000
private const val DEFAULT_WHAT_HAPPENED = "Captured diagnostic from the in-app activity log."
/** `[Bug]:` for Error entries, `[Diagnostic]:` for Info/Warning. */
fun issueTitle(entry: DiagnosticLogEntry): String = when (entry.severity) {
DiagnosticSeverity.Error -> "[Bug]: ${entry.title}"
else -> "[Diagnostic]: ${entry.title}"
}
/**
* `bug` for Error entries; `question` (an existing repo label) for
* Info/Warning so routine diagnostics don't pollute the bug queue.
*/
fun issueLabels(entry: DiagnosticLogEntry): String = when (entry.severity) {
DiagnosticSeverity.Error -> "bug"
else -> "question"
}
/**
* The actual active route for the "Connection mode" line: the role stamped
* on the entry when present, else inferred from the entry URL, else
* `unknown` — never the old unedited "LAN / Tailscale / public TLS / other"
* template text.
*/
fun connectionMode(entry: DiagnosticLogEntry): String =
entry.endpointRole
?: entry.url?.let { Connection.inferRouteRole(it) }
?: "unknown"
/**
* Markdown issue body mirroring the crash-report issue format: environment
* block + the captured entry.
*
* @param expectation the reporter's free-text "What were you expecting to
* happen?" answer collected by the Info-severity pre-flight; when
* non-blank it replaces the boilerplate "What happened" line. Runs
* through the shared diagnostics secret redaction before embedding.
*/
fun issueBody(entry: DiagnosticLogEntry, expectation: String? = null): String {
val trace = (entry.stacktrace ?: entry.detail).orEmpty().let {
if (it.length > MAX_TRACE_FOR_URL) {
it.take(MAX_TRACE_FOR_URL) + "\n… (truncated — full diagnostic copied to your clipboard)"
} else {
it
}
}
val surface = if (BuildConfig.FLAVOR.equals("sideload", ignoreCase = true)) "sideload APK" else "Google Play"
val whatHappened = DiagnosticsLog.redactReportText(expectation)
?.takeIf { it.isNotBlank() }
?: DEFAULT_WHAT_HAPPENED
return buildString {
appendLine(
"> ⚠️ Before submitting: remove any secrets, tokens, real hostnames/IPs, " +
"or personal data from the detail below.",
)
appendLine()
appendLine("### Affected area")
appendLine("Android app")
appendLine()
appendLine("### What happened?")
appendLine(whatHappened)
appendLine()
appendLine("### Environment")
appendLine("- Hermes-Relay version/tag: ${BuildConfig.VERSION_NAME} (code ${BuildConfig.VERSION_CODE})")
appendLine("- Install surface: $surface")
appendLine("- Connection mode: ${connectionMode(entry)}")
appendLine()
appendLine("### Diagnostic")
appendLine("- Title: ${entry.title}")
appendLine("- Category: ${entry.category.label}")
appendLine("- Severity: ${entry.severity.name}")
entry.endpointRole?.let { appendLine("- Route: $it") }
entry.url?.let { appendLine("- URL: $it") }
entry.elapsedMs?.let { appendLine("- Elapsed: ${it}ms") }
if (trace.isNotBlank()) {
appendLine()
appendLine("```")
appendLine(trace)
appendLine("```")
}
appendLine()
append("<sub>Captured by the Hermes-Relay in-app diagnostics log</sub>")
}
}
}
@@ -59,6 +59,37 @@ object ServerAddress {
/** True when [raw] forms a valid http(s) address once normalized. Blank → false. */
fun isValidUserInput(raw: String?): Boolean = parseUserInput(raw) != null
/**
* Advisory for a server address whose host is loopback / any-interface —
* `localhost`, `127.x.x.x`, `::1`, `0.0.0.0`. Such an address works in a
* browser *on the server* but can never reach the server from the phone,
* a recurring source of "Testing API connection" dead-ends. Accepts the
* same scheme-less input [parseUserInput] normalizes. Returns `null` for
* blank, unparseable, or non-loopback addresses. NEVER throws.
*
* Pure helper only — not yet surfaced anywhere; UI wiring is a follow-up.
*/
fun loopbackHostWarning(raw: String): String? {
val trimmed = raw.trim().trimEnd('/')
if (trimmed.isEmpty()) return null
// Bare "::1" never parses without brackets — normalize it directly.
val host = parseUserInput(trimmed)?.host?.lowercase()
?: trimmed.removePrefix("[").removeSuffix("]").lowercase().takeIf { it == "::1" }
?: return null
val loopback = host == "localhost" || host == "::1" || host == "0.0.0.0" || isLoopbackIpv4(host)
return if (loopback) {
"On your phone, localhost points at the phone itself — use the server's LAN IP or Tailscale address."
} else {
null
}
}
private fun isLoopbackIpv4(host: String): Boolean {
val labels = host.split('.')
val parts = labels.mapNotNull { it.toIntOrNull() }
return labels.size == 4 && parts.size == 4 && parts[0] == 127
}
/**
* Inline error for a server-URL / host text field, or `null` when the value
* is acceptable. Blank returns `null` so callers can gate required-ness
@@ -0,0 +1,205 @@
package com.hermesandroid.relay.viewmodel
import com.hermesandroid.relay.network.upstream.models.MessageItem
import kotlinx.coroutines.CancellationException
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Job
import kotlinx.coroutines.delay
import kotlinx.coroutines.launch
/**
* Client-side answer recovery for a sessions-endpoint chat stream that died on
* a transport error while the server kept working (issue #166).
*
* On slow local models + delegating skills, a turn can outlive the phone's SSE
* socket (screen-off / Doze / Wi-Fi power-save kill it long before OkHttp's
* read timeout). Upstream `api_server` behavior on that disconnect: the SSE
* writer dies but the agent run continues in an uncancellable executor thread
* and the FINAL ANSWER IS PERSISTED to the session store. So instead of
* finalizing the turn as an error, this poller re-reads the session transcript
* (the native upstream `/api/sessions/{id}/messages` route — standard-path
* safe, no server changes) until the answer lands.
*
* Cadence: first poll after [Timing.pollIntervalMs], doubling each poll up to
* [Timing.maxPollIntervalMs], for at most [Timing.recoveryWindowMs] of slept
* time. Elapsed time is accumulated from the delays (not wall-clock reads) so
* the loop is virtual-time friendly in tests.
*
* **Anchoring (positional, not text-only).** The pending send, if it landed,
* is the `(priorUserMessageCount + 1)`-th user-role row in the server
* transcript — i.e. the first user row AFTER the ones the client already knew
* about. The candidate answer is the last non-blank assistant row after that
* anchor. Anchoring by position (with a content-match sanity check) — rather
* than a bare `indexOfLast` text search — is what stops a short repeated
* prompt ("yes", "ok", "continue") from matching a STALE identical earlier row
* and adopting a DIFFERENT turn's (static, therefore instantly "stable")
* answer when this send never actually reached the server.
*
* Finish condition: an assistant message that postdates the anchor, is
* non-empty, and is stable across two consecutive polls (the signature also
* folds in the transcript length, so a still-running run that keeps appending
* tool rows after an intermediate assistant message defers the finish).
* Intermediate persisted rows are surfaced through [onIntermediateHistory] as
* they appear — progressive recovery.
*
* Fail-fast: when the anchor can't be established — the transcript still holds
* only the user rows the client already knew about (the POST died before the
* server persisted the turn), or the positional row's content diverges from
* the pending send (the transcript was edited/forked out from under us) — the
* poller never adopts an answer, because a wrong answer is worse than an error.
* It confirms the state across two consecutive polls (guarding against a
* transient read mid-persist) and then gives up with [GiveUpReason.RUN_NOT_FOUND]
* instead of polling to the 30-minute cap, since no answer for this turn can
* ever arrive. An EMPTY transcript carries no information (the history read
* maps fetch failures to an empty list too), so it keeps polling.
*/
class ChatStreamRecovery(
private val scope: CoroutineScope,
private val fetchHistory: suspend () -> List<MessageItem>,
private val timing: Timing = Timing(),
) {
data class Timing(
val pollIntervalMs: Long = 5_000L,
val maxPollIntervalMs: Long = 30_000L,
val recoveryWindowMs: Long = 30L * 60_000L,
)
enum class GiveUpReason {
/**
* The pending send couldn't be located as a new user row after the
* ones the client already knew about — the POST never persisted the
* turn, or the transcript diverged. No answer can arrive; resend.
*/
RUN_NOT_FOUND,
/** The recovery window elapsed without a stable answer. */
TIMED_OUT,
}
/** Whether a poll could establish the anchor for the pending send. */
private sealed interface Anchor {
/** The `(priorUserCount + 1)`-th user row exists and matches. */
data class Found(val index: Int) : Anchor
/**
* The anchor can't be proven: too few user rows (send not persisted)
* or a positional row whose content diverges (edited/forked history).
*/
data object NotEstablished : Anchor
}
private var job: Job? = null
val isActive: Boolean
get() = job?.isActive == true
/**
* Start polling. Exactly one poll loop per instance — a second [start]
* replaces the first. Exactly one terminal callback fires per loop
* ([onRecovered] or [onGaveUp]); cancellation fires none.
*
* @param priorUserMessageCount how many user-role messages the client knew
* existed BEFORE this turn's pending send (excluding the in-flight pair).
* Drives the positional anchor — see the class KDoc.
*/
fun start(
pendingUserText: String,
priorUserMessageCount: Int,
onIntermediateHistory: (List<MessageItem>) -> Unit,
onRecovered: (List<MessageItem>) -> Unit,
onGaveUp: (GiveUpReason) -> Unit,
) {
job?.cancel()
val pending = pendingUserText.trim()
val priorUsers = priorUserMessageCount.coerceAtLeast(0)
job = scope.launch {
var delayMs = timing.pollIntervalMs
var elapsedMs = 0L
var lastSignature: String? = null
var lastSurfacedCount = -1
var unanchoredPolls = 0
while (elapsedMs < timing.recoveryWindowMs) {
delay(delayMs)
elapsedMs += delayMs
delayMs = (delayMs * 2).coerceAtMost(timing.maxPollIntervalMs)
val items = try {
fetchHistory()
} catch (e: CancellationException) {
throw e
} catch (_: Exception) {
continue // unreachable — keep waiting for the network
}
if (items.isEmpty()) continue
when (val anchor = resolveAnchor(items, pending, priorUsers)) {
is Anchor.NotEstablished -> {
// The pending send isn't (verifiably) persisted as a new
// user row. A transient read while the server persists
// could look like this, so require the state to hold
// across two consecutive polls before giving up — but
// never poll to the cap for an answer that can't arrive.
lastSignature = null
if (++unanchoredPolls >= 2) {
onGaveUp(GiveUpReason.RUN_NOT_FOUND)
return@launch
}
}
is Anchor.Found -> {
unanchoredPolls = 0
val signature = answerSignature(items, anchor.index)
if (signature != null && signature == lastSignature) {
onRecovered(items)
return@launch
}
lastSignature = signature
if (items.size != lastSurfacedCount) {
lastSurfacedCount = items.size
onIntermediateHistory(items)
}
}
}
}
onGaveUp(GiveUpReason.TIMED_OUT)
}
}
fun cancel() {
job?.cancel()
job = null
}
/**
* Resolve the positional anchor for the pending send. The send, if it
* landed, is the `(priorUserCount + 1)`-th user-role row; that row must
* ALSO match the pending text (secondary sanity check for edits/forks).
* Any other shape is [Anchor.NotEstablished] — never adopt a guess.
*/
private fun resolveAnchor(
items: List<MessageItem>,
pendingUserText: String,
priorUserCount: Int,
): Anchor {
val userIndices = items.indices.filter { items[it].role == "user" }
if (userIndices.size <= priorUserCount) return Anchor.NotEstablished
val anchorPos = userIndices[priorUserCount]
if (items[anchorPos].contentText?.trim() != pendingUserText) return Anchor.NotEstablished
return Anchor.Found(anchorPos)
}
/**
* Stability signature of the candidate answer: the last non-blank
* assistant message after the anchor, or null while none exists. The
* transcript size is folded in so new rows (tool results of a
* still-running run) change the signature and defer the finish.
*/
private fun answerSignature(items: List<MessageItem>, anchor: Int): String? {
val answer = items.drop(anchor + 1).lastOrNull {
it.role == "assistant" && !it.contentText.isNullOrBlank()
} ?: return null
return "${items.size}|${answer.id}|${answer.contentText?.length}"
}
}
@@ -11,6 +11,7 @@ import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.data.AttachmentState
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.ChatSession
import com.hermesandroid.relay.data.DemoContent
import com.hermesandroid.relay.data.MediaSettings
import com.hermesandroid.relay.data.MediaSettingsRepository
import com.hermesandroid.relay.data.MessageDeliveryStatus
@@ -24,6 +25,9 @@ import com.hermesandroid.relay.data.HermesCard
import com.hermesandroid.relay.data.HermesCardAction
import com.hermesandroid.relay.data.HermesCardField
import com.hermesandroid.relay.data.HermesCardInput
import com.hermesandroid.relay.diagnostics.DiagnosticCategory
import com.hermesandroid.relay.diagnostics.DiagnosticSeverity
import com.hermesandroid.relay.diagnostics.DiagnosticsLog
import com.hermesandroid.relay.network.upstream.ActiveTurnHandle
import com.hermesandroid.relay.network.upstream.GatewayAsk
import com.hermesandroid.relay.network.upstream.GatewayChatClient
@@ -113,6 +117,26 @@ class ChatViewModel : ViewModel() {
*/
private var activeStreamIsGateway = false
private var intentionallyCancelled = false
/**
* Answer-recovery poller for a sessions-endpoint turn whose SSE transport
* died while the server kept running the turn (issue #166) — see
* [ChatStreamRecovery] and [startAnswerRecovery]. At most one per turn;
* null while idle.
*/
private var streamRecovery: ChatStreamRecovery? = null
/** Test seam for the recovery poll cadence — production uses the defaults. */
internal var recoveryTimingOverride: ChatStreamRecovery.Timing? = null
private val _recoveringAnswer = MutableStateFlow(false)
/**
* True while [streamRecovery] polls for a dropped turn's answer — drives
* the "Reconnecting to your answer…" copy on the streaming placeholder.
*/
val recoveringAnswer: StateFlow<Boolean> = _recoveringAnswer.asStateFlow()
private var firstTokenNotified = false
private var toolHistoryJob: Job? = null
private var connectionSwitchJob: Job? = null
@@ -128,6 +152,8 @@ class ChatViewModel : ViewModel() {
private val realtimeAgentModels = mutableMapOf<String, String>()
private val realtimeAgentVoices = mutableMapOf<String, String>()
private val realtimeAgentProgressKeys = mutableMapOf<String, String>()
private val terminalRealtimeAgentTurnIdsLock = Any()
private val terminalRealtimeAgentTurnIds = LinkedHashSet<String>()
private var nextInterfaceContextPrompt: String? = null
// --- Media dependencies (wired via initializeMedia from RelayApp) ---
@@ -361,6 +387,10 @@ class ChatViewModel : ViewModel() {
private val _modelProviders = MutableStateFlow<List<GatewayModelProvider>>(emptyList())
val modelProviders: StateFlow<List<GatewayModelProvider>> = _modelProviders.asStateFlow()
/** True only during an explicit user-requested dynamic model catalog refresh. */
private val _modelOptionsRefreshing = MutableStateFlow(false)
val modelOptionsRefreshing: StateFlow<Boolean> = _modelOptionsRefreshing.asStateFlow()
/** Current gateway model from `model.options`, used when no Android override is active. */
private val _gatewayCurrentModel = MutableStateFlow("")
val gatewayCurrentModel: StateFlow<String> = _gatewayCurrentModel.asStateFlow()
@@ -406,28 +436,37 @@ class ChatViewModel : ViewModel() {
/**
* Refresh the gateway's curated provider/model list (`model.options`).
* Connects the gateway on demand. No-op without a gateway client.
* [refresh] is the explicit upstream refresh path for dynamic/custom-provider catalogs;
* automatic picker opens stay on the cheap cached path.
*/
fun refreshModelOptions() {
fun refreshModelOptions(refresh: Boolean = false) {
val gateway = gatewayClient ?: run {
android.util.Log.i("ChatViewModel", "refreshModelOptions: no gateway client")
if (refresh) _modelOptionsRefreshing.value = false
return
}
if (refresh && _modelOptionsRefreshing.value) return
if (refresh) _modelOptionsRefreshing.value = true
viewModelScope.launch {
gateway.modelOptions().fold(
gateway.modelOptions(refresh = refresh).fold(
onSuccess = {
_modelProviders.value = it.providers
_gatewayCurrentModel.value = it.currentModel
_gatewayCurrentProvider.value = it.currentProvider
android.util.Log.i(
"ChatViewModel",
"model.options: ${it.providers.size} providers, " +
"model.options${if (refresh) " refresh" else ""}: ${it.providers.size} providers, " +
"${it.providers.sumOf { p -> p.models.size }} models, current=${it.currentModel}",
)
},
onFailure = {
android.util.Log.w("ChatViewModel", "model.options failed: ${it.message}")
if (refresh) {
_transientNotice.tryEmit("Couldn't refresh models: ${it.message ?: "unknown error"}")
}
},
)
if (refresh) _modelOptionsRefreshing.value = false
}
}
@@ -1110,6 +1149,23 @@ class ChatViewModel : ViewModel() {
* Default provider returns `null` so a fresh VM (before RelayApp has
* wired the flow) behaves identically to pre-profile-picker installs.
*/
/**
* Demo / Explore mode wiring. [demoModeProvider] reads
* [com.hermesandroid.relay.viewmodel.ConnectionViewModel.isDemoMode];
* [demoChatHandlerProvider] supplies the shared [ChatHandler] carrying
* the canned transcript. Both are needed because in demo there is no API
* client, so [initialize] never runs and this VM's own [chatHandler]
* stays null. Wired unconditionally from RelayApp (not inside the
* client-gated init effect).
*/
private var demoModeProvider: () -> Boolean = { false }
private var demoChatHandlerProvider: () -> ChatHandler? = { chatHandler }
fun setDemoModeWiring(isDemo: () -> Boolean, handler: () -> ChatHandler?) {
demoModeProvider = isDemo
demoChatHandlerProvider = handler
}
private var selectedProfileProvider: () -> Profile? = { null }
private var effectiveProfileProvider: () -> Profile? = { selectedProfileProvider() }
private var displayProfileProvider: () -> Profile? = {
@@ -1677,6 +1733,7 @@ class ChatViewModel : ViewModel() {
intentionallyCancelled = true
activeStream?.cancel()
activeStream = null
cancelAnswerRecovery(settleUi = false)
sessionRefreshJob?.cancel()
_isLoadingSessions.value = false
activeProfileContextKey = null
@@ -1736,6 +1793,7 @@ class ChatViewModel : ViewModel() {
stream.cancel()
}
activeStream = null
cancelAnswerRecovery(settleUi = false)
val loadGeneration = historyLoadGeneration.incrementAndGet()
sessionRefreshGeneration.incrementAndGet()
sessionRefreshJob?.cancel()
@@ -1945,9 +2003,10 @@ class ChatViewModel : ViewModel() {
val client = apiClient ?: return
val handler = chatHandler ?: return
// Cancel any in-flight stream
// Cancel any in-flight stream (and any answer-recovery poller)
activeStream?.cancel()
activeStream = null
cancelAnswerRecovery(settleUi = false)
val loadGeneration = historyLoadGeneration.incrementAndGet()
// Gateway transport: a new chat is a fresh DRAFT with NO session id.
@@ -2035,6 +2094,7 @@ class ChatViewModel : ViewModel() {
val handler = chatHandler ?: return
activeStream?.cancel()
activeStream = null
cancelAnswerRecovery(settleUi = false)
historyLoadGeneration.incrementAndGet()
val slug = name.trim().lowercase().replace(Regex("[^a-z0-9]+"), "-").trim('-').take(24)
val chatId = "t-" + slug.ifBlank { "thread" } + "-" +
@@ -2097,10 +2157,12 @@ class ChatViewModel : ViewModel() {
apiClient ?: return
val handler = chatHandler ?: return
// Cancel any in-flight stream
// Cancel any in-flight stream (and any answer-recovery poller — the
// switched-to session must not receive the old turn's reconcile).
intentionallyCancelled = true
activeStream?.cancel()
activeStream = null
cancelAnswerRecovery(settleUi = false)
val loadGeneration = historyLoadGeneration.incrementAndGet()
handler.setSessionId(sessionId)
@@ -2210,6 +2272,32 @@ class ChatViewModel : ViewModel() {
if (text.isBlank()) return
recordRecentPrompt(text)
// Demo / Explore mode: there is no server, but a silently dead Send
// button reads as broken. Echo the user's text and answer with an
// honest canned notice through the real pipeline — clientOnly
// bubbles, zero network, wiped with the rest of the transcript on
// demo exit (exitDemoMode → clearMessages).
if (demoModeProvider()) {
val demoHandler = demoChatHandlerProvider() ?: return
val now = System.currentTimeMillis()
demoHandler.addUserMessage(
ChatMessage(
id = "demo-composer-user-${java.util.UUID.randomUUID()}",
role = MessageRole.USER,
content = text.trim(),
timestamp = now,
clientOnly = true,
),
)
demoHandler.addUserMessage(
DemoContent.composerReply(
id = "demo-composer-reply-${java.util.UUID.randomUUID()}",
nowMs = now + 1L,
),
)
return
}
val client = apiClient ?: return
val handler = chatHandler ?: return
@@ -2805,6 +2893,183 @@ class ChatViewModel : ViewModel() {
)
}
// === Dropped-stream answer recovery (issue #166) ===
/**
* Terminal side effects every successfully finished turn shares — the
* normal stream completion ([startStream]'s onCompleteCb) and a recovered
* dropped-stream turn ([startAnswerRecovery]) both end here, so recovery
* finalizes with exactly the completion semantics.
*/
private fun finalizeTurnSideEffects(handler: ChatHandler, messageId: String) {
handler.onStreamComplete(messageId)
activeStream = null
_steerableTurn.value = false
_steerNotice.value = null
// Turn over — any blocked ask has been resolved server-side
// (answer, timeout, or interrupt). Timed cards self-collapse;
// an unanswered approval gets a neutral "Resolved" stamp so its
// buttons don't dead-end in "no longer active" notices.
clearPendingAsk(approvalStamp = "Resolved")
// Notify when the turn finished while the app is backgrounded —
// never for cancelled streams; errors end via onErrorCb instead.
maybeNotifyTurnComplete(handler, messageId)
// v0.4.1 polish: auto-return to Hermes-Relay if the bridge
// moved the foreground app during this run. No-op when the
// LLM already called `android_return_to_hermes` itself (in
// that case the tracker's internal flag was cleared by the
// /return_to_hermes dispatch's respond()). See BridgeRunTracker
// KDoc for the full contract.
com.hermesandroid.relay.bridge.BridgeRunTracker.notifyRunCompleted()
}
/**
* Stop any in-flight answer recovery — and ALWAYS settle the handler's
* streaming/turn-status state when a poller was actually running.
*
* When recovery is live there is NO live [activeStream] (it was nulled
* when the poller started), so nothing else fires onStreamComplete /
* onStreamError to clear the "Reconnecting to your answer…" caption and
* the global streaming flag. If abort left them set the chat would wedge
* in streaming mode (dead Stop button, frozen caption) until process
* death (issue #166).
*
* [settleUi] chooses HOW to settle:
* - `true` (a new send about to add its own placeholder) finalizes the
* leftover streaming placeholder into a completed bubble.
* - `false` (abandon paths — session/profile switch, new chat/thread,
* connection switch, user Stop, straggler-completion guard) drops the
* global streaming/turn-status flags SILENTLY: no error badge, and no
* placeholder finalize that could fight a subsequent loadMessageHistory
* or hide the message from cancelStream's Stopped-badge findLast. Those
* callers clear or reload the transcript themselves.
*/
private fun cancelAnswerRecovery(settleUi: Boolean = true) {
val hadRecovery = streamRecovery != null
streamRecovery?.cancel()
streamRecovery = null
_recoveringAnswer.value = false
if (!hadRecovery) return
chatHandler?.let { handler ->
if (settleUi) {
val streaming = handler.messages.value.findLast { it.isStreaming }
if (streaming != null) handler.onStreamComplete(streaming.id)
else handler.clearStreamingStatus()
} else {
handler.clearStreamingStatus()
}
}
}
/**
* Issue #166: on slow-model / delegating-skill turns the phone's SSE
* socket dies (screen-off, Doze, Wi-Fi power-save) long before the server
* finishes — but upstream api_server keeps executing the run after the
* SSE writer dies and PERSISTS the final answer to the session store. So
* a sessions-endpoint transport error must not finalize the turn as an
* error: poll the session transcript (native upstream
* `/api/sessions/{id}/messages` — standard-path safe) until the answer
* lands, reconciling through the normal [ChatHandler.loadMessageHistory]
* path, then finish with the same side effects as a normal completion.
* On cap expiry or a run that never started, fall back to the existing
* error UI.
*/
private fun startAnswerRecovery(
handler: ChatHandler,
sessionId: String,
pendingUserText: String,
placeholderMessageId: String,
cause: String,
) {
cancelAnswerRecovery(settleUi = false)
_recoveringAnswer.value = true
handler.setTurnStatus("Reconnecting to your answer…")
DiagnosticsLog.record(
category = DiagnosticCategory.Api,
severity = DiagnosticSeverity.Warning,
title = "Chat stream dropped — recovering the answer in the background",
detail = cause,
)
// Positional invariant for the anchor (issue #166): how many user-role
// rows the client knew about BEFORE this send. handler.messages already
// holds the in-flight pair (the just-added pending user message + the
// streaming assistant placeholder), so subtract the one pending user
// row. The pending send, once persisted, must land as the
// (priorUserCount+1)-th user row — this stops a short repeated prompt
// ("yes"/"continue") from anchoring on a stale identical earlier row.
val priorUserCount = (
handler.messages.value.count { it.role == MessageRole.USER } - 1
).coerceAtLeast(0)
val recovery = ChatStreamRecovery(
scope = viewModelScope,
fetchHistory = { loadSessionHistory(sessionId) },
timing = recoveryTimingOverride ?: ChatStreamRecovery.Timing(),
)
streamRecovery = recovery
recovery.start(
pendingUserText = pendingUserText,
priorUserMessageCount = priorUserCount,
onIntermediateHistory = { items ->
if (streamRecovery === recovery && handler.currentSessionId.value == sessionId) {
// Progressive recovery: surface already-persisted rows as
// they appear. The reload drops the (never-persisted)
// streaming placeholder, so re-add one — with a stable id —
// to keep the reconnecting indicator alive until the
// answer lands.
handler.loadMessageHistory(items)
handler.addPlaceholderMessage(
ChatMessage(
id = "recovering-$placeholderMessageId",
role = MessageRole.ASSISTANT,
content = "",
timestamp = System.currentTimeMillis(),
isStreaming = true,
agentName = handler.activeAgentName,
)
)
}
},
onRecovered = { items ->
if (streamRecovery === recovery) {
streamRecovery = null
_recoveringAnswer.value = false
if (handler.currentSessionId.value == sessionId) {
// Server-authoritative reconcile — replaces the
// placeholder with the recovered answer (the same
// reload path a normal sessions completion uses).
handler.loadMessageHistory(items)
}
finalizeTurnSideEffects(handler, placeholderMessageId)
refreshSessions()
scheduleTitleReconcile(sessionId)
drainQueue()
}
},
onGaveUp = { reason ->
if (streamRecovery === recovery) {
streamRecovery = null
_recoveringAnswer.value = false
val message = when (reason) {
ChatStreamRecovery.GiveUpReason.RUN_NOT_FOUND ->
"Connection dropped before the server received this message — please resend."
ChatStreamRecovery.GiveUpReason.TIMED_OUT ->
"Lost the connection mid-reply and the answer never arrived — check the server and try again."
}
AppAnalytics.onStreamError()
handler.onStreamError(message)
emitError(Exception(message), context = "send_message")
_queuedMessages.value = emptyList()
// Parity with the sibling stream-error branch: the turn is
// over server-side, so force-deny any still-blocked approval
// card instead of leaving its buttons dead-ended.
clearPendingAsk(approvalStamp = "deny")
}
},
)
}
fun clearQueue() {
_queuedMessages.value = emptyList()
}
@@ -2998,6 +3263,9 @@ class ChatViewModel : ViewModel() {
val trimmed = userText.trim().ifBlank { "Listening..." }
val userMessageId = UUID.randomUUID().toString()
val assistantMessageId = "realtime-agent-${UUID.randomUUID()}"
synchronized(terminalRealtimeAgentTurnIdsLock) {
terminalRealtimeAgentTurnIds.remove(assistantMessageId)
}
realtimeAgentUserMessages[assistantMessageId] = userMessageId
realtimeAgentInputTranscripts[assistantMessageId] = StringBuilder()
realtimeAgentToolCallIds[assistantMessageId] = mutableSetOf()
@@ -3031,6 +3299,94 @@ class ChatViewModel : ViewModel() {
return assistantMessageId
}
/**
* Settle a realtime turn that failed before the relay could emit a
* `voice.error`. Transport submission failures happen below the event
* layer, so without this explicit terminal path the placeholder remains
* streaming forever.
*/
fun failRealtimeAgentTurn(assistantMessageId: String, message: String) {
synchronized(terminalRealtimeAgentTurnIdsLock) {
val handler = chatHandler ?: return
removeRealtimeAgentUserPlaceholder(
handler = handler,
assistantMessageId = assistantMessageId,
)
clearRealtimeAgentTurnTracking(assistantMessageId, quarantine = true)
val existing = handler.messages.value
.firstOrNull { it.id == assistantMessageId }
?.content
?.trim()
.orEmpty()
if (existing.isBlank()) {
handler.replaceMessageContent(assistantMessageId, message)
}
handler.markError(assistantMessageId)
handler.onStreamError(message)
activeStream = null
}
}
/** Locally settle the realtime placeholder when the voice stop action wins. */
fun cancelRealtimeAgentTurnLocally(assistantMessageId: String) {
synchronized(terminalRealtimeAgentTurnIdsLock) {
val handler = chatHandler ?: return
removeRealtimeAgentUserPlaceholder(
handler = handler,
assistantMessageId = assistantMessageId,
)
clearRealtimeAgentTurnTracking(assistantMessageId, quarantine = true)
val existing = handler.messages.value
.firstOrNull { it.id == assistantMessageId }
?.content
?.trim()
.orEmpty()
if (existing.isBlank()) {
handler.replaceMessageContent(assistantMessageId, "Cancelled.")
} else {
handler.markStopped(assistantMessageId)
}
handler.onStreamComplete(assistantMessageId)
activeStream = null
}
}
private fun removeRealtimeAgentUserPlaceholder(
handler: ChatHandler,
assistantMessageId: String,
) {
val userMessageId = realtimeAgentUserMessages[assistantMessageId] ?: return
val current = handler.messages.value
.firstOrNull { it.id == userMessageId }
?.content
?.trim()
if (current == "Listening...") {
handler.removeMessage(userMessageId)
}
}
private fun clearRealtimeAgentTurnTracking(
assistantMessageId: String,
quarantine: Boolean = false,
) {
if (quarantine) synchronized(terminalRealtimeAgentTurnIdsLock) {
terminalRealtimeAgentTurnIds.add(assistantMessageId)
while (terminalRealtimeAgentTurnIds.size > 64) {
val oldest = terminalRealtimeAgentTurnIds.iterator().next()
terminalRealtimeAgentTurnIds.remove(oldest)
}
}
realtimeAgentUserMessages.remove(assistantMessageId)
realtimeAgentInputTranscripts.remove(assistantMessageId)
realtimeAgentProviderBadges.remove(assistantMessageId)
realtimeAgentToolCallIds.remove(assistantMessageId)
realtimeAgentHermesBacked.remove(assistantMessageId)
realtimeAgentProviderIds.remove(assistantMessageId)
realtimeAgentModels.remove(assistantMessageId)
realtimeAgentVoices.remove(assistantMessageId)
realtimeAgentProgressKeys.remove(assistantMessageId)
}
private fun realtimeBadges(
assistantMessageId: String,
provider: String? = null,
@@ -3065,6 +3421,17 @@ class ChatViewModel : ViewModel() {
assistantMessageId: String,
event: RealtimeVoiceEvent,
showDetailedTrace: Boolean = false,
) {
synchronized(terminalRealtimeAgentTurnIdsLock) {
if (assistantMessageId in terminalRealtimeAgentTurnIds) return
applyRealtimeAgentEventLocked(assistantMessageId, event, showDetailedTrace)
}
}
private fun applyRealtimeAgentEventLocked(
assistantMessageId: String,
event: RealtimeVoiceEvent,
showDetailedTrace: Boolean,
) {
val handler = chatHandler ?: return
val hermesSessionId = when {
@@ -3132,7 +3499,16 @@ class ChatViewModel : ViewModel() {
"hermes.tool.delta" -> {
realtimeAgentHermesBacked[assistantMessageId] = true
val name = event.toolName?.takeIf { it.isNotBlank() }
if (!name.isNullOrBlank() && !name.equals("hermes", ignoreCase = true)) {
// `_`-prefixed tools are internal machinery (upstream hides
// them from every tool surface). The gateway streams drafting
// text as a `_thinking` pseudo-tool that only ever emits
// deltas — no tool.completed — so a pill created for it spins
// "running" forever. Its text still feeds the detailed
// thinking trace below; it just never becomes a ToolCall row.
if (!name.isNullOrBlank() &&
!name.equals("hermes", ignoreCase = true) &&
!name.startsWith("_")
) {
val callId = event.toolCallId?.takeIf { it.isNotBlank() } ?: name
val seen = realtimeAgentToolCallIds
.getOrPut(assistantMessageId) { mutableSetOf() }
@@ -3201,6 +3577,10 @@ class ChatViewModel : ViewModel() {
"hermes.tool.started" -> {
realtimeAgentHermesBacked[assistantMessageId] = true
val name = event.toolName?.takeIf { it.isNotBlank() } ?: "hermes"
// Same `_`-internal-tool guard as hermes.tool.delta above —
// defensive here (the relay currently only sends started for
// real tools), since an unpaired started would pin a pill.
if (name.startsWith("_")) return
val callId = event.toolCallId?.takeIf { it.isNotBlank() } ?: name
realtimeAgentToolCallIds
.getOrPut(assistantMessageId) { mutableSetOf() }
@@ -3283,10 +3663,8 @@ class ChatViewModel : ViewModel() {
handler.onThinkingDelta(assistantMessageId, prompt)
}
}
"hermes.run.completed" -> {
// Hermes completion means the tool result is ready; provider narration ends the turn.
Unit
}
// Hermes completion means the tool result is ready; provider narration ends the turn.
"hermes.run.completed" -> Unit
"voice.response.done" -> {
if (realtimeAgentHermesBacked[assistantMessageId] != true) {
val userMessageId = realtimeAgentUserMessages[assistantMessageId]
@@ -3313,39 +3691,40 @@ class ChatViewModel : ViewModel() {
}
}
handler.onStreamComplete(assistantMessageId)
realtimeAgentUserMessages.remove(assistantMessageId)
realtimeAgentInputTranscripts.remove(assistantMessageId)
realtimeAgentProviderBadges.remove(assistantMessageId)
realtimeAgentToolCallIds.remove(assistantMessageId)
realtimeAgentHermesBacked.remove(assistantMessageId)
realtimeAgentProviderIds.remove(assistantMessageId)
realtimeAgentModels.remove(assistantMessageId)
realtimeAgentVoices.remove(assistantMessageId)
clearRealtimeAgentTurnTracking(assistantMessageId)
activeStream = null
}
"hermes.run.cancelled" -> {
handler.replaceMessageContent(assistantMessageId, "Cancelled.")
// Don't clobber a delivered answer: the cancel confirm can
// arrive after the summary already streamed into this bubble
// (chip-cancel racing completion, or a stale confirm). Only a
// bubble with no real content becomes "Cancelled."; anything
// else keeps its text and gets the Stopped badge instead.
removeRealtimeAgentUserPlaceholder(
handler = handler,
assistantMessageId = assistantMessageId,
)
val existingContent = handler.messages.value
.firstOrNull { it.id == assistantMessageId }
?.content
?.trim()
.orEmpty()
if (existingContent.isBlank()) {
handler.replaceMessageContent(assistantMessageId, "Cancelled.")
} else {
handler.markStopped(assistantMessageId)
}
handler.onStreamComplete(assistantMessageId)
realtimeAgentUserMessages.remove(assistantMessageId)
realtimeAgentInputTranscripts.remove(assistantMessageId)
realtimeAgentProviderBadges.remove(assistantMessageId)
realtimeAgentToolCallIds.remove(assistantMessageId)
realtimeAgentHermesBacked.remove(assistantMessageId)
realtimeAgentProviderIds.remove(assistantMessageId)
realtimeAgentModels.remove(assistantMessageId)
realtimeAgentVoices.remove(assistantMessageId)
clearRealtimeAgentTurnTracking(assistantMessageId, quarantine = true)
activeStream = null
}
"voice.error" -> {
removeRealtimeAgentUserPlaceholder(
handler = handler,
assistantMessageId = assistantMessageId,
)
handler.onStreamError(event.message ?: "Realtime agent failed")
realtimeAgentUserMessages.remove(assistantMessageId)
realtimeAgentInputTranscripts.remove(assistantMessageId)
realtimeAgentProviderBadges.remove(assistantMessageId)
realtimeAgentToolCallIds.remove(assistantMessageId)
realtimeAgentHermesBacked.remove(assistantMessageId)
realtimeAgentProviderIds.remove(assistantMessageId)
realtimeAgentModels.remove(assistantMessageId)
realtimeAgentVoices.remove(assistantMessageId)
clearRealtimeAgentTurnTracking(assistantMessageId, quarantine = true)
activeStream = null
}
}
@@ -3530,6 +3909,12 @@ class ChatViewModel : ViewModel() {
// fallback when no profile metadata is available.
handler.activeAgentName = currentAgentDisplayName()
// A new send always aborts any in-flight dropped-stream answer
// recovery — exactly one poller per turn (issue #166). settleUi
// finalizes the previous turn's leftover streaming placeholder so it
// can't pulse forever next to this turn's fresh one.
cancelAnswerRecovery()
// A new turn is starting: clear any leftover cancellation flag so a
// stale `true` from a PRIOR cancelled turn (the flag is sticky — a
// clean gateway cancel never fires onError to consume it) can't make
@@ -3544,6 +3929,12 @@ class ChatViewModel : ViewModel() {
// but updates when the server sends message.started with its own ID.
var currentMessageId = assistantMessageId
// The SSE endpoint this turn actually dispatched on (null on a gateway
// dispatch) — set by dispatchSse below. onErrorCb keys the dropped-
// stream answer recovery (issue #166) on "sessions": the other
// endpoints keep their existing error behavior.
var dispatchedSseEndpoint: String? = null
// Show placeholder "thinking" message immediately — filled when first delta arrives
handler.addPlaceholderMessage(
ChatMessage(
@@ -3595,28 +3986,13 @@ class ChatViewModel : ViewModel() {
handler.onTurnComplete(currentMessageId)
}
val onCompleteCb = {
handler.onStreamComplete(currentMessageId)
// Double-finalize guard: if a straggler completion arrives while
// the answer-recovery poller is running, the normal completion
// wins — stop the poller before finalizing so the turn can't
// finish twice.
cancelAnswerRecovery(settleUi = false)
finalizeTurnSideEffects(handler, currentMessageId)
AppAnalytics.onStreamComplete(lastInputTokens, lastOutputTokens)
activeStream = null
_steerableTurn.value = false
_steerNotice.value = null
// Turn over — any blocked ask has been resolved server-side
// (answer, timeout, or interrupt). Timed cards self-collapse;
// an unanswered approval gets a neutral "Resolved" stamp so its
// buttons don't dead-end in "no longer active" notices.
clearPendingAsk(approvalStamp = "Resolved")
// Notify when the turn finished while the app is backgrounded —
// never for cancelled streams; errors end via onErrorCb instead.
maybeNotifyTurnComplete(handler, currentMessageId)
// v0.4.1 polish: auto-return to Hermes-Relay if the bridge
// moved the foreground app during this run. No-op when the
// LLM already called `android_return_to_hermes` itself (in
// that case the tracker's internal flag was cleared by the
// /return_to_hermes dispatch's respond()). See BridgeRunTracker
// KDoc for the full contract.
com.hermesandroid.relay.bridge.BridgeRunTracker.notifyRunCompleted()
// Command catalog rides the now-live socket after the first real
// gateway turn — never a cold /api/ws open at composition.
@@ -3715,6 +4091,7 @@ class ChatViewModel : ViewModel() {
}
}
val onErrorCb = { errorMsg: String ->
val errorSessionId = handler.currentSessionId.value
if (intentionallyCancelled) {
intentionallyCancelled = false
// Cancellation (user Stop / session switch): suppress the
@@ -3723,6 +4100,33 @@ class ChatViewModel : ViewModel() {
// button if a cancel and a transport error race.
handler.messages.value.findLast { it.isStreaming }
?.let { handler.onStreamComplete(it.id) }
activeStream = null
_queuedMessages.value = emptyList()
_steerableTurn.value = false
_steerNotice.value = null
clearPendingAsk(approvalStamp = "deny")
} else if (
dispatchedSseEndpoint == "sessions" &&
errorSessionId != null &&
HermesApiClient.isTransportStreamError(errorMsg)
) {
// Issue #166: a transport drop on the sessions endpoint does
// NOT mean the turn failed — upstream api_server keeps running
// it and persists the final answer. Don't finalize as an
// error; recover the answer by polling the transcript. The
// send queue is deliberately KEPT: a successful recovery
// drains it exactly like a normal completion; give-up flushes
// it in the error fallback.
startAnswerRecovery(
handler = handler,
sessionId = errorSessionId,
pendingUserText = message,
placeholderMessageId = currentMessageId,
cause = errorMsg,
)
activeStream = null
_steerableTurn.value = false
_steerNotice.value = null
} else {
AppAnalytics.onStreamError()
handler.onStreamError(errorMsg)
@@ -3735,24 +4139,25 @@ class ChatViewModel : ViewModel() {
// switch, watchdog timeout) AFTER the server already finished it
// — reload history so the completed answer still surfaces
// instead of stranding the turn on its partial/errored state.
val sid = handler.currentSessionId.value
if (sid != null && (streamingEndpoint == "sessions" || streamingEndpoint == "gateway")) {
if (errorSessionId != null &&
(streamingEndpoint == "sessions" || streamingEndpoint == "gateway")
) {
viewModelScope.launch {
runCatching {
// Profile-aware read — see onCompleteCb: a bare
// getMessages 404s for a non-default-profile session
// and silently empties the transcript.
val serverMessages = loadSessionHistory(sid)
val serverMessages = loadSessionHistory(errorSessionId)
handler.loadMessageHistory(serverMessages)
}
}
}
activeStream = null
_queuedMessages.value = emptyList()
_steerableTurn.value = false
_steerNotice.value = null
clearPendingAsk(approvalStamp = "deny")
}
activeStream = null
_queuedMessages.value = emptyList()
_steerableTurn.value = false
_steerNotice.value = null
clearPendingAsk(approvalStamp = "deny")
}
// === v0.4.1 voice-intent + v0.7.x card-dispatch session sync ===
@@ -3843,6 +4248,7 @@ class ChatViewModel : ViewModel() {
// branch's per-turn fallback (gateway unreachable / not the resolved
// transport). Warns once per dispatch about any attachment it can't carry.
fun dispatchSse(endpoint: String): ActiveTurnHandle {
dispatchedSseEndpoint = endpoint
warnIfAttachmentsDropped(endpoint)
return when (endpoint) {
"runs" -> client.sendRunStream(
@@ -3925,9 +4331,33 @@ class ChatViewModel : ViewModel() {
// instead, so force any turn that has a per-turn interface context
// (set only by sendVoiceMessage) onto SSE. resolveSseFallback picks the
// best available SSE route.
//
// Synthetic sync messages (voice intents / card dispatches / provider-
// answered realtime turns) have the same gateway limitation — prompt.submit
// can't carry them — but on a gateway-primary phone "leave them for the
// next SSE turn" means *never*: the default transport is the gateway, so
// unsynced traces would defer forever and the agent never learns what
// happened in realtime voice. Drain them by forcing this one turn onto
// the sessions SSE route, but ONLY when that is strictly safe:
// - an existing session id + the sessions fallback route (a stateless
// completions/runs detour would drop THIS turn from the transcript
// to save a trace — worse than deferring), and
// - the default profile (a non-default profile's gateway session lives
// in that profile's own state.db, which the shared api_server surface
// can't see — the sessions POST would 404 and fail the user's turn).
// Cost when it fires: one turn without live gateway thinking. The synced
// traces persist server-side, so this happens at most once per batch.
val sseDrainEndpoint = resolveSseFallback(handler)
val forceSseForTraceDrain =
voiceIntentMessages != null &&
streamingEndpoint == "gateway" &&
profileName == null &&
sseDrainEndpoint == "sessions"
val effectiveEndpoint =
if (interfaceContextPrompt != null && streamingEndpoint == "gateway") {
resolveSseFallback(handler)
if ((interfaceContextPrompt != null || forceSseForTraceDrain) &&
streamingEndpoint == "gateway"
) {
sseDrainEndpoint
} else {
streamingEndpoint
}
@@ -4043,8 +4473,15 @@ class ChatViewModel : ViewModel() {
// point. Guarded per-stream so we only do the work when the
// corresponding synthetic messages were actually sent.
// Gateway turns can't carry synthetic messages (prompt.submit is
// bare text) — leave traces unsynced so the next SSE turn sends them.
if (voiceIntentMessages != null && streamingEndpoint != "gateway") {
// bare text) — leave traces unsynced so a later SSE turn (or the
// forced trace-drain above) sends them. Checked against the route
// this turn actually DISPATCHED on (effectiveEndpoint), not the
// configured transport: a gateway-configured turn forced onto SSE
// (voice interface context, trace drain) did carry the synthetic
// messages, and skipping the mark there re-sent them every turn.
// The async gateway preflight-failure fallback stays conservative:
// its traces are marked on the NEXT turn (at-least-once delivery).
if (voiceIntentMessages != null && effectiveEndpoint != "gateway") {
if (hasVoiceIntents) handler.markVoiceIntentsSynced()
if (hasCardDispatches) handler.markCardDispatchesSynced()
if (hasRealtimeTurns) handler.markRealtimeTurnsSynced()
@@ -4098,6 +4535,10 @@ class ChatViewModel : ViewModel() {
fun cancelStream() {
intentionallyCancelled = true
// User Stop also aborts a dropped-stream answer recovery. settleUi
// false: the Stopped-badge block below finalizes the placeholder
// itself (completing it here first would hide it from findLast).
cancelAnswerRecovery(settleUi = false)
activeStream?.cancel()
activeStream = null
_queuedMessages.value = emptyList()
@@ -4633,6 +5074,10 @@ class ChatViewModel : ViewModel() {
override fun onCleared() {
super.onCleared()
activeStream?.cancel()
// viewModelScope teardown already cancels the poller job; this just
// drops the reference symmetrically.
streamRecovery?.cancel()
streamRecovery = null
}
private fun appendRealtimeThinkingStatus(
File diff suppressed because it is too large Load Diff
@@ -47,9 +47,36 @@ object RealtimeTurnSyncBuilder {
return if (provenance.isBlank()) {
trace.assistantText
} else {
"${trace.assistantText}\n\n[Realtime Agent provider-native voice turn: $provenance]"
"${trace.assistantText}\n\n$PROVENANCE_PREFIX$provenance]"
}
}
/**
* Detect + strip the provenance marker [buildAssistantContent] appends to
* a synced provider-answered turn. Returns the assistant text with the
* marker removed, or null when no marker is present.
*
* Used by [com.hermesandroid.relay.network.upstream.ChatHandler.loadMessageHistory]
* so a synced turn coming back in server history renders with the quiet
* "Realtime Agent" badge instead of raw bracket noise — and so its
* superseded local clientOnly bubble can be dropped instead of showing
* the exchange twice. The marker must be the FINAL block of the content
* (a single bracket line, no embedded newline/bracket) so ordinary
* assistant prose that merely mentions the phrase is never stripped.
*/
fun stripProvenanceMarker(content: String): String? {
val trimmed = content.trimEnd()
if (!trimmed.endsWith("]")) return null
val idx = trimmed.lastIndexOf("\n\n$PROVENANCE_PREFIX")
if (idx < 0) return null
val inner = trimmed.substring(
idx + 2 + PROVENANCE_PREFIX.length,
trimmed.length - 1,
)
if ('\n' in inner || ']' in inner) return null
return trimmed.substring(0, idx).trimEnd()
}
private const val PROVENANCE_PREFIX = "[Realtime Agent provider-native voice turn: "
private const val MAX_CONTENT_CHARS = 4_000
}
@@ -109,4 +109,24 @@ class DemoContentTest {
// the demo looks the same every launch and the content is testable.
assertEquals(DemoContent.transcript(), DemoContent.transcript())
}
@Test
fun composerReplyFollowsTheDemoContentContract() {
// The canned reply for a message typed inside demo mode (composer
// no-op polish) must obey the same rules as the transcript: an
// honest offline notice, clientOnly, terminal, zero network.
val reply = DemoContent.composerReply(id = "demo-composer-reply-test", nowMs = 123L)
assertEquals("demo-composer-reply-test", reply.id)
assertEquals(123L, reply.timestamp)
assertEquals(MessageRole.ASSISTANT, reply.role)
assertTrue("composer reply must be clientOnly", reply.clientOnly)
assertFalse("composer reply must be terminal", reply.isStreaming)
assertTrue("composer reply carries the Demo badge", reply.badges.contains("Demo"))
assertTrue(
"composer reply should say it can't answer offline",
reply.content.contains("demo", ignoreCase = true),
)
assertTrue("composer reply has no attachments", reply.attachments.isEmpty())
assertEquals(DemoContent.DEMO_AGENT_NAME, reply.agentName)
}
}
@@ -0,0 +1,68 @@
package com.hermesandroid.relay.data
import androidx.datastore.core.DataStore
import androidx.datastore.preferences.core.PreferenceDataStoreFactory
import androidx.datastore.preferences.core.Preferences
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.cancel
import kotlinx.coroutines.flow.first
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Before
import org.junit.Rule
import org.junit.Test
import org.junit.rules.TemporaryFolder
import java.io.File
class VoicePreferencesRepositoryTest {
@get:Rule
val tempFolder = TemporaryFolder()
private lateinit var scope: CoroutineScope
private lateinit var dataStore: DataStore<Preferences>
private lateinit var repository: VoicePreferencesRepository
@Before
fun setUp() {
scope = CoroutineScope(Dispatchers.IO + Job())
val file: File = tempFolder.newFile("voice_preferences_test.preferences_pb")
if (file.exists()) file.delete()
dataStore = PreferenceDataStoreFactory.create(
scope = scope,
produceFile = { file },
)
repository = VoicePreferencesRepository(dataStore)
}
@After
fun tearDown() {
scope.cancel()
}
@Test
fun realtimeSelectionPersistsPerConnectionAndProfile() = runTest {
repository.setActiveScope("connection-a", "coder")
repository.setRealtimeSelection(
model = " grok-voice-think-fast-1.0 ",
voice = " leo ",
)
var settings = repository.settings.first()
assertEquals("grok-voice-think-fast-1.0", settings.realtimeModel)
assertEquals("leo", settings.realtimeVoice)
repository.setActiveScope("connection-b", "coder")
settings = repository.settings.first()
assertEquals("", settings.realtimeModel)
assertEquals("", settings.realtimeVoice)
repository.setActiveScope("connection-a", "coder")
settings = repository.settings.first()
assertEquals("grok-voice-think-fast-1.0", settings.realtimeModel)
assertEquals("leo", settings.realtimeVoice)
}
}
@@ -0,0 +1,39 @@
package com.hermesandroid.relay.network.relay
import org.junit.Assert.assertNotNull
import org.junit.Assert.assertNull
import org.junit.Test
/**
* Guards the relay-socket half of the #131 "Invalid URL host" crash class.
*
* [ConnectionManager.doConnectInternal] builds an OkHttp `Request` on a
* background coroutine, so a malformed relay URL (from a corrupt or hand-edited
* pairing payload) used to let `Request.Builder.url()` throw
* `IllegalArgumentException`, which — uncaught on the IO dispatcher — crashed
* the app (observed on Play as an `okhttp3.HttpUrl$Builder.parse` crash).
* [buildRelayRequestOrNull] must return null for such URLs so the connect path
* fails gracefully (Disconnected + diagnostic) instead of crashing.
*/
class ConnectionManagerUrlGuardTest {
@Test
fun `valid ws and wss urls build a request`() {
assertNotNull(buildRelayRequestOrNull("wss://relay.example.com:8767"))
assertNotNull(buildRelayRequestOrNull("ws://192.168.1.10:8767/path"))
assertNotNull(buildRelayRequestOrNull("wss://host.tailnet.ts.net"))
}
@Test
fun `malformed relay urls return null instead of throwing`() {
// Each makes OkHttp's Request.Builder.url() throw IllegalArgumentException
// ("Invalid URL host"): empty host, and a space inside the host.
for (bad in listOf(
"wss://",
"wss://in valid host:8767",
"wss://a b",
)) {
assertNull("expected null for malformed relay url '$bad'", buildRelayRequestOrNull(bad))
}
}
}
@@ -1350,6 +1350,106 @@ class ChatHandlerTest {
assertEquals("boom", handler.error.value)
}
// --- loadMessageHistory: synced realtime-turn provenance (marker → badge) ---
@Test
fun loadMessageHistory_stripsRealtimeProvenanceMarkerIntoBadge() {
// A provider-answered realtime turn synced by RealtimeTurnSyncBuilder
// comes back in server history with a trailing provenance marker. The
// reload must strip the bracket noise and restore the same quiet
// "Realtime Agent" badge a live turn gets.
handler.loadMessageHistory(
listOf(
MessageItem(
id = "1",
role = "assistant",
content = JsonPrimitive(
"It syncs vault metadata.\n\n" +
"[Realtime Agent provider-native voice turn: " +
"provider=xai_realtime, model=grok-voice-latest]",
),
),
),
)
val msg = handler.messages.value.single()
assertEquals("It syncs vault metadata.", msg.content)
assertTrue(msg.badges.contains("Realtime Agent"))
}
@Test
fun loadMessageHistory_noBadgeWithoutProvenanceMarker() {
handler.loadMessageHistory(
listOf(
MessageItem(id = "1", role = "assistant", content = JsonPrimitive("Plain reply")),
),
)
val msg = handler.messages.value.single()
assertEquals("Plain reply", msg.content)
assertFalse(msg.badges.contains("Realtime Agent"))
}
@Test
fun loadMessageHistory_dropsSyncedRealtimeOrphanSupersededByServerCopy() {
// Once a provider-answered turn's synced copy exists in the server
// transcript, the pre-sync local clientOnly bubble is redundant —
// preserving both would render the exchange twice.
handler.onTextDelta("realtime-agent-1", "It syncs vault metadata.")
handler.attachRealtimeTurnTrace(
"realtime-agent-1",
RealtimeTurnTrace(
userText = "What does it do?",
assistantText = "It syncs vault metadata.",
provider = "xai_realtime",
),
)
handler.markRealtimeTurnsSynced()
handler.loadMessageHistory(
listOf(
MessageItem(id = "10", role = "user", content = JsonPrimitive("What does it do?")),
MessageItem(
id = "11",
role = "assistant",
content = JsonPrimitive(
"It syncs vault metadata.\n\n" +
"[Realtime Agent provider-native voice turn: provider=xai_realtime]",
),
),
),
)
val msgs = handler.messages.value
// Only the server pair remains — the local orphan was dropped.
assertEquals(2, msgs.size)
assertTrue(msgs.none { it.id == "realtime-agent-1" })
val serverCopy = msgs.single { it.id == "11" }
assertEquals("It syncs vault metadata.", serverCopy.content)
assertTrue(serverCopy.badges.contains("Realtime Agent"))
}
@Test
fun loadMessageHistory_preservesUnsyncedRealtimeOrphan() {
// An UNSYNCED trace is still the only record of the turn — it must
// survive the reload even though it has no server row.
handler.onTextDelta("realtime-agent-1", "spoken answer")
handler.attachRealtimeTurnTrace(
"realtime-agent-1",
RealtimeTurnTrace(userText = "hi", assistantText = "spoken answer"),
)
handler.loadMessageHistory(
listOf(
MessageItem(id = "10", role = "user", content = JsonPrimitive("unrelated")),
),
)
val orphan = handler.messages.value.single { it.id == "realtime-agent-1" }
assertEquals("spoken answer", orphan.content)
assertFalse(orphan.realtimeTurn!!.syncedToServer)
}
// --- Helper ---
private fun createUserMessage(id: String, content: String) = ChatMessage(
@@ -1,5 +1,7 @@
package com.hermesandroid.relay.network.upstream
import com.hermesandroid.relay.network.upstream.models.SessionPruneFilters
import com.hermesandroid.relay.network.upstream.models.SessionPrunePreview
import kotlinx.coroutines.test.runTest
import kotlinx.serialization.json.JsonArray
import kotlinx.serialization.json.Json
@@ -57,6 +59,31 @@ class DashboardApiClientTest {
assertEquals("0.16.0", status.version)
}
@Test
fun getModelOptions_alwaysRequestsUnconfiguredProviders() = runTest {
// HRUI-022: newer upstream hides unconfigured provider skeleton rows
// unless the client opts in — without include_unconfigured=1 the
// Manage picker loses its Keys-setup affordance. Both the cached and
// the refresh path must carry the opt-in.
val body = """{"providers": []}"""
server.enqueue(MockResponse().setHeader("Content-Type", "application/json").setBody(body))
server.enqueue(MockResponse().setHeader("Content-Type", "application/json").setBody(body))
val client = DashboardApiClient(baseUrl = server.url("/").toString())
client.getModelOptions().getOrThrow()
val bare = server.takeRequest().requestUrl!!
assertEquals("/api/model/options", bare.encodedPath)
assertEquals("1", bare.queryParameter("include_unconfigured"))
assertEquals(null, bare.queryParameter("refresh"))
client.getModelOptions(refresh = true).getOrThrow()
val refreshed = server.takeRequest().requestUrl!!
assertEquals("/api/model/options", refreshed.encodedPath)
assertEquals("1", refreshed.queryParameter("include_unconfigured"))
assertEquals("1", refreshed.queryParameter("refresh"))
}
@Test
fun currentSession_onConnectionAbort_returnsFailure_doesNotThrow() = runTest {
// Reproduces the crash: a stale pooled connection aborting mid-flight
@@ -703,6 +730,213 @@ class DashboardApiClientTest {
assertEquals("/api/config/schema", server.takeRequest().path)
}
@Test
fun previewSessionPrune_postsDryRunAndParsesPreview() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody(
"""
{"ok":true,"removed":0,"matched":2,
"oldest_started_at":1000.5,"newest_started_at":2000.5,
"sessions":[
{"id":"sess-old","source":"phone","title":"Old plan","model":"claude-opus-4-8","started_at":1000.5,"message_count":3},
{"id":"sess-new","source":"phone","started_at":2000.5,"message_count":1}
]}
""".trimIndent(),
),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
val preview = client.previewSessionPrune(
SessionPruneFilters(olderThanDays = 30.0, source = "phone", profile = "mizu"),
).getOrThrow()
val request = server.takeRequest()
assertEquals("POST", request.method)
assertEquals("/api/sessions/prune", request.path)
val body = request.body.readUtf8()
// The preview MUST be a dry run — this call may never delete.
assertTrue(body.contains(""""dry_run":true"""))
assertTrue(body.contains(""""older_than_days":30.0"""))
assertTrue(body.contains(""""source":"phone""""))
assertTrue(body.contains(""""profile":"mizu""""))
assertEquals(2, preview.matched)
assertEquals(1000.5, preview.oldestStartedAt!!, 0.001)
assertEquals(2000.5, preview.newestStartedAt!!, 0.001)
assertEquals("sess-old", preview.sessions[0].id)
assertEquals(3, preview.sessions[0].messageCount)
assertEquals("Old plan", preview.sessions[0].title)
}
@Test
fun previewSessionPrune_bareFiltersOmitOptionalFields() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"ok":true,"removed":0,"matched":0,"sessions":[]}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
client.previewSessionPrune(SessionPruneFilters()).getOrThrow()
val body = server.takeRequest().body.readUtf8()
// A bare prune sends only dry_run; upstream then applies its own
// implicit ended-more-than-90-days-ago cutoff.
assertTrue(body.contains(""""dry_run":true"""))
assertFalse(body.contains("older_than_days"))
assertFalse(body.contains("source"))
assertFalse(body.contains("profile"))
assertFalse(body.contains("include_archived"))
}
@Test
fun pruneSessions_appliesWithDryRunFalseAndParsesRemoved() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"ok":true,"removed":2}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
val filters = SessionPruneFilters(olderThanDays = 30.0, source = "phone")
val preview = SessionPrunePreview(matched = 2)
val result = client.pruneSessions(filters, confirmedPreview = preview).getOrThrow()
val request = server.takeRequest()
assertEquals("POST", request.method)
assertEquals("/api/sessions/prune", request.path)
val body = request.body.readUtf8()
assertTrue(body.contains(""""dry_run":false"""))
assertTrue(body.contains(""""older_than_days":30.0"""))
assertEquals(2, result.removed)
}
@Test
fun pruneSessions_skipsServerCallWhenPreviewMatchedNothing() = runTest {
val client = DashboardApiClient(baseUrl = server.url("/").toString())
val result = client.pruneSessions(
SessionPruneFilters(olderThanDays = 30.0),
confirmedPreview = SessionPrunePreview(matched = 0),
).getOrThrow()
// Nothing matched at preview time → nothing to delete. The client must
// not fire the destructive POST at all (sessions that aged in after
// the preview are not covered by what the user confirmed).
assertEquals(0, server.requestCount)
assertEquals(0, result.removed)
}
@Test
fun exportSession_getsServerOwnedArchiveJsonScopedToProfile() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"id":"sess-old","messages":[]}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
val exported = client.exportSession("sess-old", profile = "mizu").getOrThrow()
val request = server.takeRequest()
assertEquals("GET", request.method)
assertEquals("/api/sessions/sess-old/export", request.requestUrl!!.encodedPath)
assertEquals("mizu", request.requestUrl!!.queryParameter("profile"))
assertEquals("sess-old", exported["id"].toString().trim('"'))
}
@Test
fun setSessionArchived_patchesArchivedScopedToProfile() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"ok":true,"title":"Old plan","archived":true}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
client.setSessionArchived("sess-old", archived = true, profile = "mizu").getOrThrow()
val request = server.takeRequest()
assertEquals("PATCH", request.method)
// Current upstream reads profile from the PATCH body (SessionRename
// model); the query param rides along for builds that scoped by query.
assertEquals("/api/sessions/sess-old", request.requestUrl!!.encodedPath)
assertEquals("mizu", request.requestUrl!!.queryParameter("profile"))
val body = request.body.readUtf8()
assertTrue(body.contains(""""archived":true"""))
assertTrue(body.contains(""""profile":"mizu""""))
}
@Test
fun setSessionArchived_omitsProfileForDefaultSelection() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"ok":true,"title":"","archived":false}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
client.setSessionArchived("sess-old", archived = false, profile = null).getOrThrow()
val request = server.takeRequest()
assertEquals(null, request.requestUrl!!.queryParameter("profile"))
val body = request.body.readUtf8()
assertTrue(body.contains(""""archived":false"""))
assertFalse(body.contains("profile"))
}
@Test
fun renameSession_carriesProfileInBodyAndQuery() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"ok":true,"title":"New title"}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
client.renameSession("sess-a", title = "New title", profile = "mizu").getOrThrow()
val request = server.takeRequest()
assertEquals("PATCH", request.method)
assertEquals("/api/sessions/sess-a", request.requestUrl!!.encodedPath)
assertEquals("mizu", request.requestUrl!!.queryParameter("profile"))
val body = request.body.readUtf8()
assertTrue(body.contains(""""title":"New title""""))
// Current upstream reads profile from the PATCH body, not the query.
assertTrue(body.contains(""""profile":"mizu""""))
}
@Test
fun listSessions_passesArchivedFilterThrough() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"sessions":[],"total":0,"limit":50,"offset":0}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
client.listSessions(archived = "only").getOrThrow()
val url = server.takeRequest().requestUrl!!
assertEquals("only", url.queryParameter("archived"))
}
@Test
fun listSessions_omitsArchivedParamByDefault() = runTest {
server.enqueue(
MockResponse()
.setHeader("Content-Type", "application/json")
.setBody("""{"sessions":[],"total":0,"limit":50,"offset":0}"""),
)
val client = DashboardApiClient(baseUrl = server.url("/").toString())
client.listSessions().getOrThrow()
// Default stays upstream's default (exclude) with no param, so older
// hosts that predate the archived filter see an unchanged request.
assertEquals(null, server.takeRequest().requestUrl!!.queryParameter("archived"))
}
@Test
fun parseChatDisplaySettings_mapsToolProgressNoneToOff() {
val root = Json.parseToJsonElement(
@@ -63,6 +63,14 @@ class GatewayClientHarness(
/** Methods answered with JSON-RPC -32601 — exercises the legacy-name fallback. */
val methodNotFound: MutableSet<String> = ConcurrentHashMap.newKeySet()
/** One withheld JSON-RPC ack, capturable for delayed release via [releaseAck]. */
class PendingAck(val ws: WebSocket, val method: String, val id: Long)
/** Methods whose ack is WITHHELD (queued in [pendingAcks]) instead of auto-answered —
* models upstream's fire-and-forget `prompt.submit`, whose ack can trail the turn. */
val suppressAckMethods: MutableSet<String> = ConcurrentHashMap.newKeySet()
val pendingAcks = LinkedBlockingQueue<PendingAck>()
private val wsListener = object : WebSocketListener() {
override fun onOpen(webSocket: WebSocket, response: okhttp3.Response) {
serverSockets.add(webSocket)
@@ -77,6 +85,10 @@ class GatewayClientHarness(
val params = frame["params"] as? JsonObject ?: JsonObject(emptyMap())
rpcLog.add(method to params)
if (!autoRespondEnabled) return
if (method in suppressAckMethods) {
pendingAcks.add(PendingAck(webSocket, method, id.toLong()))
return
}
if (method in methodNotFound) {
webSocket.send(
buildJsonObject {
@@ -125,6 +137,26 @@ class GatewayClientHarness(
json.parseToJsonElement("""[["/help","Show help"],["/model","Pick model"]]"""),
)
}
"model.options" -> buildJsonObject {
put("model", "gpt-5.5")
put("provider", "openai")
put(
"providers",
json.parseToJsonElement(
"""
[
{
"slug": "openai",
"name": "OpenAI",
"models": ["gpt-5.5"],
"is_current": true,
"authenticated": true
}
]
""".trimIndent(),
),
)
}
"config.get" -> when ((params["key"] as? JsonPrimitive)?.contentOrNull) {
"reasoning" -> buildJsonObject {
put("value", reasoningEffort)
@@ -210,6 +242,20 @@ class GatewayClientHarness(
error("rpc $method never arrived; saw ${rpcLog.map { it.first }}")
}
fun awaitPendingAck(): PendingAck =
pendingAcks.poll(5, TimeUnit.SECONDS) ?: error("suppressed ack never captured")
/** Release a withheld ack with a generic success result. */
fun releaseAck(ack: PendingAck) {
ack.ws.send(
buildJsonObject {
put("jsonrpc", "2.0")
put("id", ack.id)
put("result", buildJsonObject { put("ok", true) })
}.toString(),
)
}
/** Waits until [method] has been seen at least [count] times; returns the params in arrival order. */
fun awaitRpcCount(method: String, count: Int): List<JsonObject> {
val deadline = System.currentTimeMillis() + 5_000
@@ -277,24 +323,48 @@ class GatewayChatClientTest {
)
}
private fun buildClient(
rpcTimeoutMs: Long = 15_000L,
promptSubmitTimeoutMs: Long = 1_800_000L,
turnIdleTimeoutMs: Long = 180_000L,
) = GatewayChatClient(
initialDashboardClient = DashboardApiClient(
baseUrl = harness.server.url("/").toString().trimEnd('/'),
okHttpClient = OkHttpClient(),
),
okHttpClient = OkHttpClient(),
callbackDispatcher = { it() },
onGatewayUnsupported = { unsupportedMarked = true },
scope = scope,
// Keep the mid-turn reconnect window short so `failed rejoin`
// surfaces its error well within the test's await budget.
midTurnRejoinWindowMs = 3_000L,
rpcTimeoutMs = rpcTimeoutMs,
promptSubmitTimeoutMs = promptSubmitTimeoutMs,
turnIdleTimeoutMs = turnIdleTimeoutMs,
)
/**
* Swap in a client with shortened timeout seams. Mints a FRESH scope:
* shutdown() cancels the scope's Job, and the replacement client must
* still be able to launch its sendTurn coroutines.
*/
private fun rebuildClient(
rpcTimeoutMs: Long = 15_000L,
promptSubmitTimeoutMs: Long = 1_800_000L,
turnIdleTimeoutMs: Long = 180_000L,
) {
client.shutdown()
scope = CoroutineScope(SupervisorJob() + Dispatchers.IO)
client = buildClient(rpcTimeoutMs, promptSubmitTimeoutMs, turnIdleTimeoutMs)
}
@Before
fun setUp() {
harness = GatewayClientHarness()
scope = CoroutineScope(SupervisorJob() + Dispatchers.IO)
unsupportedMarked = false
client = GatewayChatClient(
initialDashboardClient = DashboardApiClient(
baseUrl = harness.server.url("/").toString().trimEnd('/'),
okHttpClient = OkHttpClient(),
),
okHttpClient = OkHttpClient(),
callbackDispatcher = { it() },
onGatewayUnsupported = { unsupportedMarked = true },
scope = scope,
// Keep the mid-turn reconnect window short so `failed rejoin`
// surfaces its error well within the test's await budget.
midTurnRejoinWindowMs = 3_000L,
)
client = buildClient()
}
@After
@@ -354,6 +424,19 @@ class GatewayChatClientTest {
assertTrue(r.preflightFailures.isEmpty())
}
@Test
fun `model options refresh flag rides gateway rpc only on explicit refresh`() = runBlocking {
val normal = client.modelOptions().getOrThrow()
val normalParams = harness.awaitRpc("model.options")
assertEquals("gpt-5.5", normal.currentModel)
assertFalse((normalParams["refresh"] as? JsonPrimitive)?.booleanOrNull == true)
val refreshed = client.modelOptions(refresh = true).getOrThrow()
val refreshParams = harness.awaitRpcCount("model.options", 2).last()
assertEquals("openai", refreshed.currentProvider)
assertTrue((refreshParams["refresh"] as? JsonPrimitive)?.booleanOrNull == true)
}
@Test
fun `foreign session events are dropped`() {
val r = Recorder()
@@ -990,4 +1073,105 @@ class GatewayChatClientTest {
val submit = harness.awaitRpc("prompt.submit")
assertFalse(submit.containsKey("truncate_before_user_ordinal"))
}
// --- HRUI-016: long / fire-and-forget prompt.submit ack semantics.
// Upstream treats prompt.submit as a long-running RPC (desktop passes a
// 30-min PROMPT_SUBMIT_REQUEST_TIMEOUT_MS at every call site) because the
// ack can trail a MoA/deep-reasoning/tool-heavy turn by minutes. A short
// ack timeout used to preflight-fail into the SSE fallback → the same
// prompt ran twice. ---
@Test
fun `slow prompt submit ack outlives the generic rpc timeout without SSE fallback`() {
// Shrink the GENERIC rpc timeout below the ack delay: if prompt.submit
// (wrongly) rode the generic timeout again, the submit would fail at
// 500ms and the preflight fallback would fire — failing this test.
rebuildClient(rpcTimeoutMs = 500L)
harness.suppressAckMethods.add("prompt.submit")
val r = Recorder()
client.sendTurn(null, "deep thought", null, r.callbacks) { r.preflightFailures += it }
val serverWs = harness.awaitServerSocket()
val ack = harness.awaitPendingAck()
// Ack arrives well after the generic rpc timeout would have fired.
Thread.sleep(1_500)
harness.releaseAck(ack)
serverWs.send(harness.eventFrame("message.delta", buildJsonObject { put("text", "42") }, "live-1"))
serverWs.send(harness.eventFrame("message.complete", buildJsonObject { put("text", "42") }, "live-1"))
assertTrue("turn never completed", r.completeLatch.await(5, TimeUnit.SECONDS))
assertEquals(listOf("42"), r.textDeltas.toList())
assertTrue("slow ack must not preflight-fail (duplicate turn)", r.preflightFailures.isEmpty())
assertTrue("slow ack must not surface a stream error, got ${r.errors}", r.errors.isEmpty())
assertEquals(1, harness.rpcLog.count { it.first == "prompt.submit" })
}
@Test
fun `turn completes when the ack never arrives and the late ack timeout does not fall back`() {
// Shrink the SUBMIT timeout so its late failure fires inside the test
// budget — after the turn has already completed via stream events.
rebuildClient(promptSubmitTimeoutMs = 1_000L)
harness.suppressAckMethods.add("prompt.submit")
val r = Recorder()
client.sendTurn(null, "hello", null, r.callbacks) { r.preflightFailures += it }
val serverWs = harness.awaitServerSocket()
harness.awaitRpc("prompt.submit")
serverWs.send(harness.eventFrame("message.delta", buildJsonObject { put("text", "Hi!") }, "live-1"))
serverWs.send(harness.eventFrame("message.complete", buildJsonObject { put("text", "Hi!") }, "live-1"))
assertTrue("turn never completed", r.completeLatch.await(5, TimeUnit.SECONDS))
// Let the shortened ack timeout fire AFTER completion — the late
// failure must not resurrect the finished turn on the SSE fallback.
Thread.sleep(1_500)
assertTrue("late ack timeout fired the SSE fallback (duplicate turn)", r.preflightFailures.isEmpty())
assertTrue(r.errors.isEmpty())
assertEquals(1, harness.rpcLog.count { it.first == "prompt.submit" })
}
@Test
fun `idle watchdog does not fire while events keep arriving slowly`() {
rebuildClient(turnIdleTimeoutMs = 1_000L)
val r = Recorder()
client.sendTurn(null, "slow drip", null, r.callbacks) { r.preflightFailures += it }
val serverWs = harness.awaitServerSocket()
harness.awaitRpc("prompt.submit")
// Each event lands inside the (shortened) idle window but the run's
// TOTAL wall-clock far exceeds it — an idle-progress watchdog stays
// quiet; a hard turn cap would have killed the turn.
repeat(8) { i ->
serverWs.send(
harness.eventFrame("message.delta", buildJsonObject { put("text", "d$i ") }, "live-1"),
)
Thread.sleep(250)
}
serverWs.send(harness.eventFrame("message.complete", buildJsonObject { put("text", "done") }, "live-1"))
assertTrue("turn never completed", r.completeLatch.await(5, TimeUnit.SECONDS))
assertTrue("watchdog fired despite live events: ${r.errors}", r.errors.isEmpty())
assertTrue(r.preflightFailures.isEmpty())
assertTrue(
"watchdog must not have interrupted a live turn",
harness.rpcLog.none { it.first == "session.interrupt" },
)
}
@Test
fun `idle watchdog fires when events stop flowing`() {
rebuildClient(turnIdleTimeoutMs = 500L)
val r = Recorder()
client.sendTurn(null, "stalls", null, r.callbacks) { r.preflightFailures += it }
val serverWs = harness.awaitServerSocket()
harness.awaitRpc("prompt.submit")
serverWs.send(harness.eventFrame("message.delta", buildJsonObject { put("text", "partial") }, "live-1"))
// …then silence: the idle watchdog must fail the turn as a STREAM
// error (never a preflight fallback — the turn started server-side)
// and interrupt the server so it stops generating.
assertTrue("watchdog never fired", r.completeLatch.await(5, TimeUnit.SECONDS))
assertTrue("expected a stream error from the watchdog", r.errors.isNotEmpty())
assertTrue(r.preflightFailures.isEmpty())
harness.awaitRpc("session.interrupt")
}
}
@@ -18,6 +18,28 @@ import org.junit.Test
*/
class HermesApiClientTest {
// --- buildApiRequestOrNull (#131 guard, streaming paths) ---
@Test
fun buildApiRequestOrNull_validUrlsBuildARequest() {
assertTrue(buildApiRequestOrNull("http://192.168.1.10:8642/api/sessions/x/chat/stream") != null)
assertTrue(buildApiRequestOrNull("https://hermes.example.com/v1/runs") != null)
}
@Test
fun buildApiRequestOrNull_malformedUrlsReturnNullInsteadOfThrowing() {
// Each would make Request.Builder.url(String) throw
// IllegalArgumentException on the streaming send path.
for (bad in listOf(
"http://", // empty host
"http://in valid host:8642/v1/runs", // space in host
"not-a-url/api/sessions/x/chat/stream", // no scheme (corrupt baseUrl)
"/api/sessions/x/chat/stream", // blank baseUrl
)) {
assertNull("expected null for malformed url '$bad'", buildApiRequestOrNull(bad))
}
}
// --- HealthCheckResult sealed interface ---
@Test
@@ -1,23 +1,102 @@
package com.hermesandroid.relay.network.upstream
import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.network.upstream.models.CreateSessionRequest
import kotlinx.serialization.encodeToString
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonArray
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.boolean
import kotlinx.serialization.json.buildJsonArray
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.contentOrNull
import kotlinx.serialization.json.jsonArray
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.putJsonObject
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Test
/**
* HRUI-001 contract tests: the sessions/runs/completions fallback payloads
* must contain ONLY fields upstream consumes (plus the documented legacy
* hint fields), synthetic phone-local history must land in a supported
* channel, and undeliverable attachments must surface explicitly through
* [ChatPayloadResult.droppedAttachments] — never a silent drop.
*/
class HermesChatPayloadsTest {
private val json = Json { ignoreUnknownKeys = true }
// --- fixtures ---
private val imageAttachment = Attachment(
contentType = "image/png",
content = "IMGB64",
fileName = "shot.png",
)
private val pdfAttachment = Attachment(
contentType = "application/pdf",
content = "PDFB64",
fileName = "doc.pdf",
)
/** OpenAI-format assistant tool-call + tool-result pair (voice intent / card dispatch shape). */
private fun toolCallPair(
callId: String,
name: String,
arguments: String,
result: String,
): List<JsonObject> = listOf(
buildJsonObject {
put("role", "assistant")
put("content", "")
putJsonArray("tool_calls") {
add(buildJsonObject {
put("id", callId)
put("type", "function")
putJsonObject("function") {
put("name", name)
put("arguments", arguments)
}
})
}
},
buildJsonObject {
put("role", "tool")
put("tool_call_id", callId)
put("content", result)
},
)
/** Plain text turn (realtime voice sync shape). */
private fun plainTurn(role: String, content: String): JsonObject = buildJsonObject {
put("role", role)
put("content", content)
}
private fun syntheticArray(entries: List<JsonObject>): JsonArray = buildJsonArray {
entries.forEach { add(it) }
}
private val voiceIntentPair = toolCallPair(
callId = "call_voiceintent_1",
name = "android_open_app",
arguments = """{"app_name":"Chrome"}""",
result = """{"ok":true,"package":"com.android.chrome"}""",
)
private val realtimeTurns = listOf(
plainTurn("user", "what's the capital of France?"),
plainTurn("assistant", "Paris."),
)
// --- CreateSessionRequest (unchanged surface) ---
@Test
fun createSessionRequest_serializesProfileWhenExplicitlySelected() {
val body = json.encodeToString(
@@ -34,6 +113,8 @@ class HermesChatPayloadsTest {
assertEquals("mizu", parsed["profile"]?.jsonPrimitive?.contentOrNull)
}
// --- sessions payload ---
@Test
fun sessionChatPayload_includesProfileModelAndSystemMessage() {
val payload = buildSessionChatStreamPayload(
@@ -41,7 +122,7 @@ class HermesChatPayloadsTest {
systemMessage = "You are Mizu.",
modelOverride = "grok-mizu",
profileName = "mizu",
)
).payload
assertEquals("hello", payload["message"]?.jsonPrimitive?.contentOrNull)
assertEquals("You are Mizu.", payload["system_message"]?.jsonPrimitive?.contentOrNull)
@@ -54,28 +135,167 @@ class HermesChatPayloadsTest {
val payload = buildSessionChatStreamPayload(
message = "hello",
profileName = null,
)
).payload
assertFalse(payload.containsKey("profile"))
}
@Test
fun runPayload_includesProfileAlongsideFallbackModelAndSystemMessage() {
fun sessionChatPayload_sendsOnlyUpstreamContractFields() {
val payload = buildSessionChatStreamPayload(
message = "hello",
systemMessage = "sys",
attachments = listOf(imageAttachment, pdfAttachment),
voiceIntentMessages = syntheticArray(voiceIntentPair + realtimeTurns),
modelOverride = "grok-mizu",
profileName = "mizu",
).payload
// Upstream parses message + system_message; model + profile are
// documented legacy hints. Nothing else may go on the wire —
// in particular no top-level `messages` or `attachments`.
assertEquals(
setOf("message", "system_message", "model", "profile"),
payload.keys,
)
}
@Test
fun sessionChatPayload_foldsSyntheticHistoryIntoSystemMessage() {
val payload = buildSessionChatStreamPayload(
message = "hello",
systemMessage = "You are Mizu.",
voiceIntentMessages = syntheticArray(voiceIntentPair + realtimeTurns),
).payload
val system = payload["system_message"]?.jsonPrimitive?.contentOrNull.orEmpty()
// Caller's per-turn system message stays first.
assertTrue(system.startsWith("You are Mizu."))
assertTrue(system.contains(SYNTHETIC_DIGEST_HEADER))
// Tool-call pair renders name + arguments + result on one line.
assertTrue(
system.contains(
"""- called android_open_app with {"app_name":"Chrome"} -> {"ok":true,"package":"com.android.chrome"}""",
),
)
// Sessions has no history channel, so plain turns join the digest.
assertTrue(system.contains("- user: what's the capital of France?"))
assertTrue(system.contains("- assistant: Paris."))
}
@Test
fun sessionChatPayload_digestAloneWhenNoSystemMessage() {
val payload = buildSessionChatStreamPayload(
message = "hello",
voiceIntentMessages = syntheticArray(voiceIntentPair),
).payload
val system = payload["system_message"]?.jsonPrimitive?.contentOrNull.orEmpty()
assertTrue(system.startsWith(SYNTHETIC_DIGEST_HEADER))
}
@Test
fun sessionChatPayload_reportsAllAttachmentsDropped() {
val result = buildSessionChatStreamPayload(
message = "hello",
attachments = listOf(imageAttachment, pdfAttachment),
)
assertEquals(listOf(imageAttachment, pdfAttachment), result.droppedAttachments)
assertFalse(result.payload.containsKey("attachments"))
}
// --- runs payload ---
@Test
fun runPayload_includesProfileAlongsideFallbackModelAndInstructions() {
val payload = buildRunStreamPayload(
message = "hello",
model = "default-model",
systemMessage = "You are Coder.",
modelOverride = "grok-coder",
profileName = "coder",
)
).payload
assertEquals("hello", payload["input"]?.jsonPrimitive?.contentOrNull)
assertEquals("grok-coder", payload["model"]?.jsonPrimitive?.contentOrNull)
assertEquals("You are Coder.", payload["system_message"]?.jsonPrimitive?.contentOrNull)
// Upstream's runs handler reads `instructions`, not `system_message`.
assertEquals("You are Coder.", payload["instructions"]?.jsonPrimitive?.contentOrNull)
assertFalse(payload.containsKey("system_message"))
assertEquals("coder", payload["profile"]?.jsonPrimitive?.contentOrNull)
assertTrue(payload["stream"]?.jsonPrimitive?.boolean == true)
}
@Test
fun runPayload_sendsOnlyUpstreamContractFields() {
val payload = buildRunStreamPayload(
message = "hello",
systemMessage = "sys",
attachments = listOf(imageAttachment, pdfAttachment),
voiceIntentMessages = syntheticArray(voiceIntentPair + realtimeTurns),
modelOverride = "grok-coder",
profileName = "coder",
).payload
assertEquals(
setOf("model", "input", "stream", "instructions", "conversation_history", "profile"),
payload.keys,
)
}
@Test
fun runPayload_splicesPlainTurnsIntoConversationHistoryAndDigestsToolPairs() {
val payload = buildRunStreamPayload(
message = "hello",
voiceIntentMessages = syntheticArray(voiceIntentPair + realtimeTurns),
).payload
// Plain realtime turns ride the upstream-parsed history channel verbatim.
val history = payload["conversation_history"]!!.jsonArray
assertEquals(2, history.size)
assertEquals("user", history[0].jsonObject["role"]?.jsonPrimitive?.contentOrNull)
assertEquals(
"what's the capital of France?",
history[0].jsonObject["content"]?.jsonPrimitive?.contentOrNull,
)
assertEquals("assistant", history[1].jsonObject["role"]?.jsonPrimitive?.contentOrNull)
assertEquals("Paris.", history[1].jsonObject["content"]?.jsonPrimitive?.contentOrNull)
// Tool-call pairs go through the instructions digest — and only them
// (plain turns must not be delivered twice).
val instructions = payload["instructions"]?.jsonPrimitive?.contentOrNull.orEmpty()
assertTrue(instructions.contains("- called android_open_app"))
assertFalse(instructions.contains("- user:"))
assertFalse(instructions.contains("- assistant:"))
}
@Test
fun runPayload_omitsConversationHistoryWhenOnlyToolPairs() {
val payload = buildRunStreamPayload(
message = "hello",
voiceIntentMessages = syntheticArray(voiceIntentPair),
).payload
assertFalse(payload.containsKey("conversation_history"))
assertTrue(
payload["instructions"]?.jsonPrimitive?.contentOrNull.orEmpty()
.contains("- called android_open_app"),
)
}
@Test
fun runPayload_reportsAllAttachmentsDropped() {
val result = buildRunStreamPayload(
message = "hello",
attachments = listOf(imageAttachment, pdfAttachment),
)
assertEquals(listOf(imageAttachment, pdfAttachment), result.droppedAttachments)
assertFalse(result.payload.containsKey("attachments"))
}
// --- completions payload ---
@Test
fun chatCompletionsPayload_usesOpenAiMessagesAndSseStream() {
val payload = buildChatCompletionsStreamPayload(
@@ -84,7 +304,7 @@ class HermesChatPayloadsTest {
systemMessage = "You are Coder.",
modelOverride = "grok-coder",
profileName = "coder",
)
).payload
assertEquals("grok-coder", payload["model"]?.jsonPrimitive?.contentOrNull)
assertEquals("coder", payload["profile"]?.jsonPrimitive?.contentOrNull)
@@ -97,4 +317,100 @@ class HermesChatPayloadsTest {
assertEquals("user", messages[1].jsonObject["role"]?.jsonPrimitive?.contentOrNull)
assertEquals("hello", messages[1].jsonObject["content"]?.jsonPrimitive?.contentOrNull)
}
@Test
fun chatCompletionsPayload_inlinesImageAttachmentsAsImageUrlParts() {
val result = buildChatCompletionsStreamPayload(
message = "what is this?",
attachments = listOf(imageAttachment),
)
val messages = result.payload["messages"]!!.jsonArray
val userContent = messages.last().jsonObject["content"]!!.jsonArray
assertEquals("text", userContent[0].jsonObject["type"]?.jsonPrimitive?.contentOrNull)
assertEquals("what is this?", userContent[0].jsonObject["text"]?.jsonPrimitive?.contentOrNull)
assertEquals("image_url", userContent[1].jsonObject["type"]?.jsonPrimitive?.contentOrNull)
assertEquals(
"data:image/png;base64,IMGB64",
userContent[1].jsonObject["image_url"]?.jsonObject?.get("url")?.jsonPrimitive?.contentOrNull,
)
// Images have a real channel here — nothing dropped.
assertTrue(result.droppedAttachments.isEmpty())
}
@Test
fun chatCompletionsPayload_splicesPlainTurnsAndDigestsToolPairs() {
val payload = buildChatCompletionsStreamPayload(
message = "hello",
voiceIntentMessages = syntheticArray(voiceIntentPair + realtimeTurns),
).payload
val messages = payload["messages"]!!.jsonArray
val roles = messages.map { it.jsonObject["role"]?.jsonPrimitive?.contentOrNull }
// system digest + spliced realtime turns + live user message; the
// tool-call pair must NOT be spliced (upstream skips tool-role
// messages and strips tool_calls, destroying the record).
assertEquals(listOf("system", "user", "assistant", "user"), roles)
assertFalse(messages.any { it.jsonObject.containsKey("tool_calls") })
val system = messages[0].jsonObject["content"]?.jsonPrimitive?.contentOrNull.orEmpty()
assertTrue(system.contains("- called android_open_app"))
assertEquals(
"what's the capital of France?",
messages[1].jsonObject["content"]?.jsonPrimitive?.contentOrNull,
)
assertEquals("Paris.", messages[2].jsonObject["content"]?.jsonPrimitive?.contentOrNull)
assertEquals("hello", messages[3].jsonObject["content"]?.jsonPrimitive?.contentOrNull)
}
@Test
fun chatCompletionsPayload_dropsOnlyNonImageAttachments() {
val result = buildChatCompletionsStreamPayload(
message = "hello",
attachments = listOf(imageAttachment, pdfAttachment),
)
assertEquals(listOf(pdfAttachment), result.droppedAttachments)
// The dead top-level `attachments` field is gone for good.
assertFalse(result.payload.containsKey("attachments"))
}
// --- digest helper edge cases ---
@Test
fun renderSyntheticHistoryDigest_nullWhenNothingRenders() {
assertNull(renderSyntheticHistoryDigest(null, includePlainTurns = true))
assertNull(renderSyntheticHistoryDigest(buildJsonArray {}, includePlainTurns = true))
// Plain turns excluded (endpoints with a real history channel) and
// no tool pairs present -> nothing to fold into the prompt.
assertNull(
renderSyntheticHistoryDigest(
syntheticArray(realtimeTurns),
includePlainTurns = false,
),
)
}
@Test
fun renderSyntheticHistoryDigest_pairsResultsByToolCallId() {
val pairA = toolCallPair("call_a", "android_open_app", """{"app_name":"Maps"}""", """{"ok":true}""")
val pairB = toolCallPair(
"call_b",
"hermes_card_action",
"""{"card_key":"k","action_value":"/approve"}""",
"""{"ok":true,"dispatched_at":1}""",
)
val digest = renderSyntheticHistoryDigest(
syntheticArray(pairA + pairB),
includePlainTurns = false,
).orEmpty()
assertTrue(digest.startsWith(SYNTHETIC_DIGEST_HEADER))
assertTrue(digest.contains("""- called android_open_app with {"app_name":"Maps"} -> {"ok":true}"""))
assertTrue(
digest.contains(
"""- called hermes_card_action with {"card_key":"k","action_value":"/approve"} -> {"ok":true,"dispatched_at":1}""",
),
)
}
}
@@ -0,0 +1,128 @@
package com.hermesandroid.relay.notifications
import androidx.datastore.core.DataStore
import androidx.datastore.preferences.core.Preferences
import androidx.datastore.preferences.core.emptyPreferences
import kotlinx.coroutines.flow.Flow
import kotlinx.coroutines.flow.MutableStateFlow
import kotlinx.coroutines.flow.first
import kotlinx.coroutines.runBlocking
import org.junit.Assert.assertEquals
import org.junit.Assert.assertNotNull
import org.junit.Assert.assertNull
import org.junit.Test
class NotificationTriggerStoreTest {
@Test
fun disabledByDefaultDoesNotMatch() = runBlocking {
val store = NotificationTriggerStore(InMemoryPreferencesDataStore())
store.saveSingleRule(
NotificationTriggerRule(appPackage = "com.slack"),
)
assertNull(store.firstMatchingRule(entry(packageName = "com.slack")))
}
@Test
fun matchesEnabledRuleByAppAndTextFilter() = runBlocking {
val store = NotificationTriggerStore(InMemoryPreferencesDataStore())
store.setMasterEnabled(true)
store.saveSingleRule(
NotificationTriggerRule(
label = "Slack from Sam",
appPackage = "com.slack",
textContains = "Sam",
),
)
assertNotNull(
store.firstMatchingRule(
entry(
packageName = "com.slack",
title = "Axiom",
text = "Sam: deploy finished",
),
),
)
assertNull(
store.firstMatchingRule(
entry(
packageName = "com.slack",
title = "Axiom",
text = "Alex: deploy finished",
),
),
)
}
@Test
fun killSwitchBlocksMatchesWithoutDeletingRule() = runBlocking {
val store = NotificationTriggerStore(InMemoryPreferencesDataStore())
store.setMasterEnabled(true)
store.saveSingleRule(NotificationTriggerRule(appPackage = "com.slack"))
store.setKillSwitch(true)
assertNull(store.firstMatchingRule(entry(packageName = "com.slack")))
assertEquals(1, store.settings.first().rules.size)
}
@Test
fun emptyFilterRuleNeverMatches() {
val rule = NotificationTriggerRule(
appPackage = " ",
titleContains = "",
textContains = null,
)
assertEquals(false, rule.matches(entry(packageName = "com.any")))
}
@Test
fun activityLogIsNewestFirstAndCapped() = runBlocking {
val store = NotificationTriggerStore(InMemoryPreferencesDataStore())
repeat(NotificationTriggerStore.MAX_ACTIVITY_LOG_ENTRIES + 2) { idx ->
store.appendActivity(
NotificationTriggerActivityEntry(
ruleId = "rule",
ruleLabel = "Rule",
action = NotificationTriggerAction.AskMe,
packageName = "pkg.$idx",
matchedAt = idx.toLong(),
result = "prompt posted",
),
)
}
val log = store.settings.first().activityLog
assertEquals(NotificationTriggerStore.MAX_ACTIVITY_LOG_ENTRIES, log.size)
assertEquals("pkg.${NotificationTriggerStore.MAX_ACTIVITY_LOG_ENTRIES + 1}", log.first().packageName)
}
private fun entry(
packageName: String,
title: String? = "Title",
text: String? = "Text",
subText: String? = null,
) = NotificationEntry(
packageName = packageName,
title = title,
text = text,
subText = subText,
postedAt = 123L,
key = "$packageName:key",
)
private class InMemoryPreferencesDataStore : DataStore<Preferences> {
private val state = MutableStateFlow<Preferences>(emptyPreferences())
override val data: Flow<Preferences> = state
override suspend fun updateData(transform: suspend (t: Preferences) -> Preferences): Preferences {
val next = transform(state.value)
state.value = next
return next
}
}
}
@@ -0,0 +1,106 @@
package com.hermesandroid.relay.screenshots
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.outlined.Forum
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Text
import androidx.compose.runtime.Composable
import androidx.compose.runtime.CompositionLocalProvider
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.platform.LocalDensity
import androidx.compose.ui.test.junit4.createComposeRule
import androidx.compose.ui.test.onNodeWithText
import androidx.compose.ui.test.onRoot
import androidx.compose.ui.test.performScrollTo
import androidx.compose.ui.unit.Density
import com.github.takahirom.roborazzi.captureRoboImage
import com.hermesandroid.relay.ui.onboarding.OnboardingPage
import com.hermesandroid.relay.ui.theme.HermesRelayTheme
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith
import androidx.test.ext.junit.runners.AndroidJUnit4
import org.robolectric.annotation.Config
import org.robolectric.annotation.GraphicsMode
/**
* Regression harness for issue #145 — onboarding slide content overflowed
* below the fold with no scroll affordance on short viewports / raised font
* scale. Renders [OnboardingPage] the way the pager hosts it (a centering
* fillMaxSize Box) at compact heights and a raised font scale, then scrolls
* to the last body line to prove every line is reachable.
*
* Render-success + reachability assertions only — no golden PNGs are
* committed; the store screenshot set is untouched.
*/
@RunWith(AndroidJUnit4::class)
@GraphicsMode(GraphicsMode.Mode.NATIVE)
@Config(qualifiers = "w320dp-h480dp-xhdpi")
class OnboardingCompactScreenshotTest {
@get:Rule
val compose = createComposeRule()
private val lastLine = "Final onboarding body line for reachability"
/** Mirrors the pager's page container: fillMaxSize Box, centered content. */
@Composable
private fun PagerHostedSlide(fontScale: Float) {
val base = LocalDensity.current
CompositionLocalProvider(
LocalDensity provides Density(base.density, fontScale = fontScale)
) {
HermesRelayTheme(appThemeId = "hermes-relay", themePreference = "dark") {
Box(
modifier = Modifier.fillMaxSize(),
contentAlignment = Alignment.Center
) {
OnboardingPage(
icon = Icons.Outlined.Forum,
title = "Chat",
description = "Your Hermes agent, streaming in real time.",
) {
Text(
text = "Live responses with tool progress, markdown, and rich cards as the agent works.",
style = MaterialTheme.typography.bodySmall,
modifier = Modifier.fillMaxWidth(),
)
Text(
text = "Switch agent profiles mid-flow — each keeps its own sessions, model, and persona.",
style = MaterialTheme.typography.bodySmall,
modifier = Modifier.fillMaxWidth(),
)
Text(
text = lastLine,
style = MaterialTheme.typography.bodySmall,
modifier = Modifier.fillMaxWidth(),
)
}
}
}
}
}
// 480dp-tall viewport at 1.5x font scale — the hero shrinks/drops and the
// body overflows; the page must scroll so the last line stays reachable.
@Test
fun compactHeight_largeFontScale_scrollsToLastLine() {
compose.setContent { PagerHostedSlide(fontScale = 1.5f) }
compose.onNodeWithText(lastLine).performScrollTo().assertExists()
compose.onRoot().captureRoboImage("build/onboarding-shots/compact_480dp_font1_5.png")
}
// 600dp-tall viewport — the 160dp shrunk-hero tier; content should render
// (and remain scroll-reachable) without dropping the hero entirely.
@Test
@Config(qualifiers = "w360dp-h600dp-xhdpi")
fun mediumHeight_defaultFontScale_rendersShrunkHero() {
compose.setContent { PagerHostedSlide(fontScale = 1.0f) }
compose.onNodeWithText(lastLine).performScrollTo().assertExists()
compose.onRoot().captureRoboImage("build/onboarding-shots/medium_600dp.png")
}
}
@@ -3,14 +3,86 @@ package com.hermesandroid.relay.ui.components
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.viewmodel.InteractionMode
import com.hermesandroid.relay.viewmodel.BackgroundRunState
import com.hermesandroid.relay.viewmodel.BackgroundRunPhase
import com.hermesandroid.relay.viewmodel.HermesConfirmationState
import com.hermesandroid.relay.viewmodel.VoiceHandoffStatus
import com.hermesandroid.relay.viewmodel.VoiceState
import com.hermesandroid.relay.viewmodel.VoiceUiState
import com.hermesandroid.relay.viewmodel.backgroundRunAfterCancelRequest
import com.hermesandroid.relay.viewmodel.preserveRealtimeTurnOnStop
import com.hermesandroid.relay.viewmodel.realtimeTranscriptState
import com.hermesandroid.relay.viewmodel.realtimeTurnActiveAfterResponseDone
import com.hermesandroid.relay.viewmodel.voiceSessionExitState
import org.junit.Assert.assertEquals
import org.junit.Assert.assertNull
import org.junit.Test
class VoiceModeOverlayStateTest {
@Test
fun providerTranscript_isTranscribingAfterMicrophoneCaptureStops() {
assertEquals(VoiceState.Transcribing, realtimeTranscriptState(micCaptureActive = false))
assertEquals(VoiceState.Listening, realtimeTranscriptState(micCaptureActive = true))
}
@Test
fun responseDone_keepsLogicalTurnActiveOnlyWhileBackgroundRunIsLive() {
assertEquals(true, realtimeTurnActiveAfterResponseDone(BackgroundRunPhase.RUNNING))
assertEquals(true, realtimeTurnActiveAfterResponseDone(BackgroundRunPhase.RECONNECTING))
assertEquals(false, realtimeTurnActiveAfterResponseDone(BackgroundRunPhase.DELIVERING))
assertEquals(false, realtimeTurnActiveAfterResponseDone(BackgroundRunPhase.DONE))
assertEquals(false, realtimeTurnActiveAfterResponseDone(null))
}
@Test
fun stop_preservesSharedTurnWhileBackgroundSummaryCanStillArrive() {
assertEquals(true, preserveRealtimeTurnOnStop(BackgroundRunPhase.RUNNING))
assertEquals(true, preserveRealtimeTurnOnStop(BackgroundRunPhase.RECONNECTING))
assertEquals(true, preserveRealtimeTurnOnStop(BackgroundRunPhase.DELIVERING))
assertEquals(false, preserveRealtimeTurnOnStop(BackgroundRunPhase.DONE))
assertEquals(false, preserveRealtimeTurnOnStop(null))
}
@Test
fun voiceExit_clearsDetachedReconnectStateBeforeNextEntry() {
val exited = voiceSessionExitState(
VoiceUiState(
voiceMode = true,
state = VoiceState.Thinking,
handoffStatus = VoiceHandoffStatus(title = "Waiting for route"),
hermesConfirmation = HermesConfirmationState(
confirmationId = "confirmation-orphaned",
message = "Approve this action?",
),
backgroundRun = BackgroundRunState(
runId = "run-orphaned",
phase = BackgroundRunPhase.RECONNECTING,
),
)
)
assertEquals(false, exited.voiceMode)
assertEquals(VoiceState.Idle, exited.state)
assertNull(exited.handoffStatus)
assertNull(exited.backgroundRun)
assertNull(exited.hermesConfirmation)
}
@Test
fun backgroundCancel_clearsChipWhenSocketRejectsRequest() {
val run = BackgroundRunState(
runId = "run-offline",
phase = BackgroundRunPhase.RECONNECTING,
)
assertNull(backgroundRunAfterCancelRequest(run, cancelSent = false))
assertEquals(
"Cancelling…",
backgroundRunAfterCancelRequest(run, cancelSent = true)?.message,
)
}
@Test
fun pendingTranscript_showsWhileThinkingBeforeChatHistoryCatchesUp() {
val text = pendingVoiceTranscriptText(
@@ -0,0 +1,179 @@
package com.hermesandroid.relay.ui.screens
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Test
/**
* HRUI-022 — `parseModelOptions` must keep the full provider catalog visible.
*
* Newer upstream only returns unconfigured provider skeleton rows when the
* client opts in via `include_unconfigured=1`; those rows arrive with EMPTY
* `models` plus picker hints (`authenticated=false`, `key_env`, `warning`).
* Dropping them silently hides every provider that still needs an API key,
* killing the Manage → Keys setup affordance.
*/
class ModelOptionsParserTest {
private fun parse(json: String): List<ModelProviderOption> =
parseModelOptions(Json.parseToJsonElement(json).jsonObject)
@Test
fun newUpstream_keepsUnconfiguredSkeletonRowAlongsideAuthenticatedProvider() {
// Shape from upstream build_models_payload(include_unconfigured=True,
// picker_hints=True): one authenticated row with models, one canonical
// skeleton row with empty models + setup-hint fields.
val options = parse(
"""
{
"providers": [
{
"slug": "openai",
"name": "OpenAI",
"is_current": true,
"is_user_defined": false,
"models": ["gpt-5.5", "gpt-5.5-mini"],
"total_models": 2,
"authenticated": true
},
{
"slug": "anthropic",
"name": "Anthropic",
"is_current": false,
"is_user_defined": false,
"models": [],
"total_models": 0,
"source": "canonical",
"authenticated": false,
"auth_type": "api_key",
"key_env": "ANTHROPIC_API_KEY",
"warning": "paste ANTHROPIC_API_KEY to activate"
}
],
"model": "gpt-5.5",
"provider": "openai"
}
""".trimIndent(),
)
// Both rows survive — the empty-models skeleton must NOT be dropped.
assertEquals(2, options.size)
val authenticated = options[0]
assertEquals("openai", authenticated.id)
assertEquals("OpenAI", authenticated.label)
assertTrue(authenticated.authenticated)
assertEquals(listOf("gpt-5.5", "gpt-5.5-mini"), authenticated.models)
// Skeleton row: greyed, sorted after authenticated rows, and keeps the
// Keys-guidance affordance (setup hint + unauthenticated flag).
val skeleton = options[1]
assertEquals("anthropic", skeleton.id)
assertEquals("Anthropic", skeleton.label)
assertFalse(skeleton.authenticated)
assertTrue(skeleton.models.isEmpty())
assertEquals("paste ANTHROPIC_API_KEY to activate", skeleton.setupHint)
}
@Test
fun oldUpstream_fullListByDefault_parsesIdentically() {
// Old upstream returned the universe without an opt-in; unauthenticated
// rows could still carry curated models. Nothing about that shape may
// parse differently after the include_unconfigured change.
val options = parse(
"""
{
"providers": [
{
"slug": "openai",
"name": "OpenAI",
"models": ["gpt-5.5"],
"authenticated": true
},
{
"slug": "xai",
"name": "xAI",
"models": ["grok-4"],
"authenticated": false,
"warning": "paste XAI_API_KEY to activate"
}
],
"model": "gpt-5.5",
"provider": "openai"
}
""".trimIndent(),
)
assertEquals(2, options.size)
assertEquals("openai", options[0].id)
assertTrue(options[0].authenticated)
// Unauthenticated-with-models keeps its catalog visible (greyed rows).
assertEquals("xai", options[1].id)
assertFalse(options[1].authenticated)
assertEquals(listOf("grok-4"), options[1].models)
assertEquals("paste XAI_API_KEY to activate", options[1].setupHint)
}
@Test
fun missingAuthenticatedHint_defaultsByModelPresence() {
// Payloads without picker hints: a row with models is assumed usable;
// an empty row can only be an unconfigured skeleton, so grey it.
val options = parse(
"""
{
"providers": [
{"slug": "openai", "name": "OpenAI", "models": ["gpt-5.5"]},
{"slug": "anthropic", "name": "Anthropic", "models": []}
]
}
""".trimIndent(),
)
assertEquals(2, options.size)
assertTrue(options[0].authenticated)
assertFalse(options[1].authenticated)
assertNull(options[1].setupHint)
}
@Test
fun idResolution_prefersSlugThenFallsBackForLegacyShapes() {
val options = parse(
"""
{
"providers": [
{"slug": "openai", "id": "ignored", "name": "OpenAI", "models": ["gpt-5.5"]},
{"id": "legacy-id", "name": "Legacy", "models": ["m1"]},
{"name": "name-only", "models": ["m2"]},
{"models": ["orphan-model"]}
]
}
""".trimIndent(),
)
// The selectable id must be the canonical provider slug when present —
// /api/model/set expects the slug, not the display label.
assertEquals(listOf("openai", "legacy-id", "name-only"), options.map { it.id })
}
@Test
fun authenticatedRowsSortAheadOfSkeletons_preservingServerOrder() {
val options = parse(
"""
{
"providers": [
{"slug": "a-skel", "name": "A", "models": [], "authenticated": false},
{"slug": "z-auth", "name": "Z", "models": ["m1"], "authenticated": true},
{"slug": "b-auth", "name": "B", "models": ["m2"], "authenticated": true}
]
}
""".trimIndent(),
)
// Authenticated first; server (canonical) order kept within each group.
assertEquals(listOf("z-auth", "b-auth", "a-skel"), options.map { it.id })
}
}
@@ -1,6 +1,7 @@
package com.hermesandroid.relay.util
import com.hermesandroid.relay.diagnostics.DiagnosticCategory
import com.hermesandroid.relay.diagnostics.DiagnosticLogEntry
import com.hermesandroid.relay.diagnostics.DiagnosticSeverity
import com.hermesandroid.relay.diagnostics.DiagnosticsLog
import java.io.IOException
@@ -92,6 +93,95 @@ class IssueReportAndDiagnosticsTest {
assertTrue(DiagnosticsLog.recent().isEmpty())
}
// --- DiagnosticIssuePrefill: severity-dependent title / labels / body ---
private fun sampleEntry(
severity: DiagnosticSeverity,
title: String = "Testing API connection",
endpointRole: String? = null,
url: String? = null,
detail: String? = null,
) = DiagnosticLogEntry(
timestampMs = 0L,
category = DiagnosticCategory.Api,
severity = severity,
title = title,
detail = detail,
endpointRole = endpointRole,
url = url,
)
@Test
fun errorEntriesKeepBugTitleAndLabel() {
val entry = sampleEntry(DiagnosticSeverity.Error, title = "API key rejected")
assertEquals("[Bug]: API key rejected", DiagnosticIssuePrefill.issueTitle(entry))
assertEquals("bug", DiagnosticIssuePrefill.issueLabels(entry))
}
@Test
fun infoAndWarningEntriesRetitleAsDiagnosticWithQuestionLabel() {
for (severity in listOf(DiagnosticSeverity.Info, DiagnosticSeverity.Warning)) {
val entry = sampleEntry(severity)
assertEquals("[Diagnostic]: Testing API connection", DiagnosticIssuePrefill.issueTitle(entry))
// "question" already exists on the repo — the prefill must not invent labels.
assertEquals("question", DiagnosticIssuePrefill.issueLabels(entry))
}
}
@Test
fun bodyUsesTheExpectationAnswerAsWhatHappened() {
val body = DiagnosticIssuePrefill.issueBody(
sampleEntry(DiagnosticSeverity.Info),
expectation = "I expected the app to connect to my server",
)
assertTrue(body.contains("### What happened?\nI expected the app to connect to my server\n"))
assertFalse(body.contains("Captured diagnostic from the in-app activity log."))
}
@Test
fun bodyFallsBackToBoilerplateWithoutAnExpectation() {
val body = DiagnosticIssuePrefill.issueBody(sampleEntry(DiagnosticSeverity.Error))
assertTrue(body.contains("### What happened?\nCaptured diagnostic from the in-app activity log.\n"))
}
@Test
fun expectationAnswerIsSecretRedactedBeforeEmbedding() {
val body = DiagnosticIssuePrefill.issueBody(
sampleEntry(DiagnosticSeverity.Info),
expectation = "it failed with token=super-secret-value somehow",
)
assertFalse(body.contains("super-secret-value"))
assertTrue(body.contains("token=[hidden]"))
}
@Test
fun connectionModeUsesTheEntryRoleWhenPresent() {
val body = DiagnosticIssuePrefill.issueBody(
sampleEntry(DiagnosticSeverity.Error, endpointRole = "tailscale", url = "http://10.0.0.5:8642"),
)
assertTrue(body.contains("- Connection mode: tailscale"))
assertFalse(body.contains("LAN / Tailscale / public TLS / other"))
}
@Test
fun connectionModeIsInferredFromTheEntryUrl() {
val lan = DiagnosticIssuePrefill.issueBody(
sampleEntry(DiagnosticSeverity.Info, url = "http://localhost:8642"),
)
assertTrue(lan.contains("- Connection mode: lan"))
val tailscale = DiagnosticIssuePrefill.issueBody(
sampleEntry(DiagnosticSeverity.Info, url = "http://100.75.1.2:8642"),
)
assertTrue(tailscale.contains("- Connection mode: tailscale"))
}
@Test
fun connectionModeIsUnknownWithoutRouteOrUrl() {
val body = DiagnosticIssuePrefill.issueBody(sampleEntry(DiagnosticSeverity.Warning))
assertTrue(body.contains("- Connection mode: unknown"))
}
@Test
fun recordErrorRedactsSecretsInTheStacktrace() {
DiagnosticsLog.clear()
@@ -101,6 +101,52 @@ class ServerAddressTest {
assertEquals("https", ServerAddress.parse("https://h.example")?.scheme)
}
// --- loopbackHostWarning: loopback addresses can't reach the server from a phone ---
@Test
fun loopbackHostWarningFlagsLoopbackAndAnyInterfaceHosts() {
val loopbacks = listOf(
"localhost",
"localhost:8642",
"http://localhost:8642",
"127.0.0.1",
"https://127.0.0.1:9119",
"::1",
"[::1]",
"http://[::1]:8642",
"0.0.0.0",
"http://0.0.0.0:8642",
)
for (value in loopbacks) {
assertNotNull("expected warning: '$value'", ServerAddress.loopbackHostWarning(value))
}
}
@Test
fun loopbackHostWarningIsNullForReachableAddresses() {
val reachable = listOf(
"192.168.1.10",
"192.168.1.10:8642",
"http://10.0.0.5:8642",
"100.64.0.1:8642",
"hermes.tail1234.ts.net",
"https://hermes.example.com:8642",
"hermes-box",
// Not loopback: hostname that merely starts with 127.
"127.evil.example.com",
)
for (value in reachable) {
assertNull("expected no warning: '$value'", ServerAddress.loopbackHostWarning(value))
}
}
@Test
fun loopbackHostWarningIsNullForBlankAndJunk() {
assertNull(ServerAddress.loopbackHostWarning(""))
assertNull(ServerAddress.loopbackHostWarning(" "))
assertNull(ServerAddress.loopbackHostWarning("not a host"))
}
// --- fieldError: inline UI message contract ---
@Test
@@ -0,0 +1,274 @@
package com.hermesandroid.relay.viewmodel
import com.hermesandroid.relay.network.upstream.models.MessageItem
import java.io.IOException
import kotlinx.coroutines.ExperimentalCoroutinesApi
import kotlinx.coroutines.test.advanceTimeBy
import kotlinx.coroutines.test.runTest
import kotlinx.serialization.json.JsonPrimitive
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Test
/**
* Virtual-time tests for [ChatStreamRecovery] — the issue #166 poller that
* recovers a dropped sessions-stream turn from the persisted transcript.
*
* Timing uses the production defaults (5s → ×2 backoff → 30s cap, 30 min
* window); `runTest` virtual time makes even the 30-minute cap instant.
*/
@OptIn(ExperimentalCoroutinesApi::class)
class ChatStreamRecoveryTest {
private val pending = "what's the weather?"
private fun user(id: String, text: String = pending) = MessageItem(
id = id,
role = "user",
content = JsonPrimitive(text),
)
private fun assistant(id: String, text: String) = MessageItem(
id = id,
role = "assistant",
content = JsonPrimitive(text),
)
/** Scripted fetch: returns [responses] in order, repeating the last one. */
private class ScriptedHistory(private val responses: List<() -> List<MessageItem>>) {
val fetchTimesMs = mutableListOf<Long>()
private var calls = 0
fun fetcher(now: () -> Long): suspend () -> List<MessageItem> = {
fetchTimesMs += now()
val step = responses[minOf(calls, responses.lastIndex)]
calls++
step()
}
val fetchCount: Int get() = calls
}
@Test
fun recoversWhenAnswerIsStableAcrossTwoConsecutivePolls() = runTest {
val script = ScriptedHistory(
listOf(
{ listOf(user("u1")) },
{ listOf(user("u1"), assistant("a1", "recovered answer")) },
{ listOf(user("u1"), assistant("a1", "recovered answer")) },
),
)
val intermediate = mutableListOf<List<MessageItem>>()
var recovered: List<MessageItem>? = null
var gaveUp: ChatStreamRecovery.GiveUpReason? = null
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(
pendingUserText = pending,
priorUserMessageCount = 0,
onIntermediateHistory = { intermediate += it },
onRecovered = { recovered = it },
onGaveUp = { gaveUp = it },
)
advanceTimeBy(5_001) // poll 1 — user only, no answer yet
assertEquals(1, intermediate.size)
assertNull(recovered)
advanceTimeBy(10_000) // poll 2 — answer appears (signature recorded)
assertEquals(2, intermediate.size)
assertNull(recovered)
advanceTimeBy(20_000) // poll 3 — unchanged → stable → recovered
assertEquals("recovered answer", recovered?.last()?.contentText)
assertNull(gaveUp)
assertFalse(recovery.isActive)
}
@Test
fun pollCadenceBacksOffExponentiallyToTheCap() = runTest {
val script = ScriptedHistory(listOf({ emptyList() }))
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(pending, 0, {}, {}, {})
// 5s, +10s, +20s, +30s, +30s… (cap) — cumulative fetch times.
advanceTimeBy(5_000 + 10_000 + 20_000 + 30_000 + 30_000 + 1)
assertEquals(
listOf(5_000L, 15_000L, 35_000L, 65_000L, 95_000L),
script.fetchTimesMs,
)
recovery.cancel()
}
@Test
fun emptyTranscriptCarriesNoInformationAndKeepsPolling() = runTest {
val script = ScriptedHistory(listOf({ emptyList() }))
var gaveUp: ChatStreamRecovery.GiveUpReason? = null
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(pending, 0, {}, {}, { gaveUp = it })
advanceTimeBy(31L * 60_000)
assertEquals(ChatStreamRecovery.GiveUpReason.TIMED_OUT, gaveUp)
assertTrue("should have kept polling to the cap", script.fetchCount > 10)
}
@Test
fun reachableTranscriptWithoutThePendingSendFailsFast() = runTest {
// Server reachable, but the turn's POST never landed: the transcript
// still holds only the PREVIOUS exchange (one user row the client
// already knew about). No new user row → no anchor → waiting cannot
// produce an answer, so give up quickly (after confirming, not on the
// very first read) rather than polling to the 30-minute cap.
val script = ScriptedHistory(
listOf(
{
listOf(
user("u0", "an earlier prompt"),
assistant("a0", "an earlier answer"),
)
},
),
)
var gaveUp: ChatStreamRecovery.GiveUpReason? = null
var recovered = false
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(pending, 1, {}, { recovered = true }, { gaveUp = it })
advanceTimeBy(5_001) // poll 1 — not established, not yet confirmed
assertNull("must confirm across a couple polls before giving up", gaveUp)
advanceTimeBy(10_000) // poll 2 — still not established → fail fast
assertEquals(ChatStreamRecovery.GiveUpReason.RUN_NOT_FOUND, gaveUp)
assertFalse("stale prior answer must never count as the recovery", recovered)
assertEquals(2, script.fetchCount)
assertFalse(recovery.isActive)
}
@Test
fun staleIdenticalEarlierUserMessageIsNotAdoptedWhenSendNeverLanded() = runTest {
// History the client already knew: "continue" → A, "tell me more" → B.
// The new send is ALSO "continue" but its POST never reached the
// server, so the transcript is unchanged. A bare `indexOfLast` text
// match would anchor on the STALE first "continue" and adopt B (static
// → instantly "stable"). The positional invariant (this send would be
// the 3rd user row, but only 2 exist) must refuse to adopt and fail.
val stale = listOf(
user("u1", "continue"),
assistant("a1", "answer A"),
user("u2", "tell me more"),
assistant("a2", "answer B"),
)
val script = ScriptedHistory(listOf({ stale }))
var recovered: List<MessageItem>? = null
var gaveUp: ChatStreamRecovery.GiveUpReason? = null
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(
pendingUserText = "continue",
priorUserMessageCount = 2, // u1 + u2 were known before this send
onIntermediateHistory = {},
onRecovered = { recovered = it },
onGaveUp = { gaveUp = it },
)
advanceTimeBy(5_000 + 10_000 + 1) // two polls of the static transcript
assertNull("must NOT adopt a different turn's answer", recovered)
assertEquals(ChatStreamRecovery.GiveUpReason.RUN_NOT_FOUND, gaveUp)
assertFalse(recovery.isActive)
}
@Test
fun repeatedShortPromptAnchorsTheNewTurnNotTheStaleOne() = runTest {
// "continue" → A is already in history; the new "continue" DID land, so
// the server appended a second identical user row + its own answer C.
// The positional anchor must pick the SECOND "continue" and adopt C —
// never the stale A that an `indexOfLast`-only heuristic risks.
val landed = listOf(
user("u1", "continue"),
assistant("a1", "answer A"),
user("u2", "continue"),
assistant("a2", "answer C"),
)
val script = ScriptedHistory(listOf({ landed }))
var recovered: List<MessageItem>? = null
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(
pendingUserText = "continue",
priorUserMessageCount = 1, // only the first "continue" was known
onIntermediateHistory = {},
onRecovered = { recovered = it },
onGaveUp = {},
)
advanceTimeBy(5_000 + 10_000 + 1) // stable across two polls
assertEquals("answer C", recovered?.last()?.contentText)
}
@Test
fun fetchExceptionKeepsPollingUntilTheServerComesBack() = runTest {
val script = ScriptedHistory(
listOf(
{ throw IOException("network still down") },
{ listOf(user("u1"), assistant("a1", "late answer")) },
{ listOf(user("u1"), assistant("a1", "late answer")) },
),
)
var recovered: List<MessageItem>? = null
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(pending, 0, {}, { recovered = it }, {})
advanceTimeBy(5_000 + 10_000 + 20_000 + 1)
assertEquals("late answer", recovered?.last()?.contentText)
}
@Test
fun growingTranscriptDefersTheFinishUntilStable() = runTest {
val script = ScriptedHistory(
listOf(
{ listOf(user("u1"), assistant("a1", "thinking about it")) },
// Run still going: a new assistant row appended → not stable.
{
listOf(
user("u1"),
assistant("a1", "thinking about it"),
assistant("a2", "final answer"),
)
},
{
listOf(
user("u1"),
assistant("a1", "thinking about it"),
assistant("a2", "final answer"),
)
},
),
)
var recovered: List<MessageItem>? = null
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(pending, 0, {}, { recovered = it }, {})
advanceTimeBy(15_001) // polls 1+2 — signature changed between them
assertNull(recovered)
advanceTimeBy(20_000) // poll 3 — stable now
assertEquals("final answer", recovered?.last()?.contentText)
}
@Test
fun cancelStopsPollingWithoutAnyTerminalCallback() = runTest {
val script = ScriptedHistory(listOf({ listOf(user("u1")) }))
var recovered = false
var gaveUp = false
val recovery = ChatStreamRecovery(this, script.fetcher { testScheduler.currentTime })
recovery.start(pending, 0, {}, { recovered = true }, { gaveUp = true })
advanceTimeBy(5_001)
assertEquals(1, script.fetchCount)
recovery.cancel()
advanceTimeBy(60L * 60_000)
assertEquals("no polls after cancel", 1, script.fetchCount)
assertFalse(recovered)
assertFalse(gaveUp)
assertFalse(recovery.isActive)
}
}
@@ -0,0 +1,111 @@
package com.hermesandroid.relay.viewmodel
import com.hermesandroid.relay.network.upstream.ChatHandler
import com.hermesandroid.relay.network.upstream.HermesApiClient
import com.hermesandroid.relay.network.relay.RealtimeVoiceEvent
import okhttp3.mockwebserver.Dispatcher
import okhttp3.mockwebserver.MockResponse
import okhttp3.mockwebserver.MockWebServer
import okhttp3.mockwebserver.RecordedRequest
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.robolectric.RobolectricTestRunner
import org.robolectric.annotation.Config
@RunWith(RobolectricTestRunner::class)
@Config(sdk = [34])
class ChatViewModelRealtimeTurnTest {
private lateinit var server: MockWebServer
private lateinit var handler: ChatHandler
private lateinit var viewModel: ChatViewModel
@Before
fun setUp() {
server = MockWebServer().apply {
dispatcher = object : Dispatcher() {
override fun dispatch(request: RecordedRequest): MockResponse =
MockResponse().setResponseCode(404)
}
start()
}
handler = ChatHandler()
viewModel = ChatViewModel().also {
it.initialize(HermesApiClient(server.url("/").toString(), "test-key"), handler)
}
}
@After
fun tearDown() {
server.shutdown()
}
@Test
fun transportFailureSettlesRealtimePlaceholder() {
val assistantId = viewModel.startRealtimeAgentTurn(userText = "", chatSessionId = "session-1")
viewModel.failRealtimeAgentTurn(
assistantId,
"Voice connection was interrupted. Tap the mic to try again.",
)
val assistant = handler.messages.value.single { it.id == assistantId }
assertFalse(handler.isStreaming.value)
assertFalse(assistant.isStreaming)
assertEquals("Voice connection was interrupted. Tap the mic to try again.", assistant.content)
assertTrue("Error" in assistant.badges)
assertFalse(handler.messages.value.any { it.content == "Listening..." })
}
@Test
fun localStopSettlesRealtimePlaceholder() {
val assistantId = viewModel.startRealtimeAgentTurn(userText = "", chatSessionId = "session-1")
viewModel.cancelRealtimeAgentTurnLocally(assistantId)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(type = "voice.response.delta", delta = "late response", raw = "{}"),
)
val assistant = handler.messages.value.single { it.id == assistantId }
assertFalse(handler.isStreaming.value)
assertFalse(assistant.isStreaming)
assertEquals("Cancelled.", assistant.content)
assertFalse(handler.messages.value.any { it.content == "Listening..." })
}
@Test
fun normalResponseCompletionAllowsLaterBackgroundSummaryOnSameTurn() {
val assistantId = viewModel.startRealtimeAgentTurn(userText = "Check Hermes", chatSessionId = "session-1")
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(type = "voice.response.delta", delta = "I'll check.", raw = "{}"),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(type = "voice.response.done", raw = "{}"),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "voice.response.started",
provider = "xai_realtime",
raw = "{}",
),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(type = "voice.response.delta", delta = " Final answer.", raw = "{}"),
)
val assistant = handler.messages.value.single { it.id == assistantId }
assertTrue(handler.isStreaming.value)
assertEquals("I'll check. Final answer.", assistant.content)
}
}
@@ -0,0 +1,279 @@
package com.hermesandroid.relay.viewmodel
import android.os.Looper
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.network.upstream.ChatHandler
import com.hermesandroid.relay.network.upstream.HermesApiClient
import java.time.Duration
import java.util.concurrent.atomic.AtomicInteger
import okhttp3.mockwebserver.Dispatcher
import okhttp3.mockwebserver.MockResponse
import okhttp3.mockwebserver.MockWebServer
import okhttp3.mockwebserver.RecordedRequest
import okhttp3.mockwebserver.SocketPolicy
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNotNull
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Assert.fail
import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.robolectric.RobolectricTestRunner
import org.robolectric.Shadows.shadowOf
import org.robolectric.annotation.Config
/**
* End-to-end coverage for the issue #166 dropped-stream answer recovery:
* a real [HermesApiClient] against a [MockWebServer] whose sessions SSE
* route dies mid-turn (transport error, NOT a server error), with the
* `/api/sessions/{id}/messages` route scripted so the poller first finds
* no answer and then the persisted final answer.
*
* Runs under Robolectric so the client's main-thread Handler dispatch and
* `viewModelScope` (Dispatchers.Main → Robolectric main looper) are real;
* the test drives time by idling the main looper. Recovery cadence is
* shrunk via [ChatViewModel.recoveryTimingOverride] so polls take
* milliseconds instead of seconds.
*/
@RunWith(RobolectricTestRunner::class)
@Config(sdk = [34])
class ChatViewModelStreamRecoveryTest {
private companion object {
const val SESSION_ID = "s1"
const val PROMPT = "hello there"
const val RECOVERED = "recovered answer"
const val USER_ONLY_BODY =
"""{"data":[{"id":"u1","role":"user","content":"$PROMPT","timestamp":1000.0}]}"""
const val WITH_ANSWER_BODY =
"""{"data":[""" +
"""{"id":"u1","role":"user","content":"$PROMPT","timestamp":1000.0},""" +
"""{"id":"a1","role":"assistant","content":"$RECOVERED","timestamp":1001.0}]}"""
const val EMPTY_BODY = """{"data":[]}"""
}
private lateinit var server: MockWebServer
private lateinit var vm: ChatViewModel
private lateinit var handler: ChatHandler
private val messagesRequests = AtomicInteger(0)
private val streamRequests = AtomicInteger(0)
/** GET /messages body for the [n]-th poll (1-based). */
@Volatile
private var messagesBodyFor: (Int) -> String = { USER_ONLY_BODY }
/** Non-first POST /chat/stream responses complete cleanly when true. */
@Volatile
private var secondStreamSucceeds = false
@Before
fun setUp() {
server = MockWebServer()
server.dispatcher = object : Dispatcher() {
override fun dispatch(request: RecordedRequest): MockResponse {
val path = request.path ?: return MockResponse().setResponseCode(404)
return when {
request.method == "POST" &&
path == "/api/sessions/$SESSION_ID/chat/stream" -> {
val n = streamRequests.incrementAndGet()
if (n > 1 && secondStreamSucceeds) {
MockResponse()
.setResponseCode(200)
.setHeader("Content-Type", "text/event-stream")
.setBody("data: [DONE]\n\n")
} else {
// Transport death mid-turn: advertise a full body
// but cut the socket halfway — the SSE listener's
// onFailure fires with an IOException, exactly the
// Doze / Wi-Fi power-save drop class.
MockResponse()
.setResponseCode(200)
.setHeader("Content-Type", "text/event-stream")
.setBody(": keepalive\n\n: keepalive\n\n: keepalive\n\n")
.setSocketPolicy(SocketPolicy.DISCONNECT_DURING_RESPONSE_BODY)
}
}
request.method == "GET" &&
path == "/api/sessions/$SESSION_ID/messages" -> {
MockResponse()
.setResponseCode(200)
.setHeader("Content-Type", "application/json")
.setBody(messagesBodyFor(messagesRequests.incrementAndGet()))
}
else -> MockResponse().setResponseCode(404)
}
}
}
server.start()
handler = ChatHandler()
vm = ChatViewModel()
vm.streamingEndpoint = "sessions"
vm.recoveryTimingOverride = ChatStreamRecovery.Timing(
pollIntervalMs = 150,
maxPollIntervalMs = 150,
recoveryWindowMs = 60_000,
)
vm.initialize(HermesApiClient(server.url("/").toString(), "test-key"), handler)
handler.setSessionId(SESSION_ID)
idle()
}
@After
fun tearDown() {
runCatching { server.shutdown() }
idle()
}
/** Run queued main-looper tasks and advance the Robolectric clock a bit. */
private fun idle(ms: Long = 200) {
shadowOf(Looper.getMainLooper()).idleFor(Duration.ofMillis(ms))
}
/**
* Idle the main looper (advancing virtual time so coroutine delays fire)
* while real IO threads make progress, until [condition] holds.
*/
private fun awaitCondition(what: String, timeoutMs: Long = 20_000, condition: () -> Boolean) {
val deadline = System.currentTimeMillis() + timeoutMs
while (System.currentTimeMillis() < deadline) {
idle()
if (condition()) return
Thread.sleep(15)
}
fail("Timed out waiting for: $what")
}
/** Settle, snapshot the poll counter, idle on, and assert it stayed put. */
private fun assertPollingStopped() {
// Let any already-in-flight poll (or post-turn reconcile read) land.
Thread.sleep(150)
idle(500)
val settled = messagesRequests.get()
repeat(10) {
idle(300)
Thread.sleep(15)
}
assertEquals("poller must not keep hitting /messages", settled, messagesRequests.get())
}
@Test
fun streamKilledMidTurn_recoversPersistedAnswerAndCompletesPlaceholder() {
// Poll 1 has no answer yet; the answer is persisted from poll 2 on.
messagesBodyFor = { n -> if (n < 2) USER_ONLY_BODY else WITH_ANSWER_BODY }
vm.sendMessage(PROMPT)
idle()
awaitCondition("streaming placeholder shown") {
handler.messages.value.any { it.role == MessageRole.ASSISTANT && it.isStreaming }
}
awaitCondition("recovery started after the transport drop") {
vm.recoveringAnswer.value
}
awaitCondition("recovered answer reconciled + turn finalized") {
!vm.recoveringAnswer.value &&
!handler.isStreaming.value &&
handler.messages.value.any {
it.role == MessageRole.ASSISTANT && it.content == RECOVERED && !it.isStreaming
}
}
// Stability requires the answer on two consecutive polls, and the
// first poll had none — at least 3 polls total.
assertTrue("expected >= 3 polls, saw ${messagesRequests.get()}", messagesRequests.get() >= 3)
assertNull("a recovered turn must not surface an error", handler.error.value)
assertFalse(
"no streaming placeholder may survive recovery",
handler.messages.value.any { it.isStreaming },
)
}
@Test
fun userCancelDuringRecovery_abortsThePoller() {
messagesBodyFor = { USER_ONLY_BODY } // answer never arrives
vm.sendMessage(PROMPT)
awaitCondition("recovery started") { vm.recoveringAnswer.value }
awaitCondition("at least one poll issued") { messagesRequests.get() >= 1 }
vm.cancelStream()
idle()
assertFalse(vm.recoveringAnswer.value)
assertFalse(handler.isStreaming.value)
assertNull("user cancel is not an error", handler.error.value)
assertTrue(
"the cancelled placeholder should carry the Stopped badge",
handler.messages.value.any { it.role == MessageRole.ASSISTANT && "Stopped" in it.badges },
)
assertPollingStopped()
}
@Test
fun sessionSwitchDuringRecovery_leavesChatNotStreamingWithNoStuckStatus() {
messagesBodyFor = { USER_ONLY_BODY } // answer never arrives
vm.sendMessage(PROMPT)
awaitCondition("recovery started") { vm.recoveringAnswer.value }
awaitCondition("at least one poll issued") { messagesRequests.get() >= 1 }
// Switch away while recovery is polling — there is NO live stream, so
// aborting the poller must itself settle the streaming/turn-status
// state instead of wedging the chat "streaming forever". (s2 has no
// scripted /messages route → the history load 404s to empty, which is
// fine: this asserts the abort settles, not that s2 loads anything.)
vm.switchSession("s2")
idle()
assertFalse("recovery poller must be aborted", vm.recoveringAnswer.value)
assertFalse("chat must not be stuck streaming", handler.isStreaming.value)
assertNull("no stuck 'Reconnecting…' turn status", handler.turnStatus.value)
assertNull("a silent abandon is not an error", handler.error.value)
assertPollingStopped()
}
@Test
fun newSendDuringRecovery_abortsThePoller() {
messagesBodyFor = { USER_ONLY_BODY } // first turn's answer never arrives
secondStreamSucceeds = true
vm.sendMessage(PROMPT)
awaitCondition("recovery started") { vm.recoveringAnswer.value }
awaitCondition("at least one poll issued") { messagesRequests.get() >= 1 }
vm.sendMessage("a second question")
idle()
assertFalse("a new send must abort the poller", vm.recoveringAnswer.value)
awaitCondition("second turn completes") { !handler.isStreaming.value }
assertPollingStopped()
}
@Test
fun recoveryWindowExpiry_fallsBackToTheErrorUi() {
// Empty transcript = "no information" → the poller keeps trying until
// the (shrunken) window expires, then the existing error UI fires.
messagesBodyFor = { EMPTY_BODY }
vm.recoveryTimingOverride = ChatStreamRecovery.Timing(
pollIntervalMs = 100,
maxPollIntervalMs = 100,
recoveryWindowMs = 350,
)
vm.sendMessage(PROMPT)
awaitCondition("recovery started") { vm.recoveringAnswer.value }
awaitCondition("cap expiry surfaces the error") { handler.error.value != null }
assertFalse(vm.recoveringAnswer.value)
assertFalse(handler.isStreaming.value)
assertNotNull(handler.error.value)
assertPollingStopped()
}
}
@@ -61,6 +61,58 @@ class RealtimeTurnSyncBuilderTest {
assertFalse(RealtimeTurnSyncBuilder.hasUnsynced(history))
}
// --- stripProvenanceMarker ---
@Test
fun stripProvenanceMarker_roundTripsBuilderOutput() {
val history = listOf(
chatMessage(
id = "rt-1",
role = MessageRole.ASSISTANT,
realtimeTurn = RealtimeTurnTrace(
userText = "What does the Bitwarden integration do?",
assistantText = "It syncs vault metadata.",
provider = "xai_realtime",
model = "grok-voice-latest",
voice = "leo",
),
),
)
val assistantContent = RealtimeTurnSyncBuilder.buildSyntheticMessages(history)[1]
.jsonObject["content"]?.jsonPrimitive?.content.orEmpty()
assertEquals(
"It syncs vault metadata.",
RealtimeTurnSyncBuilder.stripProvenanceMarker(assistantContent),
)
}
@Test
fun stripProvenanceMarker_returnsNullWithoutMarker() {
assertEquals(null, RealtimeTurnSyncBuilder.stripProvenanceMarker("Just a normal reply."))
assertEquals(null, RealtimeTurnSyncBuilder.stripProvenanceMarker(""))
}
@Test
fun stripProvenanceMarker_ignoresMarkerPhraseMidText() {
// The phrase appears but is NOT a final bracket block — prose that
// mentions it (or a marker followed by more text) must not be stripped.
val midText = "The [Realtime Agent provider-native voice turn: provider=x] " +
"marker is how sync works."
assertEquals(null, RealtimeTurnSyncBuilder.stripProvenanceMarker(midText))
val markerThenMore = "Answer.\n\n[Realtime Agent provider-native voice turn: " +
"provider=x]\n\nMore prose after."
assertEquals(null, RealtimeTurnSyncBuilder.stripProvenanceMarker(markerThenMore))
}
@Test
fun stripProvenanceMarker_toleratesTrailingWhitespace() {
val content = "Answer.\n\n[Realtime Agent provider-native voice turn: provider=x]\n "
assertEquals("Answer.", RealtimeTurnSyncBuilder.stripProvenanceMarker(content))
}
private fun chatMessage(
id: String,
role: MessageRole,
@@ -0,0 +1,240 @@
package com.hermesandroid.relay.voice
import android.app.Application
import android.util.Log
import com.hermesandroid.relay.network.relay.RealtimeAgentSessionControl
import com.hermesandroid.relay.network.relay.VoiceHandoffEvent
import com.hermesandroid.relay.viewmodel.BackgroundRunPhase
import com.hermesandroid.relay.viewmodel.BackgroundRunState
import com.hermesandroid.relay.viewmodel.VoiceViewModel
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkStatic
import io.mockk.verify
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.ExperimentalCoroutinesApi
import kotlinx.coroutines.async
import kotlinx.coroutines.test.UnconfinedTestDispatcher
import kotlinx.coroutines.test.advanceTimeBy
import kotlinx.coroutines.test.resetMain
import kotlinx.coroutines.test.runCurrent
import kotlinx.coroutines.test.runTest
import kotlinx.coroutines.test.setMain
import okhttp3.WebSocket
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
import java.util.concurrent.CountDownLatch
import java.util.concurrent.TimeUnit
import java.util.concurrent.atomic.AtomicReference
@OptIn(ExperimentalCoroutinesApi::class)
class VoiceViewModelRealtimeSessionFenceTest {
private val mainDispatcher = UnconfinedTestDispatcher()
private fun awaitBlocked(thread: AtomicReference<Thread?>): Boolean {
val deadlineNanos = System.nanoTime() + TimeUnit.SECONDS.toNanos(2)
while (System.nanoTime() < deadlineNanos) {
if (thread.get()?.state == Thread.State.BLOCKED) return true
Thread.sleep(5L)
}
return false
}
@Before
fun setUp() {
Dispatchers.setMain(mainDispatcher)
mockkStatic(Log::class)
every { Log.i(any(), any<String>()) } returns 0
every { Log.w(any(), any<String>()) } returns 0
every { Log.e(any(), any<String>()) } returns 0
every { Log.d(any(), any<String>()) } returns 0
}
@After
fun tearDown() {
Dispatchers.resetMain()
unmockkStatic(Log::class)
}
@Test
fun staleRealtimeCallbackCannotRepopulateNewVoiceSession() = runTest {
val viewModel = VoiceViewModel(mockk<Application>(relaxed = true))
viewModel.enterVoiceMode()
val staleGeneration = viewModel.realtimeSessionGenerationForTest()
viewModel.exitVoiceMode()
viewModel.enterVoiceMode()
viewModel.recordVoiceHandoffForTest(
staleGeneration,
VoiceHandoffEvent(label = "Waiting for route"),
)
assertTrue(viewModel.uiState.value.voiceMode)
assertNull(viewModel.uiState.value.handoffStatus)
assertNull(viewModel.uiState.value.backgroundRun)
}
@Test
fun olderRouteTransitionCannotOverwriteConfirmedReconnect() = runTest {
val viewModel = VoiceViewModel(mockk<Application>(relaxed = true))
viewModel.enterVoiceMode()
val generation = viewModel.realtimeSessionGenerationForTest()
viewModel.recordVoiceHandoffForTest(
generation,
VoiceHandoffEvent(
label = "Voice reconnected",
active = false,
success = true,
transitionRevision = 4L,
),
)
viewModel.recordVoiceHandoffForTest(
generation,
VoiceHandoffEvent(
label = "Waiting for route",
transitionRevision = 3L,
),
)
assertEquals("Voice reconnected", viewModel.uiState.value.handoffStatus?.title)
assertTrue(viewModel.uiState.value.handoffStatus?.success == true)
}
@Test
fun terminalHandoffCannotLeaveBackgroundRunReconnecting() = runTest {
val socket = mockk<WebSocket>(relaxed = true)
val viewModel = VoiceViewModel(mockk<Application>(relaxed = true))
viewModel.enterVoiceMode()
val generation = viewModel.realtimeSessionGenerationForTest()
viewModel.seedBackgroundRunForTest(
run = BackgroundRunState(
runId = "run-terminal-handoff",
phase = BackgroundRunPhase.RECONNECTING,
),
control = RealtimeAgentSessionControl(socket),
)
viewModel.recordVoiceHandoffForTest(
generation,
VoiceHandoffEvent(
label = "Voice handoff failed",
active = false,
transitionRevision = 5L,
),
)
assertNull(viewModel.uiState.value.backgroundRun)
assertEquals("Voice handoff failed", viewModel.uiState.value.handoffStatus?.title)
}
@Test
fun exitSerializesWithCallbackThatAlreadyPassedGenerationCheck() = runTest {
val viewModel = VoiceViewModel(mockk<Application>(relaxed = true))
viewModel.enterVoiceMode()
val generation = viewModel.realtimeSessionGenerationForTest()
val reporterEntered = CountDownLatch(1)
val releaseReporter = CountDownLatch(1)
val exitStarted = CountDownLatch(1)
val exitThread = AtomicReference<Thread?>()
viewModel.setVoiceHandoffReporterForTest {
reporterEntered.countDown()
releaseReporter.await(2, TimeUnit.SECONDS)
}
val callback = async(Dispatchers.Default) {
viewModel.recordVoiceHandoffForTest(
generation,
VoiceHandoffEvent(label = "Waiting for route"),
)
}
assertTrue(reporterEntered.await(2, TimeUnit.SECONDS))
val exit = async(Dispatchers.Default) {
exitThread.set(Thread.currentThread())
exitStarted.countDown()
viewModel.exitVoiceMode()
}
assertTrue(exitStarted.await(2, TimeUnit.SECONDS))
assertTrue("Exit must be waiting on the session lock", awaitBlocked(exitThread))
releaseReporter.countDown()
callback.await()
exit.await()
assertNull(viewModel.uiState.value.handoffStatus)
assertNull(viewModel.uiState.value.backgroundRun)
}
@Test
fun exitDetachesRunPromotedByCallbackAlreadyHoldingSessionLock() = runTest {
val socket = mockk<WebSocket>()
every { socket.send(any<String>()) } returns true
val viewModel = VoiceViewModel(mockk<Application>(relaxed = true))
viewModel.enterVoiceMode()
val generation = viewModel.realtimeSessionGenerationForTest()
val reporterEntered = CountDownLatch(1)
val releaseReporter = CountDownLatch(1)
val exitStarted = CountDownLatch(1)
val exitThread = AtomicReference<Thread?>()
viewModel.setVoiceHandoffReporterForTest {
reporterEntered.countDown()
releaseReporter.await(2, TimeUnit.SECONDS)
viewModel.seedBackgroundRunForTest(
run = BackgroundRunState(
runId = "run-promoted-during-exit",
phase = BackgroundRunPhase.RUNNING,
),
control = RealtimeAgentSessionControl(socket),
)
}
val callback = async(Dispatchers.Default) {
viewModel.recordVoiceHandoffForTest(
generation,
VoiceHandoffEvent(label = "Waiting for route"),
)
}
assertTrue(reporterEntered.await(2, TimeUnit.SECONDS))
val exit = async(Dispatchers.Default) {
exitThread.set(Thread.currentThread())
exitStarted.countDown()
viewModel.exitVoiceMode()
}
assertTrue(exitStarted.await(2, TimeUnit.SECONDS))
assertTrue("Exit must be waiting on the session lock", awaitBlocked(exitThread))
releaseReporter.countDown()
callback.await()
exit.await()
verify(exactly = 0) { socket.send(match<String> { it.contains("response.cancel") }) }
assertNull(viewModel.uiState.value.backgroundRun)
assertNull(viewModel.uiState.value.handoffStatus)
}
@Test
fun queuedCancelDismissesWhenAcknowledgementNeverArrives() = runTest {
val socket = mockk<WebSocket>()
every { socket.send(any<String>()) } returns true
val viewModel = VoiceViewModel(mockk<Application>(relaxed = true))
viewModel.seedBackgroundRunForTest(
run = BackgroundRunState(
runId = "run-cancel-timeout",
phase = BackgroundRunPhase.RECONNECTING,
),
control = RealtimeAgentSessionControl(socket),
)
viewModel.cancelBackgroundRun()
assertEquals("Cancelling…", viewModel.uiState.value.backgroundRun?.message)
advanceTimeBy(5_001L)
runCurrent()
assertNull(viewModel.uiState.value.backgroundRun)
assertNull(viewModel.uiState.value.handoffStatus)
}
}
+35 -12
View File
@@ -39,14 +39,20 @@ import {
shouldAdvertiseComputerUse
} from '../tools/handlerSet.js'
import { DesktopToolRouter } from '../tools/router.js'
import { RelayTransport } from '../transport/RelayTransport.js'
import { PROMPT_SUBMIT_REQUEST_TIMEOUT_MS, RelayTransport } from '../transport/RelayTransport.js'
// (getSession is imported above with the other remoteSessions exports so we
// can render the endpoint-role banner without changing the saveSession/auth
// persistence path.)
const READY_TIMEOUT_MS = 60_000
const TURN_TIMEOUT_MS = 10 * 60_000
// Idle-progress watchdog, NOT a hard turn cap: the timer is re-armed on every
// gateway event, so a turn only dies after this long with NO events at all.
// A long MoA/tool-heavy turn that keeps streaming lives indefinitely —
// upstream treats prompt.submit as fire-and-forget (completion arrives via
// message.complete, not the RPC ack), so wall-clock capping a healthy turn
// killed legitimate work.
const TURN_IDLE_TIMEOUT_MS = 10 * 60_000
function flag(args: ParsedArgs, name: string): string | null {
const v = args.flags[name]
@@ -193,15 +199,22 @@ function runOneTurn(
let detach: (() => void) | null = null
const promise = new Promise<void>((resolve, reject) => {
const timer = setTimeout(() => {
if (settled) {
return
}
settled = true
detach?.()
reject(new Error(`turn timeout after ${TURN_TIMEOUT_MS}ms`))
}, TURN_TIMEOUT_MS)
timer.unref?.()
// Idle-progress watchdog: re-armed on every gateway event below. Fires
// only when the stream has gone completely silent for the window.
let timer: ReturnType<typeof setTimeout> | undefined
const armIdleTimer = () => {
clearTimeout(timer)
timer = setTimeout(() => {
if (settled) {
return
}
settled = true
detach?.()
reject(new Error(`turn idle timeout — no gateway events for ${TURN_IDLE_TIMEOUT_MS}ms`))
}, TURN_IDLE_TIMEOUT_MS)
timer.unref?.()
}
armIdleTimer()
const handler = (ev: GatewayEvent) => {
renderer.handle(ev)
@@ -209,6 +222,10 @@ function runOneTurn(
return
}
// Any event (delta, tool progress, status, heartbeat) proves the turn
// is alive — push the idle deadline out.
armIdleTimer()
if (ev.type === 'message.complete') {
settled = true
clearTimeout(timer)
@@ -232,7 +249,13 @@ function runOneTurn(
detach = () => gw.off('event', handler)
gw.on('event', handler)
gw.request<PromptSubmitResponse>('prompt.submit', { session_id: sessionId, text: prompt }).catch((e: unknown) => {
// Long-running RPC: the ack can trail the turn by minutes (see
// PROMPT_SUBMIT_REQUEST_TIMEOUT_MS) — liveness is the idle watchdog's job.
gw.request<PromptSubmitResponse>(
'prompt.submit',
{ session_id: sessionId, text: prompt },
PROMPT_SUBMIT_REQUEST_TIMEOUT_MS
).catch((e: unknown) => {
if (settled) {
return
}
+2 -2
View File
@@ -26,8 +26,8 @@ export class GatewayClient extends EventEmitter {
this.transport.start()
}
request<T = unknown>(method: string, params: Record<string, unknown> = {}): Promise<T> {
return this.transport.request<T>(method, params)
request<T = unknown>(method: string, params: Record<string, unknown> = {}, timeoutMs?: number): Promise<T> {
return this.transport.request<T>(method, params, timeoutMs)
}
drain() {
+22 -2
View File
@@ -25,6 +25,18 @@ const MAX_BUFFERED_EVENTS = 2000
const REQUEST_TIMEOUT_MS = Math.max(30000, parseInt(process.env.HERMES_RELAY_RPC_TIMEOUT_MS ?? '120000', 10) || 120000)
const AUTH_TIMEOUT_MS = Math.max(5000, parseInt(process.env.HERMES_RELAY_AUTH_TIMEOUT_MS ?? '15000', 10) || 15000)
// `prompt.submit` ack ceiling — mirrors upstream desktop's
// PROMPT_SUBMIT_REQUEST_TIMEOUT_MS (apps/desktop/src/hermes.ts, upstream
// commit 164144183). The submit is effectively fire-and-forget: turn
// completion arrives via stream events (message.complete), NOT the RPC
// return, and MoA/deep-reasoning/tool-heavy turns can take minutes to ack.
// Bounding the ack by the generic REQUEST_TIMEOUT_MS (120s) killed
// legitimately long turns. Matches the backend's agent-turn ceiling
// (agent.gateway_timeout = 1800s), so this only fires when the turn would
// have been abandoned server-side anyway. Callers pass it as the per-call
// `timeoutMs` on `request('prompt.submit', …)`.
export const PROMPT_SUBMIT_REQUEST_TIMEOUT_MS = 1_800_000
// Reconnect knobs — mirrored from Android ConnectionManager.kt.
const RECONNECT_BASE_MS = 1000
const RECONNECT_MAX_MS = 30_000
@@ -1014,7 +1026,15 @@ export class RelayTransport extends EventEmitter implements Transport {
return this.logs.tail(Math.max(1, limit)).join('\n')
}
request<T = unknown>(method: string, params: Record<string, unknown> = {}): Promise<T> {
/** Send one JSON-RPC request. `timeoutMs` overrides the generic
* env-tunable default for long-running RPCs — `prompt.submit` passes
* PROMPT_SUBMIT_REQUEST_TIMEOUT_MS because its ack can trail the turn by
* minutes; everything else should omit it. */
request<T = unknown>(
method: string,
params: Record<string, unknown> = {},
timeoutMs: number = REQUEST_TIMEOUT_MS
): Promise<T> {
if (!this.ws) {
return Promise.reject(new Error('relay transport not connected'))
}
@@ -1026,7 +1046,7 @@ export class RelayTransport extends EventEmitter implements Transport {
const id = `r${++this.reqId}`
return new Promise<T>((resolve, reject) => {
const timeout = setTimeout(this.onTimeout, REQUEST_TIMEOUT_MS, id)
const timeout = setTimeout(this.onTimeout, timeoutMs, id)
timeout.unref?.()
this.pending.set(id, {
+5 -2
View File
@@ -13,8 +13,11 @@ import type { GatewayEvent } from '../gatewayTypes.js'
export interface Transport {
/** Drop in-flight state and start the carrier (open socket). */
start(): void
/** Send a JSON-RPC request; resolves with `result` or rejects with the server's error. */
request<T = unknown>(method: string, params?: Record<string, unknown>): Promise<T>
/** Send a JSON-RPC request; resolves with `result` or rejects with the
* server's error. `timeoutMs` optionally overrides the transport's generic
* request timeout for long-running RPCs (e.g. `prompt.submit`, whose ack
* can legitimately trail the turn by minutes). */
request<T = unknown>(method: string, params?: Record<string, unknown>, timeoutMs?: number): Promise<T>
/** Attach event/exit listeners. */
on(event: 'event', handler: (ev: GatewayEvent) => void): void
on(event: 'exit', handler: (code: number | null) => void): void
+28 -8
View File
@@ -35,8 +35,13 @@ import { renderVoicePage } from './voicePage.js'
import type { GatewayClient } from './gatewayClient.js'
import type { GatewayEvent, PromptSubmitResponse } from './gatewayTypes.js'
import { rpcErrorMessage } from './lib/rpc.js'
import { PROMPT_SUBMIT_REQUEST_TIMEOUT_MS } from './transport/RelayTransport.js'
const TURN_TIMEOUT_MS = 5 * 60_000
// Idle-progress watchdog, NOT a hard turn cap: re-armed on every gateway
// event, so a voice turn only dies after this long with NO events at all.
// The prompt.submit ack itself is bounded separately (it can trail the turn
// by minutes — see PROMPT_SUBMIT_REQUEST_TIMEOUT_MS).
const TURN_IDLE_TIMEOUT_MS = 5 * 60_000
export interface VoiceServerOptions {
/** Relay bearer token. Used as `Authorization: Bearer <token>` on voice routes. */
@@ -207,12 +212,19 @@ async function handleTurn(req: IncomingMessage, res: ServerResponse, ctx: Handle
let completed = false
const settle = new Promise<void>((resolve, reject) => {
const timer = setTimeout(() => {
if (completed) return
cleanup()
reject(new Error(`turn timeout after ${TURN_TIMEOUT_MS}ms`))
}, TURN_TIMEOUT_MS)
timer.unref?.()
// Idle-progress watchdog: re-armed on every gateway event so a long
// tool-heavy turn that keeps streaming is never wall-clock capped.
let timer: ReturnType<typeof setTimeout> | undefined
const armIdleTimer = () => {
clearTimeout(timer)
timer = setTimeout(() => {
if (completed) return
cleanup()
reject(new Error(`turn idle timeout — no gateway events for ${TURN_IDLE_TIMEOUT_MS}ms`))
}, TURN_IDLE_TIMEOUT_MS)
timer.unref?.()
}
armIdleTimer()
const onAbort = () => {
if (completed) return
@@ -224,6 +236,8 @@ async function handleTurn(req: IncomingMessage, res: ServerResponse, ctx: Handle
const handler = (ev: GatewayEvent) => {
if (completed) return
// Any event proves the turn is alive — push the idle deadline out.
armIdleTimer()
if (ev.type === 'message.delta') {
const t = ev.payload?.text ?? ''
if (t) {
@@ -254,8 +268,14 @@ async function handleTurn(req: IncomingMessage, res: ServerResponse, ctx: Handle
ctx.opts.gateway.on('event', handler)
abort.signal.addEventListener('abort', onAbort)
// Long-running RPC: the ack can trail the turn by minutes — liveness
// is the idle watchdog's job, not the ack timeout's.
ctx.opts.gateway
.request<PromptSubmitResponse>('prompt.submit', { session_id: ctx.opts.sessionId, text: transcript })
.request<PromptSubmitResponse>(
'prompt.submit',
{ session_id: ctx.opts.sessionId, text: transcript },
PROMPT_SUBMIT_REQUEST_TIMEOUT_MS
)
.catch((e: unknown) => {
if (completed) return
completed = true
+24 -18
View File
@@ -126,29 +126,34 @@ POST /api/sessions/{session_id}/fork
```
The native upstream list envelope is `{"object":"list","data":[...]}`.
Older fork/bootstrap builds may return `items`, `sessions`, or `messages`;
clients should continue accepting those as compatibility shapes.
Older fork builds and pre-retirement bootstrap versions (the current bootstrap
no longer injects session CRUD at all) may return `items`, `sessions`, or
`messages`; clients should continue accepting those as compatibility shapes.
### Chat (Non-Streaming)
```
POST /api/sessions/{session_id}/chat
Content-Type: application/json
Body: {
"message": "Hello",
"model": "claude-opus-4-6", // optional override
"system_message": "...", // optional ephemeral system prompt
"enabled_toolsets": ["hermes-cli"], // optional
"disabled_toolsets": [], // optional
"skip_context_files": false, // optional
"skip_memory": false, // optional
"attachments": [ // optional image attachments
{
"contentType": "image/png",
"content": "<base64-data>"
}
]
"message": "Hello", // required (alias: "input") — plain string, or
// OpenAI-style content parts (text + image_url,
// including data:image/... URLs)
"system_message": "..." // optional ephemeral per-turn system prompt
// (alias: "instructions"; string only)
}
```
`message`/`input` and `system_message`/`instructions` are the ONLY body
fields current native upstream parses on the session chat endpoints
(`_handle_session_chat` / `_handle_session_chat_stream` in
`gateway/platforms/api_server.py`). Legacy fork builds additionally honored
`model`, `attachments` (`{contentType, content}` base64 objects),
`enabled_toolsets`, `disabled_toolsets`, `skip_context_files`, and
`skip_memory` — native upstream ignores all of them. The Android client
still sends `model` + `profile` as best-effort hints for those builds but
nothing else (HRUI-001).
```
-> {
"session_id": "sess_...",
"run_id": "run_...",
@@ -287,10 +292,11 @@ GET /api/memory?target=memory // or target=user
```
GET /v1/skills # native upstream read-only list
GET /v1/toolsets # native upstream toolset inventory
GET /api/skills # legacy compatibility list, optional ?category= filter
GET /api/skills/{name}
GET /api/skills/{name} # legacy detail view (bootstrap compatibility)
```
> `/api/skills/categories` was removed from upstream as dead code (commit 8d023e43) and is not re-injected by the bootstrap.
> The legacy `GET /api/skills` list was retired from the bootstrap — use native
> `/v1/skills`. `/api/skills/categories` was removed from upstream as dead code
> (commit 8d023e43) and is not re-injected by the bootstrap.
## Capability Detection
+31 -12
View File
@@ -569,14 +569,17 @@ We considered four options:
- `install.sh` step 2 — copies the `.pth` into the venv site-packages
**Removal path** is now per surface:
1. Sessions: once the supported Hermes baseline includes #33134, remove the
sessions compatibility handlers and any docs that require bootstrap for
history/chat. Until then, verify native `/api/sessions/*` routes win.
2. Read-only skills/toolsets: clients should prefer native `/v1/skills` and
`/v1/toolsets` from #33016. Retire `/api/skills` list dependence; keep legacy
detail/toggle only if the UI still needs it.
3. Config/memory/available-models: remove those compatibility handlers only
after stable core APIs exist or the dependent Android surfaces are redesigned.
1. Sessions: **done (2026-07-08, HRUI-002).** The sessions CRUD/messages/fork
handlers were removed from the bootstrap with no pre-#33134 fallback kept;
native `/api/sessions/*` (#33134) is the only provider. Older core builds
degrade via the client capability probe to `/v1/chat/completions`/`/v1/runs`.
2. Read-only skills/toolsets: **done (2026-07-08, HRUI-002).** The legacy
`GET /api/skills` list handler was removed; clients use native `/v1/skills`
and `/v1/toolsets` (#33016). Legacy detail (`/api/skills/{name}`) and the
501 toggle stub remain — no native equivalent exists.
3. Config/memory/available-models/session search: remove those compatibility
handlers only after stable core APIs exist or the dependent Android surfaces
are redesigned.
4. Slash middleware: remove after native API-server slash preprocessing exists.
5. Full cleanup: delete `hermes_relay_bootstrap/`, delete
`hermes_relay_bootstrap.pth`, remove the `.pth` install block, and update
@@ -1266,6 +1269,8 @@ We already had a working precedent: the `MEDIA:` marker in assistant text gives
Each [com.hermesandroid.relay.data.HermesCardDispatch] carries a `syncedToServer` flag. On the next chat send, `CardDispatchSyncBuilder.buildSyntheticMessages` materializes every unsynced dispatch into an OpenAI-format `assistant` (with `tool_calls`) + `tool` (with `tool_call_id`) pair under a synthetic tool name `hermes_card_action` — a namespaced name the upstream dispatcher will never try to execute, it's a historical audit record only. The arguments object carries `card_key` / `action_value` / `card_type` / `card_title` / `action_label` / `action_mode` / `action_style` so the LLM has enough context to describe the interaction even if the card itself gets trimmed from rolling window memory. Pairs are spliced into the same request-body slot as voice-intent synthetic messages (`voiceIntentMessages` param — name is historical, the param accepts any synthetic-message JsonArray); once the API client accepts the request, `ChatHandler.markCardDispatchesSynced` flips every dispatch's flag so subsequent turns don't re-emit. Commit-timing matches the voice-intent path exactly (post-handoff) so a thrown request-building exception leaves dispatches unsynced for the next try.
*Update (2026-07, HRUI-001):* the top-level `messages` array these pairs originally rode on the sessions/runs SSE payloads was never parsed by native upstream — the context was silently dropped on those transports. The payload builders (`HermesChatPayloads.kt`) now deliver synthetic history through channels upstream actually consumes: tool-call pairs render as a plain-text digest folded into the per-turn ephemeral system prompt (`system_message` on sessions, `instructions` on runs, the `system` message on completions), and plain realtime-voice turns ride a real history channel where one exists (completions `messages` splice, runs `conversation_history`). The digest is per-turn context, not persisted server-side session history.
**Phase B (deferred — not v0.7.x).**
Contribute a `gateway/rich_cards.py` helper upstream + Discord/Slack adapter translations. Discord gains its first real embed usage; Slack reuses the existing Block Kit path. Plain-text platforms (Signal, SMS) fall back to a markdown render of the same card. Same playbook as the current compatibility-overlay model: ship locally while the shape is proving out, then retire the local marker path once a released core build exposes the native card surface. Held until real phone-side card usage surfaces concrete fidelity issues worth translating for.
@@ -1655,8 +1660,10 @@ override:
promotion vs. silent + visual only.
- `progress_spoken_after_ms` / `progress_repeat_ms` - reuse existing
`_HERMES_SPOKEN_PROGRESS_*` knobs, now configurable.
- `result_delivery` - `speak_when_idle` (default) vs. `notify_then_speak`
(chime/visual, speak on user re-engage) vs. `visual_only`.
- `result_delivery` - `speak_verbatim` (default; the realtime provider reads
the authoritative answer word for word, with relay TTS as the validator's
fallback) vs. `speak_when_idle` (provider/model summary), `notify_then_speak`
(chime/visual, speak on user re-engage), or `visual_only`.
- `max_background_runs` - concurrent background runs per session (default 1 for
the MVP; the existing single-`hermes_task` field assumes 1).
@@ -1680,8 +1687,20 @@ idle with `turn_detection: None` + resume TTL; relay-host probe retained as a
regression check, not a precondition). The premise was also superseded in
implementation: Tier B closes the pending provider call with an interim ack
rather than holding an open response, so the socket only sees the normal
between-turns idle gap — no provider needs the `must-reopen` fallback today, and
default-on is unblocked.
between-turns idle gap. This unblocked default-on at the short-window scale.
See the 2026-07-08 revision below for xAI's later 900s idle-expiry behavior.
**Phase 0 revision (2026-07-08).** The xAI verdict was scoped to between-turn
idle and broke at the 15-minute scale: a live event log showed xAI closing a
quiet conversation with "timed out after 900.0 seconds due to inactivity"
(~896s of zero events after a background-run summary finished speaking).
Follow-up probe runs proved no keepalive works: neither uncommitted silent PCM
nor acknowledged `session.update` pings reset xAI's 900s timer. The broker now
treats an idle timeout as routine provider-session expiry: it logs the expiry,
closes the attached Android websocket cleanly while idle, emits no `voice.error`,
and lets the next user turn open a fresh provider conversation seeded from the
durable Hermes session. Full findings live in `docs/realtime-voice-poc.md` →
"Idle tolerance" → "Revision (2026-07-08)".
**Rules.**
@@ -129,7 +129,7 @@ Relay-side `realtime_voice` config (source of truth, per-profile override via
| `spoken_handoff` | `true` | Speak "I've started that" on promotion |
| `progress_spoken_after_ms` | `15000` | Reuse `_HERMES_SPOKEN_PROGRESS_AFTER_SECONDS` |
| `progress_repeat_ms` | `30000` | Reuse `_HERMES_SPOKEN_PROGRESS_REPEAT_SECONDS` |
| `result_delivery` | `speak_when_idle` | vs `notify_then_speak` / `visual_only` |
| `result_delivery` | `speak_verbatim` | vs `speak_when_idle` / `notify_then_speak` / `visual_only` |
| `max_background_runs` | `1` | Fixed at 1 this plan |
---
@@ -171,8 +171,10 @@ rather than holding it conversational. Capture that in the ADR's Phase 0 line.
provider call with an interim ack* rather than holding an open response, so the
"hold the floor conversational while a run completes" worst case this spike
guarded against does not occur — the socket only sees the normal between-turns
idle gap. Both providers are `hold-floor-ok`, so the default-on gate is satisfied
(no provider needs the `must-reopen` fallback today).
idle gap. Both providers were `hold-floor-ok` for the short-window Phase 0 gate,
so default-on was satisfied. Later 2026-07-08 xAI probes revised the long-idle
behavior to `must-reopen` after the provider's 900s conversation-inactivity
expiry.
---
@@ -0,0 +1,103 @@
# OpenAI Realtime / live voice — research notes (2026-07-08)
Research snapshot mapping OpenAI's realtime voice offering (as of July 2026)
onto the hermes-relay realtime agent, to scope next-release-candidate work.
Companion TODO items live under "OpenAI realtime provider — next-RC roadmap"
in `TODO.md`.
## Repo starting point
`plugin/relay/realtime_agent/providers/openai.py` is a complete,
connection-oriented `RealtimeAgentProvider` implementing the full
`RealtimeAgentConnection` interface (send_audio / commit_audio / send_text /
clear_audio / cancel_response / send_tool_result / request_response / events)
and is wired into the broker alongside xAI. It already uses the GA session
shape (`session.type: "realtime"`, `output_modalities`,
`audio.input/output.format`), per-response `instructions` overrides,
`gpt-realtime-whisper` transcription, and 24kHz PCM16.
Gaps: the default model constant is `gpt-realtime-2` (superseded 2026-07-06
by `gpt-realtime-2.1` / `gpt-realtime-2.1-mini`, drop-in protocol), and no
recorded live voice round has exercised the OpenAI path — every forensics
session in `realtime-agent-runs/` is grok-voice/xAI.
## Findings
**Model lineup** (speech-to-speech, single-model, not cascaded):
- `gpt-realtime` — GA snapshot `gpt-realtime-2025-08-28`, 32k context,
WebRTC/WebSocket/SIP.
<https://developers.openai.com/api/docs/models/gpt-realtime>
- `gpt-realtime-2` (GPT-5-class reasoning), `gpt-realtime-translate`,
`gpt-realtime-whisper`.
<https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/>
- `gpt-realtime-2.1` + `gpt-realtime-2.1-mini` (2026-07-06): configurable
reasoning effort, better interruption/noise/alphanumeric handling,
p95 latency −25%.
**Transports / audio:** WebRTC, WebSocket, SIP; PCM16 @ 24kHz (g711 for
telephony). The relay uses WebSocket + PCM16 @ 24kHz — aligned.
**Session lifecycle:** hard **60-minute wall-clock cap** regardless of
activity (raised from 30). `turn_detection.idle_timeout_ms` exists but only
under `server_vad` — N/A to the relay, which runs `turn_detection: null` and
owns turn-taking. No resume token: an in-flight response survives a socket
drop, but a capped/dropped session must be rebuilt via
`conversation.item.create` / `response.create.input` / `item_reference`.
<https://developers.openai.com/api/docs/guides/conversation-state>
**Turn-taking:** `server_vad` / `semantic_vad` (content-based end-of-turn) /
`none`. Barge-in = `input_audio_buffer.speech_started` →
`conversation.item.truncate` when provider VAD is on. The relay uses `none`
and owns the floor (`RealtimeFloor`) — deliberate; provider VAD/truncate
unused. <https://platform.openai.com/docs/guides/realtime-vad>
**Per-response instructions + out-of-band responses:**
`response.create.instructions` override (already used by the broker), plus
out-of-band responses — `"conversation": "none"` with a custom `"input"`
array (`item_reference`, new messages, or `[]`) and `"metadata"`. Strictly
more control than xAI exposes; directly relevant to exact-answer delivery.
<https://developers.openai.com/api/docs/guides/realtime-conversations>
**Function calling:** standard `function_call` → `function_call_output`; GA
allows the session to continue while a function call is pending (async,
unlike the Responses API). Hosted MCP and image input supported (the relay
keeps non-Hermes tools off).
<https://developers.openai.com/blog/realtime-api>
**Voices / pricing:** `alloy/ash/ballad/coral/echo/sage/shimmer/verse` +
`marin` + `cedar` (Realtime-exclusive). Per 1M tokens — 2.1: audio in $32 /
out $64 / cached $0.40; text $4/$16. 2.1-mini: audio in $10 / out $20 /
cached $0.30. <https://developers.openai.com/api/docs/pricing>
## OpenAI Realtime vs xAI Grok Voice (as the relay uses them)
| Dimension | OpenAI (gpt-realtime-2.1) | xAI (grok-voice-latest) |
|---|---|---|
| Session close | Hard 60-min wall-clock cap, activity-independent | 900s conversation-inactivity close (verified live 4×; active turns stay alive) + ~30-min hard cap per docs |
| Idle knob | `idle_timeout_ms` (server_vad only — N/A) | None; no message resets the 900s timer (verified live) |
| Reconnect/resume | No token; rebuild conversation items on reconnect | Session ends → reseed from the durable Hermes session (current handling) |
| Turn detection | server_vad / semantic_vad / none | none (relay-driven) |
| Duplex/barge-in | `speech_started` + `item.truncate` when VAD on | Provider events; relay owns the floor either way |
| Per-response instructions | Yes, plus out-of-band `conversation:"none"` + `input` | Yes (`instructions`); xAI-only `force_message` provides exact TTS without model inference |
| Async function calls | Yes — session continues while a call is pending | Unverified |
| Pricing | Per-token (see above) | Flat $0.05/min + tool/text tokens separate |
| Reasoning control | 2.1 configurable effort; mini reasons before speaking | No exposed knob |
## Capabilities that could obsolete current workarounds
1. **Forced-summary validation fragility** — grok-voice spoke deferral filler
in most live rounds. xAI Exact mode now uses its provider-native
`force_message` event and bypasses model compliance entirely; OpenAI still
needs the out-of-band response experiment carrying the Hermes answer as
explicit `input` context. The blocklist/validator remains a safety net.
2. **Delivery-note hack** — async function calling makes it possible to leave
a promoted `hermes_run_task` pending and complete it with a real late
`function_call_output`, so the provider's own history reads "done" and
`native_pending_delivery_note` becomes unnecessary (on OpenAI).
3. **Keepalive reasoning doesn't transfer** — OpenAI's failure mode is a
wall-clock cap that can cut an ACTIVE session, unlike xAI's inactivity
timer; the current idle-close handling is xAI-shaped
(`_PROVIDER_IDLE_CLOSE_WS_REASON`) and a focused pass over the broker's
close handling should confirm how a cap-close surfaces before scheduling
the reconnect work.
+5 -3
View File
@@ -85,10 +85,12 @@ This app is a community project and is not affiliated with or endorsed by NousRe
Paste into Play Console → **What's new** (≤500 characters):
```
v1.2.5 — Stability + Try the demo.
v1.4.0 — Realtime voice that finishes the job.
• Fixed a crash that could close the app when a non-URL value (like a label or a line copied from the docs) was entered in a server address field — it now shows an inline error instead.
• New: Try the demo — explore an offline preview of the chat experience with no server or setup, right from the first screen.
• Long voice tasks can queue, keep running while you ask quick follow-ups, and deliver answers in the selected realtime voice.
• Voice sessions recover more reliably after background or route changes and clear stale task states.
• Refresh model catalogs on demand; add opt-in notification rules and multi-device Bridge targeting.
• Safer startup, server-address handling, long chat turns, and credential media access.
```
## Category
+38 -5
View File
@@ -1153,7 +1153,7 @@ strategy:
|---|---|---|
| `hold-floor-ok` | Socket survives idle; post-idle turn clean | Hold the provider session open during the background run (default) |
| `needs-keepalive` | Survives but post-idle turn degraded | Hold open + send a minimal keep-alive; revalidate the first post-idle turn |
| `must-reopen` | Socket closes/errors while idle | Detach the run but close+reopen (or resume) the provider socket on completion |
| `must-reopen` | Socket closes/errors while idle, or no keepalive resets provider expiry | Let the provider socket expire cleanly and reopen/reseed on the next user turn |
### Findings
@@ -1168,7 +1168,8 @@ speaking and the next `response.create`, across the resume TTL).
| Provider | Date | Basis | Windows | Post-idle turn | Verdict |
|---|---|---|---|---|---|
| OpenAI (`gpt-realtime-2`) | 2026-05-24 | empirical (probe) | 10s, 20s, 30s | audio returned, no error | **`hold-floor-ok`** |
| xAI (`grok-voice-latest`) | 2026-05-24 | existing production behavior + protocol parity | between-turn idle in daily use | clean (no idle-close reports) | **`hold-floor-ok`** |
| xAI (`grok-voice-latest`) | 2026-05-24 | existing production behavior + protocol parity | between-turn idle in daily use | clean (no idle-close reports) | ~~`hold-floor-ok`~~ (revised below) |
| xAI (`grok-voice-latest`) | 2026-07-08 | empirical (live log + four relay-host probes) | 900s continuous silence | conversation closed server-side; silent PCM and `session.update` pings also timed out at 900.0s | **`must-reopen`** |
**OpenAI — empirical.** Ran `realtime-provider-idle-probe.py --provider openai
--windows 10,20,30` against the live API. The session stayed open across all
@@ -1188,9 +1189,41 @@ degrading on those idle gaps. Background promotion does not lengthen the
*open-response* duration (the call is closed with an interim ack), so it does not
introduce a new idle condition beyond what xAI already tolerates today. Running
the probe on the relay host is retained as a **regression check**, not a
precondition. If it ever returns `needs-keepalive`/`must-reopen`, set that
provider's `realtime_voice` override accordingly; the per-provider setting
surface already supports it.
precondition. This short-window verdict was later revised for xAI's 900s
continuous-silence expiry; see the revision below.
**Conclusion.** Both verdicts are `hold-floor-ok`, so `promotion_enabled`
defaults **on**. Phase 0's gate is satisfied; default-on is no longer blocked.
### Revision (2026-07-08) — xAI must reopen after ~900s idle expiry
The regression case the 2026-05-24 xAI verdict reserved for has occurred. A
live relay event log showed a session dying while quiet: after a background-run
summary finished speaking, a complete gap of zero events for ~896s ended in
`voice.error` — "xAI Realtime error: Conversation timed out after 900.0
seconds due to inactivity". The 2026-05-24 verdict was scoped to *between-turn*
idle in normal use (tens of seconds); it never probed the 15-minute scale,
because "daily use" never leaves a session silent that long. The client's
manual turn-taking (`turn_detection: None`, mic streamed only while the user
talks) means nothing on the provider socket resets xAI's inactivity timer
during a background wait or an open-but-silent session.
**Further revision (2026-07-08 PM) — no keepalive works.** Four relay-host
probe runs completed against live xAI: the repro died at exactly 900.0s,
silent-PCM appends at 240s/480s/720s also died at exactly 900.0s, and valid
`session.update` pings at 240s/480s/720s (each acknowledged by the server)
still timed out at exactly 900.0s. xAI's conversation-inactivity timer counts
only real conversation items — no side-channel message resets it.
**Final design (shipped in the broker, 2026-07-08 PM).** Keepalive is retired.
An xAI idle timeout is treated as routine provider-session expiry: the broker
logs `voice.realtime_agent.provider_idle_close`, closes the attached Android
websocket cleanly while the user is idle, and does not emit a `voice.error`.
The next user turn opens a fresh provider conversation and is seeded from the
durable Hermes session rather than relying on the expired provider-side
conversation history. This is equivalent to reopening before the deadline for
context preservation, but avoids a timer race and churn while the user is not
speaking.
OpenAI's tolerance at the 900s scale remains unprobed (2026-05-24 ran only
10/20/30s).
+12 -6
View File
@@ -704,7 +704,7 @@ to a tracked background task so the provider event pump stays responsive instead
of blocking on the run:
```json
{"type":"hermes.run.promoted","event_id":20,"source":"hermes","run_id":"...","tier":"promoted","promote_after_ms":6000,"spoken_handoff":true,"result_delivery":"speak_when_idle","call_id":"call-1"}
{"type":"hermes.run.promoted","event_id":20,"source":"hermes","run_id":"...","tier":"promoted","promote_after_ms":6000,"spoken_handoff":true,"result_delivery":"speak_verbatim","call_id":"call-1"}
{"type":"hermes.run.background_completed","event_id":41,"source":"hermes","run_id":"...","ok":true,"tool_count":2}
```
@@ -714,9 +714,12 @@ of blocking on the run:
provider speak a brief "I'm on it" line.
- When the background run finishes, the relay emits
`hermes.run.background_completed`, waits for the audio **floor** to be idle,
then injects the result through the same forced-summary path so the provider
speaks the answer exactly once. `result_delivery` selects `speak_when_idle`
(default), `notify_then_speak`, or `visual_only`.
then delivers the answer exactly once. `result_delivery` selects
`speak_verbatim` (default: the realtime provider reads the authoritative
answer word for word — relay TTS speaks it only as the validator's
fallback), `speak_when_idle` (provider/model summary), `notify_then_speak`,
or `visual_only`. All spoken modes keep the session's realtime voice; the
forced-summary validator + relay-TTS fallback backstop off-script responses.
- `hermes_run_task(mode="background")` skips the grace window and detaches
immediately (`tier:"durable"`), even when grace-period promotion is disabled.
- `hermes.run.progress` carries two extra fields while a run is in flight:
@@ -734,8 +737,11 @@ on `GET/PATCH /voice/realtime-agent/config` as a `promotion` block:
`spoken_handoff`, `progress_spoken_after_ms`, `progress_repeat_ms`,
`result_delivery`, and `max_background_runs`. The default-on path is safe because
it closes the pending call rather than holding an open provider response; the
`scripts/realtime-provider-idle-probe.py` verdict (see `docs/realtime-voice-poc.md`)
confirms per-provider socket survival across the between-turns idle gap.
`scripts/realtime-provider-idle-probe.py` verdict (see
`docs/realtime-voice-poc.md`) confirms per-provider socket behavior. For xAI's
900s conversation-inactivity timeout, the relay treats the provider close as
routine expiry and opens a fresh provider conversation on the next user turn
instead of surfacing a voice error.
If a provider-native realtime turn answers directly without Hermes, Android
keeps the local user/assistant bubbles and marks that provider-only assistant
+3 -2
View File
@@ -38,8 +38,9 @@ relay-token voice fallback use relay pairing.
When in doubt, start with Tailscale. It meets the `tailscale serve`
contract PR #9295 will eventually land upstream — when that happens, our
helper detects the canonical flag and no-ops with a log line, same
auto-retire pattern `hermes_relay_bootstrap/` uses for session-API
endpoints.
auto-retire pattern `hermes_relay_bootstrap/` uses for its remaining
compatibility endpoints (its session-API endpoints have since been fully
retired in favor of native upstream).
## Setup
+19 -12
View File
@@ -29,36 +29,43 @@ AccessibilityService-backed Device Control (screen reading, taps, typing, screen
- The server relay only accepts one phone at a time
- All tool commands are proxied through the relay — the phone is never directly exposed
## Known Limitations (Prototype)
## Known Limitations
### No Encryption
WebSocket connections may use `ws://` (plaintext) instead of `wss://` (TLS). This means:
### Plaintext `ws://` legs are possible
The relay can run without TLS (`hermes relay start --no-ssl`), in which case connections use `ws://` (plaintext) instead of `wss://` (TLS). On a plaintext leg:
- Commands, screen content, and screenshots travel unencrypted
- Anyone on the network path between phone and server can intercept traffic
- **Mitigation**: Use over a trusted network, or set up a reverse proxy with TLS (nginx/caddy)
The clients do not accept this silently: each pairing candidate carries a `transport_hint`, and the Android app keeps plain `ws://` disabled until the operator turns on "Allow plain (unencrypted) connections" and acknowledges the one-time warning dialog. TLS connections get trust-on-first-use SPKI certificate pinning on both the Android app and the desktop CLI.
- **Mitigation**: Use plaintext only on a trusted network, or front the relay with Tailscale (`hermes-relay-tailscale enable`) or a TLS reverse proxy (nginx/caddy)
Hermes API bearer tokens on `/voice/*` are stricter than the legacy WebSocket path: non-loopback API-bearer requests are rejected unless the request is HTTPS, comes through a trusted HTTPS reverse-proxy signal, or the operator has explicitly enabled the insecure dev escape hatch with `hermes relay insecure-api-key on` or the startup env var.
### Full Device Access
Once paired, the agent has unrestricted access to:
### Broad device access once Device Control is enabled
On a `sideload` build with Bridge enabled, the agent can:
- Read all screen content (any app)
- Tap, type, swipe anywhere
- Open any app
- Take screenshots
- Read installed app list
There is no granular permission system — the agent can access banking apps, messages, etc.
- **Mitigation**: Only pair with trusted Hermes instances. Disconnect when not in use.
This access is fenced by several shipped rails rather than a per-app Android permission model:
- **Per-channel grants with TTLs** on the relay session token — bridge access can be excluded or time-boxed at pair time and revoked from any client.
- **Bridge safety rails** (`BridgeSafetyManager`): a per-app blocklist, destructive-verb confirmation prompts (fail-closed on `/call` and `/send_sms`), and an auto-disable timer.
- **Master toggle + status overlay** so control is visible and can be cut instantly on the phone.
### No Command Audit Log
There is no persistent log of what commands the agent executed on the phone.
- **Mitigation**: The relay logs commands to stdout when run with INFO logging.
The rails are deny-lists and confirmations, not sandboxing — a permissive configuration still exposes sensitive apps.
- **Mitigation**: Only pair with trusted Hermes instances, keep the blocklist populated, and disconnect (or let auto-disable fire) when not in use.
### Command audit is local, not centralized
The Bridge screen keeps an on-device activity log of executed device-control commands, and the relay logs commands to its service log at INFO level. There is no centralized, tamper-evident audit trail across surfaces.
- **Mitigation**: Review the Bridge activity log on the phone and the relay service log (`journalctl`) on the host.
### Remote connectivity
Hermes-Relay does **not** ship its own application-layer crypto. The operator owns both endpoints, and the trust model assumes TLS is terminated somewhere on the path that the operator already controls — Tailscale (managed TLS + tailnet ACL identity), a reverse proxy with Let's Encrypt, a WireGuard / other VPN, or a Cloudflare Tunnel. See [`docs/remote-access.md`](remote-access.md) for the decision matrix and setup recipes per mode.
Multi-endpoint pairing (ADR 24) makes "same phone, different networks" a first-class case: a single QR carries `lan` / `tailscale` / `public` candidates in strict-priority order and the phone re-probes reachability on every network change. Per-candidate `transport_hint` drives the plaintext-`ws://` consent dialog — explicit operator consent is still required for any unencrypted leg. The TOFU cert pin is keyed by `host:port`, so two endpoints pointing at the same hostname share a pin (correct — same cert, same pin) while distinct hostnames each get their own.
Multi-endpoint pairing (ADR 24) makes "same phone, different networks" a first-class case: a single QR carries `lan` / `tailscale` / `public` candidates in strict-priority order and the phone re-probes reachability on every network change. Per-candidate `transport_hint` drives the plaintext-`ws://` gating — an unencrypted leg is used only after the operator has enabled and acknowledged plain connections. The TOFU cert pin is keyed by `host:port`, so two endpoints pointing at the same hostname share a pin (correct — same cert, same pin) while distinct hostnames each get their own.
## Recommendations for Production Use
+2 -2
View File
@@ -810,7 +810,7 @@ utilities.
- `GET /voice/config` — provider availability + current settings from `tts:` / `stt:` in `~/.hermes/config.yaml`. When the basic TTS provider is Gemini or xAI, the response includes a `tts.enhanced` capability block (voices/models/audio-tag support + `supports_persona`/`supports_language` flags) so the app renders a per-request enhanced-voice picker. The Vanilla Hermes dashboard `/api/audio/speak` has no per-request surface — enhanced voice there stays config-only via Manage `PUT /api/config`.
- `GET/PATCH /voice/output/config`, `POST /voice/output/session`, and `GET /voice/output/{session_id}` — relay-mediated streaming TTS renderer sessions. Android sends final Hermes text or brokered tool-status text and receives mono PCM deltas for direct `AudioTrack` playback. Session responses include resumable-session metadata and PCM events carry `event_id`/`audio_event_id`, so short route changes during stable speech playback can resume and replay missed audio without re-rendering. Config responses include provider option metadata (`providers[].models`, `providers[].voices`, `providers[].languages`, `providers[].sample_rates`) for first-class dropdowns.
- `GET/PATCH /voice/realtime/config`, `POST /voice/realtime/session`, and `GET /voice/realtime/{session_id}` — relay-mediated realtime provider-agent sessions for lab/dev experiments. Android can send PCM input events and receives mono PCM provider deltas for direct `AudioTrack` playback. Realtime config responses expose the same provider option shape where known.
- `GET/PATCH /voice/realtime-agent/config`, `POST /voice/realtime-agent/session`, and `GET /voice/realtime-agent/{session_id}` — experimental Hermes-brokered Realtime Agent engine. The broker binds active profile/chat session/auth, streams Android mic PCM to a native realtime provider such as `xai_realtime` or `openai_realtime`, normalizes provider transcript/audio/function-call events, mirrors Hermes session/tool/confirmation events into Android, and returns compact Hermes tool results to the provider for concise spoken follow-up. Session responses include resumable-session metadata (`resume_token`, `resume_supported`, `resume_ttl_ms`); server events carry `event_id`, audio deltas carry `audio_event_id`, and Android can resume a detached session through the current `effectiveRelayUrl` after short Wi-Fi/cellular/LAN/Tailscale changes without starting a second Hermes run. The only provider-facing tool surface is `hermes_run_task`, `hermes_get_status`, `hermes_cancel`, and `hermes_confirm`.
- `GET/PATCH /voice/realtime-agent/config`, `POST /voice/realtime-agent/session`, and `GET /voice/realtime-agent/{session_id}` — experimental Hermes-brokered Realtime Agent engine. The broker binds active profile/chat session/auth, streams Android mic PCM to a native realtime provider such as `xai_realtime` or `openai_realtime`, normalizes provider transcript/audio/function-call events, mirrors Hermes session/tool/confirmation events into Android, and returns compact Hermes tool results to the provider for concise spoken follow-up. Session responses include resumable-session metadata (`resume_token`, `resume_supported`, `resume_ttl_ms`); server events carry `event_id`, audio deltas carry `audio_event_id`, and Android can resume a detached session through the current `effectiveRelayUrl` after short Wi-Fi/cellular/LAN/Tailscale changes without starting a second Hermes run. A replacement route is usable only after relay `voice.session.resumed` confirmation; socket generation + resume-episode claims reject stale failure/close/fatal callbacks, unacknowledged input is replayed atomically, and each route-loss episode owns a bounded retry budget that starts at loss rather than session prewarm. Terminal exhaustion detaches session-owned reconnect UI so a stopped retry loop cannot leave an active task pill behind. The only provider-facing tool surface is `hermes_run_task`, `hermes_get_status`, `hermes_cancel`, and `hermes_confirm`.
- `GET /voice/output/providers/{provider_id}/options`, `GET /voice/realtime/providers/{provider_id}/options`, and `GET /voice/realtime-agent/providers/{provider_id}/options` — provider-specific option refresh before saving. Android calls these when a provider is selected so dynamic account-backed choices can be fetched by the relay without exposing provider secrets. xAI refreshes built-in/paginated custom voices when API/OAuth auth is available; ElevenLabs refreshes voices/models/languages with its API key; OpenAI uses static documented voice choices. Realtime Agent provider payloads include `supports_realtime_agent_native` so render/lab-only realtime support is not confused with native speech-to-speech Hermes tooling. Responses include `schema_version`, grouped voice metadata, recommended/custom flags, and model/voice compatibility hints when known. Unknown or unauthenticated discovery falls back to static provider metadata plus manual entry.
- `POST /voice/output/providers/{provider_id}/validate`, `POST /voice/realtime/providers/{provider_id}/validate`, and `POST /voice/realtime-agent/providers/{provider_id}/validate` — pre-save validation for provider/model/voice/sample-rate selections. Unknown manual IDs return warnings; explicit incompatibilities return blocking errors.
- Voice-output provider defaults are relay-owned under `voice_output:` in `~/.hermes-relay/config.yaml` (or `RELAY_VOICE_OUTPUT_CONFIG`), then overridden by `RELAY_VOICE_OUTPUT_*` env vars for temporary tests. Authenticated operator clients may patch safe defaults (`enabled`, `provider`, `model`, `voice`, `sample_rate`, `language`, `codec`, `optimize_streaming_latency`, `text_normalization`, `auto_speech_tags`, `fallback_enabled`) through the relay. With `?profile=<name>`, the patch writes that profile's `voice_output:` section. Provider secrets and local auth paths stay server-side. `auto_speech_tags` is an xAI enhanced-voice control: when the renderer is `xai_tts` the relay applies `upstream_voice.apply_xai_speech_tags()` (upstream's inline/wrapping tone markers) to each chunk before rendering, so the streaming path matches the basic `/voice/synthesize` tone behavior. The `voice_lab` renderer set is xai/openai/elevenlabs — there is no Gemini streaming provider, so Gemini enhanced voice is `/voice/synthesize`-only.
@@ -880,7 +880,7 @@ Current Android dependency versions. Source of truth is `gradle/libs.versions.to
| ML Kit Barcode Scanning | 17.3.0 | QR pairing scan |
| CameraX | 1.6.0 | QR camera preview |
| xterm.js | 5.x | Terminal emulator (WebView) |
| aiohttp | 3.9+ | Server relay |
| aiohttp | 3.14.1+ | Server relay |
| libtmux | 0.37+ | tmux session management |
| gradle-play-publisher | 4.0.0 | Automated Play Console upload (optional) |
+2 -2
View File
@@ -4,8 +4,8 @@ Improvements that would benefit hermes-relay (and other frontends) if added to [
## Current Upstream PR Alignment
- PR #33134 (`feat(api-server): session control API — sessions/chat/fork/SSE-stream`) merged the canonical upstream path for `/api/sessions/*`, message history, fork, chat, and chat stream. It salvaged the useful portion of PR #29302, which superseded the older broad PR #8556. Hermes-Relay should prefer these native routes and keep the bootstrap only as an older-build compatibility overlay.
- PR #33016 (`feat(api-server): add GET /v1/skills and /v1/toolsets`) merged the canonical read-only skill/toolset discovery path. Hermes-Relay should prefer `/v1/skills` and `/v1/toolsets` over legacy `/api/skills` list shapes.
- PR #33134 (`feat(api-server): session control API — sessions/chat/fork/SSE-stream`) merged the canonical upstream path for `/api/sessions/*`, message history, fork, chat, and chat stream. It salvaged the useful portion of PR #29302, which superseded the older broad PR #8556. Hermes-Relay uses these native routes exclusively; the bootstrap's session handlers were retired (2026-07-08) and the bootstrap no longer provides any `/api/sessions` CRUD/messages/fork fallback.
- PR #33016 (`feat(api-server): add GET /v1/skills and /v1/toolsets`) merged the canonical read-only skill/toolset discovery path. Hermes-Relay uses `/v1/skills` and `/v1/toolsets`; the bootstrap's legacy `GET /api/skills` list injection was retired (2026-07-08). Only the legacy detail view (`/api/skills/{name}`) and toggle stub remain bootstrap compatibility routes.
- PR #8199 (`feat(api): add native audio transcription and speech endpoints`) is the canonical upstream path for core STT/TTS execution through `/v1/audio/transcriptions` and `/v1/audio/speech`. Hermes-Relay should keep `/voice/*` as the paired-device facade but eventually call those native core endpoints internally before falling back to private helper imports.
- PR #29364 (`feat: add API server audio endpoints`) should not become a competing `/api/audio/*` API if #8199 remains the accepted audio base. Rework it as a discovery/compatibility follow-up or close it after confirming the upstream maintainer preference.
+8 -7
View File
@@ -1,6 +1,6 @@
# Upstream Hermes Integration Sync
Last reviewed: 2026-06-07
Last reviewed: 2026-07-08 (bootstrap sessions/skills-list retirement)
This document tracks how Hermes-Relay integrates with Hermes upstream surfaces, which
parts use supported extension points, and which parts are compatibility layers that
@@ -40,9 +40,9 @@ relay, dashboard, Android app, desktop app, bootstrap package, or user docs.
| Agent tools | Tool Gateway tools registered through plugin context | `ctx.register_tool(...)` in `plugin/__init__.py`; schemas and handlers in `plugin/tools/*` | Aligned with custom transports | Tool registration should stay in `register(ctx)`; transport details stay behind handlers. |
| Dashboard tab and plugin API | Dashboard plugin manifest plus plugin API routes under the Hermes dashboard plugin mount | `plugin/dashboard/manifest.json`, `plugin/dashboard/plugin_api.py` | Aligned wrapper | Dashboard routes may proxy relay state, but discovery and mounting should stay upstream-native. |
| Chat and model API | OpenAI-compatible API server routes such as `/v1/chat/completions`, `/v1/models`, `/v1/capabilities`, `/health`, and supported streaming routes | Android `HermesApiClient`, relay docs, Web API docs | Mixed | Prefer `/v1/capabilities` when present, then targeted probes for mixed-version fallback. |
| Sessions API | Native API-server session controls merged in NousResearch/hermes-agent PR #33134 (`/api/sessions`, messages, fork, chat, chat stream) | Android `HermesApiClient`; older-build compatibility overlay in `plugin/hermes_relay_bootstrap/*` | Native upstream with fallback | Prefer native `/api/sessions/*`. Bootstrap must skip native routes per method/path and only inject missing compatibility routes for old core builds. |
| Skills and toolsets discovery | Native read-only `/v1/skills` and `/v1/toolsets` merged in NousResearch/hermes-agent PR #33016 | Android `HermesApiClient.getSkills()` prefers `/v1/skills`; desktop/CLI tool surfaces should prefer `/v1/toolsets` where applicable | Native upstream with legacy fallback | Retire `/api/skills` list dependence from clients; keep legacy detail/toggle only where no native equivalent exists. |
| Config, memory, legacy skills, available-models APIs | Not stable current upstream API-server routes as of the 2026-06-16 source check against `55cb4103` | `plugin/hermes_relay_bootstrap/*`, `docs/HERMES-WEBAPI-REFERENCE.md` | Compatibility layer | Keep separate from the sessions/skills retirement path. Do not keep the bootstrap solely for sessions or read-only skill lists once supported baselines include #33134/#33016. |
| Sessions API | Native API-server session controls merged in NousResearch/hermes-agent PR #33134 (`/api/sessions`, messages, fork, chat, chat stream) | Android `HermesApiClient` against native routes only | Native upstream | The bootstrap's sessions CRUD/messages/fork injection is retired (2026-07-08) with no old-build fallback; only `/api/sessions/search` remains bootstrap-provided. Pre-#33134 builds degrade via the client capability probe. |
| Skills and toolsets discovery | Native read-only `/v1/skills` and `/v1/toolsets` merged in NousResearch/hermes-agent PR #33016 | Android `HermesApiClient.getSkills()` prefers `/v1/skills`; desktop/CLI tool surfaces should prefer `/v1/toolsets` where applicable | Native upstream | The bootstrap's legacy `GET /api/skills` list injection is retired (2026-07-08); legacy detail (`/api/skills/{name}`) and the toggle stub remain because no native equivalent exists. |
| Config, memory, legacy skill detail/toggle, available-models, session search APIs | Not stable current upstream API-server routes as of the 2026-06-16 source check against `55cb4103` | `plugin/hermes_relay_bootstrap/*`, `docs/HERMES-WEBAPI-REFERENCE.md` | Compatibility layer | These are now the only routes the bootstrap injects. Each retires individually when core exposes a stable equivalent or the dependent UX is removed. |
| Mobile, desktop, and terminal relay transport | No general upstream plugin WSS transport for persistent remote clients in current public docs | `plugin/relay/server.py`, `plugin/relay/channels/*` | Custom | Keep the relay protocol documented and avoid leaking relay-only assumptions into upstream API clients. |
| Pairing QR and relay session minting | No upstream pairing or device-registration method for remote mobile clients in current public docs | `plugin/pair.py`, relay `/pairing/*`, Android QR parser | Custom | QR payloads should keep API credentials (`key`), dashboard URL (`dashboard_url`), and relay credentials (`relay.code`) as separate fields. |
| Basic STT/TTS over HTTP | Proposed upstream API-server audio endpoints in PR #8199 (`/v1/audio/transcriptions`, `/v1/audio/speech`) | Relay `/voice/config`, `/voice/transcribe`, `/voice/synthesize`; Android `RelayVoiceClient`; `plugin/relay/upstream_voice.py` | Custom wrapper pending upstream replacement | Keep `/voice/*` as the relay auth/session compatibility facade. Once core audio endpoints land, prefer proxying to native `/v1/audio/*` for STT/TTS work before falling back to private helper imports. |
@@ -54,7 +54,7 @@ relay, dashboard, Android app, desktop app, bootstrap package, or user docs.
| Deviation | Owner files | Why it exists | Guard or fallback | Retirement condition |
| --- | --- | --- | --- | --- |
| API bootstrap route and middleware injection | `plugin/hermes_relay_bootstrap/*`; repo-root `hermes_relay_bootstrap/*` is a legacy import shim | Older native installs need session/config/skills/memory endpoints and slash-command preprocessing before upstream exposes stable equivalents. Current upstream already covers sessions plus read-only skills/toolsets. | Method/path feature detection skips native upstream routes and injects only missing compatibility gaps; upstream-module checks skip middleware when native slash preprocessing exists. | Retire per surface: sessions once the supported Hermes baseline includes #33134; read-only skill lists once clients use `/v1/skills`; config/memory/legacy skill detail/toggle/available-models after stable core replacements or local UX removal; slash middleware after native preprocessing exists. |
| API bootstrap route and middleware injection | `plugin/hermes_relay_bootstrap/*`; repo-root `hermes_relay_bootstrap/*` is a legacy import shim | Native installs still lack config, memory, legacy skill detail/toggle, available-models, and session-search endpoints plus slash-command preprocessing. Sessions CRUD/messages/fork and read-only skill lists retired from the bootstrap 2026-07-08 — native upstream (#33134/#33016) owns them with no bootstrap fallback. | Method/path feature detection skips native upstream routes and injects only missing compatibility gaps; upstream-module checks skip middleware when native slash preprocessing exists. | Retire per remaining surface: config/memory/legacy skill detail/toggle/available-models/session search after stable core replacements or local UX removal; slash middleware after native preprocessing exists. |
| Plugin CLI shim fallback | `plugin/__init__.py`, `plugin/cli.py`, install scripts | Current upstream wires third-party plugin CLI commands into the top-level parser, but older supported Hermes builds and scripts may still call the dashed shims. | Prefer `ctx.register_cli_command` / plugin-provided `hermes pair` on current upstream after Hermes-Relay is installed and enabled; standalone shims stay as compatibility wrappers. | Remove shims only after the supported Hermes baseline includes the upstream CLI discovery fix and release/install docs have switched away from the dashed names. |
| Relay HTTP and WSS server | `plugin/relay/server.py`, `plugin/relay/channels/*` | Mobile, desktop, terminal, media, push, and bridge features need persistent client channels and relay-owned session state. | Keep upstream API calls separate from relay session calls and document the protocol in `docs/relay-protocol.md`. | Replace pieces only when upstream provides equivalent remote-client transport or platform adapters. |
| Pairing schema with `relay.code` | `plugin/pair.py`, Android pairing parser, relay `/pairing/*` | An API bearer key authenticates Hermes API calls but does not create relay sessions or describe WSS endpoints. | QR payloads carry direct API credentials, optional dashboard URL, and relay credentials as separate families. | Remove custom pairing when upstream offers native remote-device registration and relay discovery. |
@@ -71,9 +71,10 @@ relay, dashboard, Android app, desktop app, bootstrap package, or user docs.
standard chat/model/health API paths without a fork-only requirement.
- Enhanced management features may require the bootstrap compatibility package only
for surfaces that still lack upstream equivalents. Sessions and read-only skill
lists should be treated as native-upstream-first.
lists are native-upstream-only — the bootstrap no longer injects them at all.
- The bootstrap must compose with partially-upgraded Hermes core builds. Native
routes win per method/path; missing compatibility routes may still be injected.
routes win per method/path; the remaining compatibility routes may still be
injected where missing.
- Relay-specific features must authenticate through relay sessions or explicitly
documented Hermes API bearer checks; do not treat API bearer auth and relay pairing
auth as the same thing.
+2 -2
View File
@@ -22,7 +22,7 @@ Verified upstream source snapshot:
| `/v1/capabilities` | Upstream API server | No | Capability probe | Source of truth for API-server features; current upstream advertises no audio API. |
| `/v1/chat/completions` | Upstream API server | No | Chat fallback | OpenAI-compatible streaming. Tool events may degrade to inline annotations. |
| `/v1/runs`, `/v1/runs/{id}/events` | Upstream API server | No | Chat fallback | Structured run events and stop/approval support. |
| `/api/sessions/*` | Upstream API server | No | Session CRUD and SSE chat | Native upstream session list/create/read/update/delete/messages/fork/chat/chat-stream. Bootstrap is old-build fallback only. |
| `/api/sessions/*` | Upstream API server/dashboard | No | Session CRUD, SSE chat, export, archive, and bulk cleanup | Native upstream session list/create/read/update/delete/messages/fork/chat/chat-stream. Newer hosts also expose single-session JSON export via `GET /api/sessions/{id}/export`, soft archive via `PATCH /api/sessions/{id}`, and guarded bulk cleanup via `POST /api/sessions/prune`; Android must dry-run prune first and show the matched count/span before destructive apply. The bootstrap no longer injects any session CRUD/messages/fork routes (retired in favor of native #33134); only `/api/sessions/search` remains a bootstrap compatibility route. |
| `/v1/skills`, `/v1/toolsets` | Upstream API server | No | Discovery | Read-only API-server skill/toolset inventory. |
| Dashboard `/api/status`, `/api/auth/me` | Upstream dashboard | No | Manage auth | Dashboard cookie/session path; separate from API bearer. |
| Dashboard `/api/auth/ws-ticket`, `/api/ws` | Upstream dashboard/tui_gateway | No | Preferred chat transport | Vanilla Hermes gateway chat path with live reasoning/thinking events. |
@@ -31,7 +31,7 @@ Verified upstream source snapshot:
| `/pairing/*`, `/sessions`, `/voice/*`, `/desktop/*`, `/media/*`, `/notifications/*` on Relay | Hermes-Relay plugin/server | Yes | Relay pairing, terminal, bridge, relay voice, desktop tools | Owned by `plugin/relay/server.py`; Android must gate behind Relay readiness/session grants. |
| Dashboard `/api/plugins/hermes-relay/*` | Hermes-Relay dashboard plugin | Yes for live data | Relay dashboard tab | FastAPI plugin backend proxies loopback requests to the Relay server. |
| `hermes relay doctor` | Hermes-Relay plugin CLI | No for diagnostics | Operator/agent diagnostics | Reports vanilla upstream Hermes route reachability, plugin layout, Relay loopback state, and legacy bootstrap presence. |
| `hermes_relay_bootstrap` routes | Legacy compatibility monkeypatch | No, but non-upstream | Fallback only | Installed via `.pth` by legacy installer. Keep only for older Hermes builds or compatibility-only gaps. |
| `hermes_relay_bootstrap` routes | Legacy compatibility monkeypatch | No, but non-upstream | Fallback only | Installed via `.pth` by legacy installer. Injects only compatibility-only gaps: session search, memory, legacy skill detail/toggle, config, available-models, slash middleware. Sessions CRUD and skill/toolset lists are native upstream and retired from the bootstrap. |
## Client capability gate (build flavor)
+13 -7
View File
@@ -1,23 +1,24 @@
[versions]
appVersionName = "1.2.6"
appVersionCode = "20"
appVersionName = "1.4.0"
appVersionCode = "22"
agp = "9.2.1"
kotlin = "2.4.0"
compose-bom = "2026.06.00"
compose-bom = "2026.06.01"
navigation-compose = "2.9.8"
okhttp = "5.4.0"
kotlinx-serialization = "1.11.0"
kotlinx-coroutines = "1.10.2"
mockk = "1.14.9"
kotlinx-coroutines = "1.11.0"
mockk = "1.14.11"
robolectric = "4.16.1"
konsist = "0.17.3"
security-crypto = "1.1.0"
tink-android = "1.16.0"
lifecycle = "2.11.0"
activity-compose = "1.13.0"
core-ktx = "1.18.0"
core-ktx = "1.19.0"
datastore = "1.2.1"
splashscreen = "1.2.0"
markdown-renderer = "0.42.0"
markdown-renderer = "0.43.0"
coil = "3.5.0"
haze = "1.7.2"
mlkit-barcode = "17.3.0"
@@ -116,6 +117,11 @@ kotlinx-coroutines-android = { group = "org.jetbrains.kotlinx", name = "kotlinx-
# Security
security-crypto = { group = "androidx.security", name = "security-crypto", version.ref = "security-crypto" }
# Pin Tink ahead of what security-crypto pulls transitively: older Tink's
# HybridConfig.<clinit> calls List.removeFirst()/removeLast(), which the Play
# console flags as an Android-15 / SDK-35 crash risk on API < 35. 1.16.0 replaces
# them and keeps the AndroidKeysetManager/AEAD API EncryptedSharedPreferences uses.
tink-android = { group = "com.google.crypto.tink", name = "tink-android", version.ref = "tink-android" }
# DataStore
datastore-preferences = { group = "androidx.datastore", name = "datastore-preferences", version.ref = "datastore" }
+1 -1
View File
@@ -1,6 +1,6 @@
distributionBase=GRADLE_USER_HOME
distributionPath=wrapper/dists
distributionUrl=https\://services.gradle.org/distributions/gradle-9.6.0-bin.zip
distributionUrl=https\://services.gradle.org/distributions/gradle-9.6.1-bin.zip
networkTimeout=10000
retries=0
retryBackOffMs=500
+19
View File
@@ -446,6 +446,25 @@ for stale in "$HERMES_HOME/plugins/hermes-android" "$HERMES_HOME/hermes-agent/pl
fi
done
# Remove any OTHER plugin dir that also declares `name: hermes-relay`. The
# gateway loader dedups discovered plugins by manifest name, so a stale
# duplicate (a backup copy from an older installer, or a leftover native
# install) can win the dedup and make the gateway load stale code — silently
# ignoring every later deploy. Keep only the canonical symlink created above.
plugins_dir="$(dirname "$PLUGIN_LINK")"
canonical_name="$(basename "$PLUGIN_LINK")"
if [ -d "$plugins_dir" ]; then
for entry in "$plugins_dir"/*; do
[ -e "$entry" ] || continue
[ "$(basename "$entry")" = "$canonical_name" ] && continue
if [ -f "$entry/plugin.yaml" ] \
&& grep -Eq '^[[:space:]]*name:[[:space:]]*["'\'']?hermes-relay["'\'']?[[:space:]]*$' "$entry/plugin.yaml"; then
rm -rf "$entry"
ok "Removed duplicate hermes-relay plugin dir: $entry"
fi
done
fi
# Dashboard plugin toggle. The hermes-agent dashboard auto-discovers plugins
# via `dashboard/manifest.json`. We flip visibility by renaming the manifest
# file — no separate config lives anywhere else, and the same state is
+21 -3
View File
@@ -1,4 +1,14 @@
"""Plugin-owned lifecycle for the legacy Hermes-Relay compatibility hook."""
"""Plugin-owned lifecycle for the legacy Hermes-Relay compatibility hook.
The hook (`hermes_relay_bootstrap.pth`) loads the plugin-owned bootstrap
package at interpreter startup. The bootstrap now injects compatibility-only
API surfaces — session search, memory, legacy skill detail/toggle, config,
available-models — plus the slash-command middleware. It no longer provides
sessions CRUD/messages/fork or read-only skill/toolset lists; those are
native upstream (hermes-agent PR #33134 / #33016) and were retired from the
bootstrap. Vanilla Hermes chat, Manage, and dashboard voice never need this
hook.
"""
from __future__ import annotations
@@ -150,8 +160,11 @@ def collect_compat_status(
"standard_path_requires_compat": False,
"plugin_bootstrap_init": str(_plugin_bootstrap_init(plugin_dir)),
"recommendation": (
"Leave compat uninstalled on modern Hermes unless an older server "
"still needs compatibility-only API routes or slash middleware."
"Leave compat uninstalled unless the server still needs the "
"compatibility-only API routes (session search, memory, legacy "
"skill detail/toggle, config, available-models) or slash "
"middleware. Sessions and skill/toolset lists are native "
"upstream and no longer bootstrap-provided."
),
}
@@ -266,6 +279,11 @@ def render_compat_text(data: dict[str, Any]) -> str:
lines.append(f" {item['path']}")
lines.append("")
lines.append("Standard chat, Manage, and dashboard voice do not require compat.")
lines.append(
"Compat covers only session search, memory, legacy skill "
"detail/toggle, config, available-models, and slash middleware; "
"sessions and skill/toolset lists are native upstream."
)
return "\n".join(lines)
+1 -1
View File
@@ -3,7 +3,7 @@
"label": "Relay",
"description": "Paired devices, bridge activity, media inspection, and remote access for hermes-relay",
"icon": "Activity",
"version": "1.3.0",
"version": "1.4.0",
"tab": {
"path": "/relay",
"position": "after:skills"
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "hermes-relay-dashboard",
"version": "1.3.0",
"version": "1.4.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "hermes-relay-dashboard",
"version": "1.3.0",
"version": "1.4.0",
"devDependencies": {
"esbuild": "^0.25.12",
"qrcode": "^1.5.4"
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "hermes-relay-dashboard",
"version": "1.3.0",
"version": "1.4.0",
"private": true,
"description": "Hermes-Relay dashboard plugin frontend (IIFE bundle). Loaded verbatim by the hermes-agent dashboard via the Plugin SDK global.",
"scripts": {
+157 -1
View File
@@ -4,6 +4,11 @@ The doctor command is intentionally local and read-only. It gives operators and
agents one stable surface for checking which parts of the Relay install are
plugin-owned, which standard upstream Hermes routes are reachable, and whether
the legacy bootstrap monkeypatch is still installed.
The bootstrap it reports on is compatibility-only: session search, memory,
legacy skill detail/toggle, config, available-models, and slash middleware.
Sessions CRUD/messages/fork and read-only skill/toolset lists are native
upstream (PR #33134 / #33016) and are no longer bootstrap-provided.
"""
from __future__ import annotations
@@ -86,6 +91,13 @@ def _default_relay_port() -> int:
return DEFAULT_RELAY_PORT
def _default_plugins_dir() -> Path:
"""The Hermes user-plugins directory (`~/.hermes/plugins` or `$HERMES_HOME/plugins`)."""
home = os.environ.get("HERMES_HOME")
base = Path(home) if home else Path.home() / ".hermes"
return base / "plugins"
def _route_exists_status(status: int | None) -> bool:
"""Treat auth and method errors as evidence that a route exists."""
if status is None:
@@ -139,6 +151,58 @@ def http_probe(
}
_MANAGE_SURFACE_FIX = (
"Manage, dashboard voice, and gateway chat need the dashboard/Manage "
"surface. Start it with `hermes dashboard` (default port 9119) and point "
"the dashboard URL at it; `hermes serve` is the headless backend command "
"and is not the Manage surface."
)
def classify_manage_surface(
status_probe: dict[str, Any],
capabilities_probe: dict[str, Any],
) -> dict[str, Any]:
"""Classify what the configured dashboard URL is actually serving.
``status_probe`` is ``GET /api/status`` on the dashboard base (the
dashboard/Manage liveness marker); ``capabilities_probe`` is
``GET /v1/capabilities`` on the same base, which only the API server
answers. The classification catches the two common misconfigurations:
pointing the dashboard slot at the API server, and pointing it at a host
where no dashboard is running at all.
"""
if status_probe.get("ok"):
return {
"kind": "dashboard",
"summary": "dashboard URL answers the dashboard/Manage surface",
"recommendation": None,
}
if capabilities_probe.get("ok"):
return {
"kind": "api-server",
"summary": (
"dashboard URL answers like the Hermes API server, "
"not the dashboard/Manage surface"
),
"recommendation": _MANAGE_SURFACE_FIX,
}
if status_probe.get("status") is None and capabilities_probe.get("status") is None:
return {
"kind": "unreachable",
"summary": "no dashboard/Manage surface reachable at the dashboard URL",
"recommendation": _MANAGE_SURFACE_FIX,
}
return {
"kind": "unknown",
"summary": (
"dashboard URL is reachable but did not answer like the "
"dashboard/Manage surface"
),
"recommendation": _MANAGE_SURFACE_FIX,
}
def _site_dirs(site_dirs: Iterable[Path] | None = None) -> list[Path]:
if site_dirs is not None:
return [Path(p) for p in site_dirs]
@@ -192,6 +256,36 @@ def _plugin_manager_layout(plugin_dir: Path) -> dict[str, Any]:
}
def _duplicate_plugin_dirs(plugins_dir: Path, plugin_name: str) -> list[str]:
"""Names of extra plugin directories that declare the same ``plugin_name``.
The gateway plugin loader dedups discovered plugins by manifest ``name``.
When more than one directory under the plugins dir declares the same name
(e.g. a leftover backup copy from an older installer, or a second native
install), the loader can pick the stale copy and load old code — silently
ignoring every later deploy. Distinct real targets sharing a name are the
hazard; two links to the *same* target are harmless (deduped by real path).
"""
if not plugins_dir.is_dir():
return []
try:
entries = sorted(plugins_dir.iterdir())
except OSError:
return []
by_target: dict[str, str] = {}
for entry in entries:
if _read_simple_manifest(entry / "plugin.yaml").get("name") != plugin_name:
continue
try:
target = str(entry.resolve())
except OSError:
target = str(entry)
by_target.setdefault(target, entry.name)
if len(by_target) <= 1:
return []
return sorted(by_target.values())
def _relay_import_chain() -> dict[str, Any]:
"""Import the relay server module chain under the CURRENT package layout.
@@ -227,6 +321,7 @@ def collect_doctor_report(
timeout: float = 2.0,
probe: Probe = http_probe,
site_dirs: Iterable[Path] | None = None,
plugins_dir: Path | None = None,
) -> dict[str, Any]:
"""Collect a stable, JSON-serializable diagnostic report."""
manifest = _read_simple_manifest(PLUGIN_DIR / "plugin.yaml")
@@ -246,6 +341,22 @@ def collect_doctor_report(
method="POST",
timeout=timeout,
)
# Same base as the dashboard URL on purpose: only the API server answers
# /v1/capabilities, so a hit here means the dashboard slot points at the
# API server instead of the dashboard/Manage surface.
dashboard_capabilities = probe(
_url(dashboard_base, "/v1/capabilities"),
method="GET",
timeout=timeout,
)
# HEAD, never POST: /api/sessions/prune deletes sessions. A POST-only
# FastAPI route answers HEAD with 405 when present and 404 when absent.
dashboard_sessions_prune = probe(
_url(dashboard_base, "/api/sessions/prune"),
method="HEAD",
timeout=timeout,
)
manage_surface = classify_manage_surface(dashboard_status, dashboard_capabilities)
relay_info = probe(
f"http://127.0.0.1:{int(port)}/relay/info",
method="GET",
@@ -253,6 +364,9 @@ def collect_doctor_report(
)
layout = _plugin_manager_layout(PLUGIN_DIR)
plugins_root = plugins_dir if plugins_dir is not None else _default_plugins_dir()
plugin_name = manifest.get("name", PLUGIN_NAME)
duplicate_dirs = _duplicate_plugin_dirs(plugins_root, plugin_name)
bootstrap = _bootstrap_status(site_dirs)
relay_import = _relay_import_chain()
checks: list[dict[str, str]] = []
@@ -269,6 +383,19 @@ def collect_doctor_report(
"ok" if layout["has_dashboard_manifest"] else "warn",
"dashboard manifest is present under the plugin root",
)
_check(
checks,
"plugin-name-unique",
"warn" if duplicate_dirs else "ok",
(
f"multiple plugin directories under {plugins_root} declare name "
f"'{plugin_name}' ({', '.join(duplicate_dirs)}) — the loader dedups by "
"name, so a stale copy can win and the gateway loads old code. Remove "
"the extras and keep only the hermes-relay entry."
)
if duplicate_dirs
else "no duplicate plugin directories share this plugin name",
)
_check(
checks,
"api-capabilities",
@@ -285,6 +412,23 @@ def collect_doctor_report(
if dashboard_status.get("ok")
else "dashboard not reachable from this host",
)
_check(
checks,
"dashboard-manage-surface",
"ok" if manage_surface["kind"] == "dashboard" else "warn",
manage_surface["summary"],
)
_check(
checks,
"dashboard-session-prune",
"ok" if dashboard_sessions_prune.get("exists") else "warn",
"server-backed session cleanup route (/api/sessions/prune) exists"
if dashboard_sessions_prune.get("exists")
else (
"server-backed session cleanup route (/api/sessions/prune) was not "
"detected; bulk cleanup degrades to per-session deletes"
),
)
_check(
checks,
"dashboard-audio",
@@ -328,7 +472,8 @@ def collect_doctor_report(
checks,
"legacy-bootstrap",
"warn" if bootstrap["installed"] else "ok",
"legacy bootstrap monkeypatch is installed"
"legacy bootstrap monkeypatch is installed (compatibility-only "
"surfaces; sessions and skill/toolset lists are native upstream)"
if bootstrap["installed"]
else "legacy bootstrap monkeypatch is not installed",
)
@@ -339,6 +484,8 @@ def collect_doctor_report(
"name": manifest.get("name", PLUGIN_NAME),
"version": manifest.get("version", ""),
"layout": layout,
"plugins_dir": str(plugins_root),
"duplicate_dirs": duplicate_dirs,
},
"standard": {
"api_url": api_base,
@@ -348,6 +495,9 @@ def collect_doctor_report(
"status": dashboard_status,
"audio_transcribe": dashboard_audio,
"ws_ticket": dashboard_ws_ticket,
"capabilities": dashboard_capabilities,
"sessions_prune": dashboard_sessions_prune,
"manage_surface": manage_surface,
},
},
"relay": {
@@ -410,6 +560,12 @@ def render_doctor_text(report: dict[str, Any]) -> str:
status = str(check.get("status", "unknown")).upper()
lines.append(f" [{status}] {check.get('id')}: {check.get('summary')}")
manage_surface = standard.get("dashboard", {}).get("manage_surface", {})
recommendation = manage_surface.get("recommendation")
if recommendation:
lines.append("")
lines.append(f"Fix: {recommendation}")
lines.append("")
lines.append(
"Note: standard chat, Manage, and dashboard voice should work against "
+17 -16
View File
@@ -6,29 +6,30 @@ that waits for `aiohttp.web` to be imported, then replaces `web.Application`
with a thin subclass that detects when hermes-agent's `APIServerAdapter`
attaches itself to a fresh app and:
1. **Injects missing compatibility routes** — older hermes-agent builds may
lack `/api/sessions/*`, `/api/memory`, `/api/skills`, `/api/config`, and
`/api/available-models`. Current upstream already has the session API and
read-only `/v1/skills` + `/v1/toolsets`, so native routes win per method/path
and the bootstrap only fills gaps.
1. **Injects missing compatibility-only routes** — surfaces with no native
upstream API-server replacement yet: `/api/sessions/search`, `/api/memory`,
`/api/skills/{name}` (legacy detail), `PUT /api/skills/toggle` (501 stub),
`/api/config`, and `/api/available-models`. Native routes still win per
method/path if any of these ever land in core.
Retired and no longer injected (native upstream owns them): sessions
CRUD/messages/fork (`/api/sessions/*`, PR #33134) and the legacy read-only
`GET /api/skills` list (native `/v1/skills` + `/v1/toolsets`, PR #33016).
Older pre-#33134 core builds no longer get bootstrap-provided session CRUD;
clients degrade via capability probing to `/v1/chat/completions` / `/v1/runs`.
2. **Installs slash-command middleware** — an aiohttp middleware that intercepts
`/v1/chat/completions` and `/v1/runs` to handle gateway slash commands
(`/help`, `/commands`, `/profile`, `/provider`) and return decline notices
for stateful commands (`/model`, `/new`, `/retry`, etc.), preventing the
LLM from hallucinating responses for them. This mirrors the upstream
Stage 1 preprocessor from `gateway/platforms/api_server_slash.py`.
Stage 1 preprocessor from `gateway/platforms/api_server_slash.py` and
skips itself when that native module exists.
Chat streaming prefers upstream's native
`/api/sessions/{session_id}/chat/stream` endpoint when it is advertised. Older
builds that only get bootstrap-provided session CRUD fall back to standard
`/v1/chat/completions` or `/v1/runs` paths.
This module retires per surface, not as one broad PR cleanup. Sessions can go
once the supported hermes-agent baseline includes PR #33134, read-only skills
should use PR #33016's `/v1/skills`, and the remaining config/memory/legacy
skill/available-model/slash-command surfaces need stable replacements or local
UX removal before the package and `.pth` hook can disappear.
This module keeps retiring per surface, not as one broad cleanup. The
remaining config/memory/legacy skill detail+toggle/available-models/session
search/slash-command surfaces need stable native replacements or deliberate
local UX removal before the package and `.pth` hook can disappear entirely.
"""
from __future__ import annotations
+64 -239
View File
@@ -1,34 +1,31 @@
"""Compatibility handlers from the pre-upstream Hermes-Relay API branch.
"""Compatibility-only handlers for surfaces upstream does not serve natively.
The original broad branch was superseded upstream. Current Hermes main has
native session controls via PR #33134 and read-only skills/toolsets via PR
#33016; these handlers remain for older core builds and for compatibility-only
surfaces that do not yet have stable API-server replacements.
This module began as a mirror of the management endpoints from the pre-upstream
Hermes-Relay fork branch. Upstream hermes-agent has since absorbed the major
surfaces natively, and the bootstrap retires per surface as that happens. The
split as of the sessions/skills retirement (HRUI-002):
This file mirrors the management endpoints from the fork branch, adapted to
take the `APIServerAdapter` instance as an explicit parameter rather than
relying on `self`. That keeps the patch loosely coupled to upstream's class
shape — we don't bind methods onto the adapter, just register closures that
capture an `adapter` reference.
RETIRED — native upstream owns these; the bootstrap no longer injects them,
not even as a fallback for pre-#33134 core builds (older builds degrade to
`/v1/chat/completions` / `/v1/runs` via the client's capability probe):
Endpoints injected (all bearer-auth gated via `adapter._check_auth`):
GET/POST /api/sessions, GET/PATCH/DELETE /api/sessions/{id},
GET /api/sessions/{id}/messages, POST /api/sessions/{id}/fork
— native session control API, PR #33134
POST /api/sessions/{id}/chat + /chat/stream
— native via PR #33134 (never injected here; see note below)
GET /api/skills (legacy read-only list)
— superseded by native `/v1/skills` + `/v1/toolsets`, PR #33016
STILL INJECTED — genuine compatibility gaps with no native API-server
replacement yet (all bearer-auth gated via `adapter._check_auth`):
GET /api/sessions — list sessions
POST /api/sessions — create a new session
GET /api/sessions/search?q=... — full-text message search
GET /api/sessions/{session_id} — fetch one session
GET /api/sessions/{session_id}/messages — fetch session messages
PATCH /api/sessions/{session_id} — rename / update metadata
DELETE /api/sessions/{session_id} — delete a session
POST /api/sessions/{session_id}/fork — clone a session
GET /api/memory — read memory state
POST /api/memory — append memory entry
PATCH /api/memory — replace memory entry
DELETE /api/memory — remove memory entry
GET /api/skills — list skills (optional ?category=)
GET /api/skills/{name} — fetch skill body
GET /api/skills/{name} — fetch skill body (legacy detail)
PUT /api/skills/toggle — STUB (501 Not Implemented).
Registered so the Android client's
capability probe observes the route
@@ -36,42 +33,40 @@ Endpoints injected (all bearer-auth gated via `adapter._check_auth`):
instead of missing. See
``toggle_skill`` docstring for the
upstream gap explanation.
GET /api/config — read model + config
PATCH /api/config — update model/provider/base_url
GET /api/available-models — provider model list
NOT injected:
Handlers take the `APIServerAdapter` instance as an explicit parameter rather
than relying on `self`. That keeps the patch loosely coupled to upstream's
class shape — we don't bind methods onto the adapter, just register closures
that capture an `adapter` reference.
- `POST /api/sessions/{session_id}/chat/stream` — native upstream provides
this in PR #33134. The bootstrap does not inject a chat-stream handler for
older builds because that path requires coordinating with `_create_agent` /
`run_conversation` — the fork's riskiest cross-cutting dependencies. Clients
should fall back to `/v1/chat/completions` or `/v1/runs` when chat streaming
is not advertised.
Notes on surfaces that were never injected:
- `POST /api/sessions/{session_id}/chat/stream` — even before the sessions
retirement, the bootstrap never injected a chat-stream handler because that
path requires coordinating with `_create_agent` / `run_conversation` — the
fork's riskiest cross-cutting dependencies. Clients fall back to
`/v1/chat/completions` or `/v1/runs` when chat streaming is not advertised.
- `GET /api/skills/categories` — removed from upstream as dead code in commit
8d023e43 ("refactor: remove dead code — 1,784 lines across 77 files"). The
app does not call this endpoint; skill browsing uses `/api/skills?category=`.
Re-injecting it would require importing a symbol that no longer exists.
app does not call this endpoint. Re-injecting it would require importing a
symbol that no longer exists.
Removal note: upstream is moving toward focused native surfaces rather than one
large frontend API patch. As each method/path lands in hermes-agent, route
registration below skips that native route and keeps only the missing
compatibility gaps. Cleanup should therefore happen per surface: sessions can
retire once the supported core baseline includes PR #33134, read-only skill
lists should use `/v1/skills` from PR #33016, while config/memory/legacy skill
detail/toggle/available-models remain until core exposes stable equivalents or
Hermes-Relay stops depending on them.
Removal note: registration stays method/path-aware (`_add_route_if_missing`)
so native upstream routes always win if any of the remaining paths ever land
in core. Each remaining surface retires individually when core exposes a
stable equivalent or Hermes-Relay stops depending on it: config, memory,
legacy skill detail/toggle, available-models, and session search.
"""
from __future__ import annotations
import json
import logging
import uuid
from typing import Any, Dict, List, Optional
from typing import Any, Dict
logger = logging.getLogger(__name__)
@@ -96,7 +91,9 @@ def _resolve_upstream():
curated_models_for_provider,
list_available_providers,
)
from tools.skills_tool import skill_view, skills_list
# Only the legacy skill *detail* view remains bootstrap territory; the
# read-only list surface retired in favor of native /v1/skills (#33016).
from tools.skills_tool import skill_view
# MemoryStore lives at tools/memory_tool.py upstream. We import it lazily
# because it pulls in a chain of optional deps that we don't want to crash
@@ -115,7 +112,6 @@ def _resolve_upstream():
"save_config": save_config,
"curated_models_for_provider": curated_models_for_provider,
"list_available_providers": list_available_providers,
"skills_list": skills_list,
"skill_view": skill_view,
}
@@ -169,20 +165,6 @@ def _get_memory_store(adapter, upstream):
# Pure helpers (no adapter coupling)
# ---------------------------------------------------------------------------
def _normalize_session_record(session: Optional[Dict[str, Any]]) -> Optional[Dict[str, Any]]:
"""Parse serialized session fields into API-friendly JSON."""
if session is None:
return None
normalized = dict(session)
model_config = normalized.get("model_config")
if model_config:
try:
normalized["model_config"] = json.loads(model_config)
except (TypeError, json.JSONDecodeError):
pass
return normalized
def _current_model_settings(config: Dict[str, Any]) -> Dict[str, Any]:
"""Extract model/provider/base_url/api_mode from config.yaml."""
model_cfg = config.get("model")
@@ -214,64 +196,16 @@ def _parse_int(value: Any, default: int, minimum: int = 0) -> int:
# ---------------------------------------------------------------------------
# Sessions handlers
# Session search handler
# ---------------------------------------------------------------------------
#
# The only surviving `/api/sessions*` surface. Sessions CRUD, messages, and
# fork retired with native upstream PR #33134; full-text message search has
# no native API-server equivalent, so it stays a compatibility injection.
def _make_sessions_handlers(adapter, upstream):
def _make_session_search_handlers(adapter, upstream):
web = upstream["web"]
async def list_sessions(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
try:
limit = _parse_int(request.query.get("limit"), 50)
offset = _parse_int(request.query.get("offset"), 0)
except ValueError as exc:
return web.json_response({"error": str(exc)}, status=400)
source = (request.query.get("source") or "").strip() or None
db = _get_session_db(adapter, upstream)
items = [
_normalize_session_record(item)
for item in db.list_sessions_rich(source=source, limit=limit, offset=offset)
]
total = db.session_count(source=source)
return web.json_response({"items": items, "total": total})
async def create_session(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
try:
body = await request.json()
except (json.JSONDecodeError, Exception):
return web.json_response({"error": "Invalid JSON in request body"}, status=400)
title = body.get("title")
source = str(body.get("source") or "api_server").strip() or "api_server"
model = body.get("model")
system_prompt = body.get("system_prompt")
session_id = f"sess_{uuid.uuid4().hex}"
db = _get_session_db(adapter, upstream)
try:
db.create_session(
session_id=session_id,
source=source,
model=model,
system_prompt=system_prompt,
)
if title is not None:
db.set_session_title(session_id, str(title))
except ValueError as exc:
return web.json_response({"error": str(exc)}, status=400)
except Exception as exc:
return web.json_response({"error": str(exc)}, status=500)
session = _normalize_session_record(db.get_session(session_id))
return web.json_response({"session": session})
async def search_sessions(request):
auth_err = adapter._check_auth(request)
if auth_err:
@@ -289,114 +223,8 @@ def _make_sessions_handlers(adapter, upstream):
results = db.search_messages(query=query, limit=limit, offset=offset)
return web.json_response({"query": query, "count": len(results), "results": results})
async def get_session(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
session_id = request.match_info["session_id"]
db = _get_session_db(adapter, upstream)
session = _normalize_session_record(db.get_session(session_id))
if session is None:
return web.json_response({"error": "Session not found"}, status=404)
return web.json_response({"session": session})
async def get_session_messages(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
session_id = request.match_info["session_id"]
db = _get_session_db(adapter, upstream)
if db.get_session(session_id) is None:
db.ensure_session(session_id, source="web")
items = db.get_messages(session_id)
return web.json_response({"items": items, "total": len(items)})
async def update_session(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
session_id = request.match_info["session_id"]
db = _get_session_db(adapter, upstream)
if db.get_session(session_id) is None:
return web.json_response({"error": "Session not found"}, status=404)
try:
body = await request.json()
except (json.JSONDecodeError, Exception):
return web.json_response({"error": "Invalid JSON in request body"}, status=400)
try:
if "title" in body:
db.set_session_title(session_id, body.get("title"))
if "system_prompt" in body:
db.update_system_prompt(session_id, body.get("system_prompt"))
if "end_reason" in body:
db.end_session(session_id, str(body.get("end_reason") or "updated"))
except ValueError as exc:
return web.json_response({"error": str(exc)}, status=400)
except Exception as exc:
return web.json_response({"error": str(exc)}, status=500)
session = _normalize_session_record(db.get_session(session_id))
return web.json_response({"session": session})
async def delete_session(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
session_id = request.match_info["session_id"]
db = _get_session_db(adapter, upstream)
deleted = db.delete_session(session_id)
if not deleted:
return web.json_response({"error": "Session not found"}, status=404)
return web.json_response({"ok": True})
async def fork_session(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
session_id = request.match_info["session_id"]
db = _get_session_db(adapter, upstream)
original = db.get_session(session_id)
if original is None:
return web.json_response({"error": "Session not found"}, status=404)
forked_id = f"sess_{uuid.uuid4().hex}"
try:
db.create_session(
session_id=forked_id,
source=original.get("source") or "api_server",
model=original.get("model"),
system_prompt=original.get("system_prompt"),
user_id=original.get("user_id"),
parent_session_id=session_id,
)
for message in db.get_messages(session_id):
db.append_message(
session_id=forked_id,
role=message.get("role"),
content=message.get("content"),
tool_name=message.get("tool_name"),
tool_calls=message.get("tool_calls"),
tool_call_id=message.get("tool_call_id"),
token_count=message.get("token_count"),
finish_reason=message.get("finish_reason"),
reasoning=message.get("reasoning"),
)
except Exception as exc:
return web.json_response({"error": str(exc)}, status=500)
session = _normalize_session_record(db.get_session(forked_id))
return web.json_response({"session": session, "forked_from": session_id})
return {
"list_sessions": list_sessions,
"create_session": create_session,
"search_sessions": search_sessions,
"get_session": get_session,
"get_session_messages": get_session_messages,
"update_session": update_session,
"delete_session": delete_session,
"fork_session": fork_session,
}
@@ -512,21 +340,17 @@ def _make_memory_handlers(adapter, upstream):
# ---------------------------------------------------------------------------
# Skills handlers
# Skills handlers (legacy detail + toggle stub only)
# ---------------------------------------------------------------------------
#
# The legacy read-only list (`GET /api/skills`) retired in favor of native
# `/v1/skills` + `/v1/toolsets` (PR #33016). The per-skill detail view and
# the 501 toggle stub remain: neither has a native API-server equivalent.
def _make_skills_handlers(adapter, upstream):
web = upstream["web"]
skills_list = upstream["skills_list"]
skill_view = upstream["skill_view"]
async def list_skills(request):
auth_err = adapter._check_auth(request)
if auth_err:
return auth_err
category = (request.query.get("category") or "").strip() or None
return web.json_response(json.loads(skills_list(category=category)))
async def view_skill(request):
auth_err = adapter._check_auth(request)
if auth_err:
@@ -577,7 +401,6 @@ def _make_skills_handlers(adapter, upstream):
)
return {
"list_skills": list_skills,
"view_skill": view_skill,
"toggle_skill": toggle_skill,
}
@@ -737,29 +560,29 @@ def register_routes(app, adapter) -> int:
aiohttp keeps mutable until `AppRunner.setup()` freezes it shortly after
`connect()` returns. Native upstream routes win per method/path.
Only compatibility-only surfaces are registered here. Sessions CRUD,
messages, and fork (native via PR #33134) and the legacy read-only skill
list (native `/v1/skills` via PR #33016) are retired — see the module
docstring for the full split.
Returns the number of compatibility routes actually added.
"""
upstream = _resolve_upstream()
sessions = _make_sessions_handlers(adapter, upstream)
search = _make_session_search_handlers(adapter, upstream)
memory = _make_memory_handlers(adapter, upstream)
skills = _make_skills_handlers(adapter, upstream)
config = _make_config_handlers(adapter, upstream)
routes = [
("GET", "/api/sessions", sessions["list_sessions"]),
("POST", "/api/sessions", sessions["create_session"]),
("GET", "/api/sessions/search", sessions["search_sessions"]),
("GET", "/api/sessions/{session_id}", sessions["get_session"]),
("GET", "/api/sessions/{session_id}/messages", sessions["get_session_messages"]),
("PATCH", "/api/sessions/{session_id}", sessions["update_session"]),
("DELETE", "/api/sessions/{session_id}", sessions["delete_session"]),
("POST", "/api/sessions/{session_id}/fork", sessions["fork_session"]),
# Full-text message search — no native API-server equivalent.
("GET", "/api/sessions/search", search["search_sessions"]),
# Memory CRUD — no native API-server equivalent.
("GET", "/api/memory", memory["get_memory"]),
("POST", "/api/memory", memory["add_memory"]),
("PATCH", "/api/memory", memory["replace_memory"]),
("DELETE", "/api/memory", memory["delete_memory"]),
("GET", "/api/skills", skills["list_skills"]),
# Legacy skill detail — native /v1/skills is list-only.
("GET", "/api/skills/{name}", skills["view_skill"]),
]
# Stubbed 501 — see `toggle_skill` docstring. Registered so the
@@ -770,6 +593,8 @@ def register_routes(app, adapter) -> int:
routes.extend(
[
# Model/config + provider model list — dashboard web_server has
# equivalents, but the API server does not.
("GET", "/api/config", config["get_config"]),
("PATCH", "/api/config", config["update_config"]),
("GET", "/api/available-models", config["available_models"]),

Some files were not shown because too many files have changed in this diff Show More