Compare commits

...
67 changed files with 9845 additions and 459 deletions
+11 -2
View File
@@ -138,8 +138,8 @@ jobs:
# The broad Gradle `test` aggregate currently hangs in deferred JVM test
# suites tracked by issue #32. Keep CI release-relevant until that suite is
# split: pairing URL derivation plus connection switching are the stable
# Android regression slice for the active release work.
# split: run the stable connection slice plus focused Chat/Voice state,
# parser, layout, and accessibility regressions for the active release.
- name: Run focused Android unit tests
run: |
./gradlew :app:testSideloadDebugUnitTest \
@@ -149,6 +149,15 @@ jobs:
--tests com.hermesandroid.relay.util.ServerAddressTest \
--tests com.hermesandroid.relay.util.IssueReportAndDiagnosticsTest \
--tests com.hermesandroid.relay.viewmodel.ChatStreamRecoveryTest \
--tests com.hermesandroid.relay.viewmodel.ChatViewModelRealtimeTurnTest \
--tests com.hermesandroid.relay.network.relay.RealtimeVoiceEventParsingTest \
--tests com.hermesandroid.relay.voice.VoiceCommandInterpreterTest \
--tests com.hermesandroid.relay.data.VoiceModePresetTest \
--tests com.hermesandroid.relay.ui.components.BackgroundTaskCardTest \
--tests com.hermesandroid.relay.ui.components.DotMatrixIndicatorTest \
--tests com.hermesandroid.relay.ui.components.AttachmentGalleryLayoutTest \
--tests com.hermesandroid.relay.ui.components.MarkdownStreamingParserTest \
--tests com.hermesandroid.relay.ui.screens.ChatUnreadStateTest \
--console=plain
# Upload reports only for failures. Successful PR report uploads add
+20
View File
@@ -6,6 +6,26 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this
## [Unreleased]
## [1.4.1] - 2026-07-11
### Added
- **Background work is visible in Standard Chat.** A live process strip opens a mobile process sheet with running or recent state, output, elapsed time, Stop, and Dismiss controls. It remains compatible with older Hermes servers that do not expose process details.
- **Background work has a clearer Chat home.** Realtime work appears as a titled task card with working, waiting, delivery, and completion states, queued work, and an expandable tool timeline.
- **Multi-image messages open as galleries.** Adjacent images render in a compact grid and open at the selected image in a swipeable viewer while preserving sensitive-media reveal and original-file actions.
- **Voice gains commands and presets.** Spoken commands can stop speech, cancel background work, pause or resume listening, repeat a result, or start Standard voice chat. Hands-free, Low latency, Careful tools, and Quiet presets tune existing interaction settings.
### Changed
- **Streaming Chat content stays steadier and more readable.** Settled prose and headings adopt final Markdown styling during generation, wide tables scroll with readable columns, the thinking indicator respects system motion and TalkBack settings, and the jump-to-bottom control counts unread messages.
- **Offline Demo mode no longer starts Voice.** The mic action now explains locally that a Hermes connection is required.
### Fixed
- **An in-flight Chat turn survives reopening the app.** Session-backed replies restore partial text, live reasoning, lifecycle status, tool/subagent cards, background-task state, and unanswered approval or clarification cards. Current Hermes gateways reattach to the same running turn; older or finished sessions reconcile from history without duplicating the prompt or losing the final answer.
- **Realtime Agent delivery is protected.** Hermes results use exact provider speech where supported, delivery validation, generation-safe confirmation, and a single relay-TTS fallback if the provider closes or rejects delivery. Voice commands no longer leave synthetic cancellation turns or mute a later background answer.
- **Standard Chat receives background-process completions automatically.** When Hermes completes detached work and starts a follow-up turn on the originating Gateway session, Android shows the unsolicited assistant stream in the open conversation and reconciles history after a cold reconnect. The synthetic process prompt is rendered as a compact process notice rather than a user-authored message.
## [1.4.0] - 2026-07-09
### Added
+177
View File
@@ -1,5 +1,182 @@
# Hermes-Relay — Dev Log
## 2026-07-10 — In-flight Chat turns recover across app recreation
Current upstream Hermes can keep a running Dashboard/TUI Gateway session alive
after its WebSocket transport disappears. `session.activate` rebinds an exact
live session id, while `session.resume` can reuse a live session by its durable
session key and returns `running`, `status`, and an `inflight` snapshot containing
the user prompt plus partial assistant text. Those fields are enough to recover
the live worker and transcript tail, but upstream intentionally does not persist
Android's reasoning presentation, tool-card lifecycle, pending ask card, or
client-owned background-task UI.
Android now checkpoints one session-backed in-flight turn in the shared app
DataStore. The snapshot is scoped by connection/profile and session, expires
after 24 hours, and contains the user/assistant pair, partial answer, reasoning,
tool and subagent states, lifecycle caption, background-task state, and the
server-issued half of an interactive ask. Entered passwords/secrets are never
written. Mutations are debounced during streaming and flushed immediately when
the app backgrounds, when a tool/ask/session boundary changes, and during
orderly ViewModel teardown.
Returning to Chat first restores the rich local snapshot. Gateway sessions then
activate the saved live id before accepting new deltas; if activation is absent
or the live id has expired, Android resumes by durable session id. A running
payload binds the normal event mapper so reasoning, tool, status, ask, and
completion callbacks continue on the same bubble. A settled or unreachable
worker uses the existing bounded, positionally anchored history recovery, which
also covers sessions-SSE transport loss and route handoff. Explicit Stop and
session/profile/connection changes retain their prior interrupt semantics and
clear the checkpoint; lifecycle teardown detaches without sending
`session.interrupt`.
Regression coverage includes checkpoint round-trip/corruption/expiry, rich
ChatHandler rehydration without duplicated repeated prompts, exact activation,
durable-resume fallback, idle-session settlement, detach-without-interrupt, and
a Robolectric reopen that continues reasoning/tool events and clears the saved
turn after authoritative completion.
## 2026-07-10 — Gateway background processes become visible Chat activity
Current upstream Hermes exposes a session-scoped process registry over the same
Dashboard/TUI Gateway socket Android already uses for Chat. `process.list`
returns running and recently finished entries plus a bounded output tail;
`process.kill` stops one process after verifying session ownership;
`agent.terminal.output`, process status events, and terminal/process tool
completion provide refresh and live-output signals. There is no structured
process-start event, so Android follows the official Desktop reconciliation
recipe: load after session prewarm, refresh on relevant events, and poll every
five seconds only while a process remains running. Method-not-found is treated
as an unsupported optional surface instead of a Chat transport failure.
Live phone verification exposed a start-discovery gap: the Gateway emitted the
assistant turn that confirmed a new process ID but no terminal/process
`tool.complete` event, so Android could not begin the running-only poll and first
found the row from a later reconnect/completion snapshot. Every exact-session
`message.complete` now invalidates the process snapshot as a low-cost fallback;
ordinary tool/status events remain the faster path when upstream emits them.
Chat now exposes that state through a compact composer-adjacent background strip
and a current-chat bottom sheet. Running and recent rows show command, elapsed
time, completion/exit state, expandable live or snapshot output, exact-process
Stop, and local Dismiss. Session/client generations reject stale responses after
a chat or connection switch, and profile context is part of the ownership key
because isolated profile databases can reuse stored session IDs. A newer async
prewarm invalidates an older resume before it can replace the live session.
Reconnects repopulate from `process.list`; the five-second safety poll pauses in
the background unless the user explicitly enabled Gateway keep-alive, so it
cannot reopen the socket after the normal background grace close. Raw process
output is length-only in logcat, never placed in notifications, and ANSI/control
sequences are removed before the plain-text mobile viewer renders it.
Hermes intentionally persists a completed process notification as synthetic
user-role input before starting the agent's follow-up turn. Android now recognizes
the upstream formatter shape and presents that history item as a compact,
expandable process notice rather than a human-authored bubble, while preserving
its wire/history role and excluding it from edit-and-resend behavior.
Focused Gateway transport, process-controller, notification-parser, and output
viewer tests pass on both Android product flavors. Google Play and sideload debug
lint report zero errors, and the sideload debug APK assembles successfully for
physical-device validation.
## 2026-07-10 — Gateway background completions return to ordinary Chat
Upstream Hermes already associates a detached process with the originating
Dashboard/TUI Gateway session. When that process completes, its notification
poller injects a synthetic user event, runs a follow-up agent turn, emits the
normal `message.start` / delta / completion lifecycle, and persists the reply.
Android discarded that lifecycle because `GatewayChatClient` only allocated a
turn mapper after a phone-initiated `sendTurn()`; with no request-scoped
`activeTurn`, every event returned before session filtering or UI dispatch.
The Gateway client now accepts a server-initiated turn only when an explicit
event session exactly matches its active live session and the open Chat still
matches the corresponding stored session. It allocates a fresh mapper, binds a
real cancellable turn handle into `ChatViewModel`, and reuses the normal text,
thinking, tool, ask, status, completion, notification, queue, and authoritative
history-reconciliation paths. Foreign or untagged events remain fail-closed.
The mapper also collapses the adjacent duplicate `message.start` pair currently
emitted by the upstream completion poller, preventing a phantom boundary or
duplicate placeholder.
After a user Stop, the client retains a short exact-session drain tombstone for
the interrupted turn. Its late deltas/terminal event are ignored before a
same-session next prompt is submitted, so canceled output cannot reappear as an
unsolicited answer or prematurely complete the newer turn.
A cold foreground prewarm now refreshes the exact resumed session when no turn
is active, recovering a completion that may have finished while the Gateway
socket was closed without overwriting another session or live response.
Regressions cover no-`sendTurn()` delivery, exact-session filtering, duplicate
starts, error recovery, Chat rendering/finalization, Stop-to-interrupt behavior,
late canceled terminals, queued-send draining, and disconnected history recovery.
## 2026-07-09 — Android 1.4.1 Chat and Voice enhancement batch
Chat now represents a promoted realtime background run as one first-class
assistant turn. The same row moves through queued, running, waiting, delivering,
complete, failed, or cancelled state and owns its tool detail and authoritative
answer. Run IDs retain the initiating assistant-row identity across later turns,
including local Voice commands, so delayed progress or delivery cannot settle a
newer placeholder. Local pause, resume, stop, repeat, and cancel commands are
removed from Chat history, quarantine their provider acknowledgement, and cannot
become the target of Retry. The authoritative answer still persists through the
existing session history, and the in-flight Chat checkpoint now preserves the
client-only task-card metadata across a cold restart.
Streaming Markdown can promote blank-terminated prose and headings without
waiting for the final response, while structurally ambiguous lists, quotes,
tables, HTML, and fences remain in the raw tail. GFM tables now wrap in readable
minimum-width columns inside a horizontally scrollable surface. Contiguous image
attachments render as a bounded gallery with selected-page full-screen paging,
sensitive-action gating, original-byte Share/Save behavior, and no adjacent
full-resolution preload. The thinking indicator follows app and system motion
settings plus TalkBack, the jump-to-bottom affordance reports unread messages,
and the Demo mic explains locally that Voice requires a real connection.
Voice now intercepts only exact, final-transcript commands in states where the
action is safe. Standard Voice supports a rearmed new-chat command; realtime
new-chat remains gated until a persistent WebSocket can be rebound safely. Four
presets compose existing Voice settings without replacing manual controls or
silently enabling experimental barge-in. Preset application updates the relay
first and rolls it back if local persistence fails, with an explicit recovery
message if rollback also fails. Relay event parsing accepts the documented and
legacy field aliases used by current broker events.
Foreground Hermes results now use the same forced-summary lifecycle as protected
background delivery. Non-structured verbatim results take the provider's exact
text path where supported; structured results use constrained instructions.
Provider send or response-request failures emit one authoritative fallback before
the terminal error, and each delivery emits one completion boundary. Delivery
confirmation is generation-scoped so an alarm from an older response cannot
invalidate a newer one. Voice-command response suppression is callback-local,
and forced deliveries remain audible after pause, stop, or background-cancel
commands.
Verification passed the focused Google Play and sideload Chat/Voice unit suites,
including a forced clean rerun of the cross-turn ownership regression. Both
Android lint flavors passed. The realtime route, promotion, validation, xAI, and
OpenAI provider slice passed 94/94 tests. Device validation remains for gallery
gestures, reduced-motion/TalkBack behavior, cross-turn background delivery,
command phrasing, all four presets, provider failure fallback, and the existing
route-loss/audio quality release gates.
## 2026-07-09 — Android and plugin 1.4.0 released
`android-v1.4.0` and `plugin-v1.4.0` were published from the same release
commit. The plugin wheel, source archive, and checksum file were downloaded and
verified after publication. The Android release exposes only the intended
sideload APK, Google Play AAB, and matching checksum file; both downloaded
artifacts matched their recorded hashes, and the APK/AAB certificate digests
matched the release signer.
Google Play accepted Android versionCode 22 as a production draft. The draft
was then promoted to `completed`, starting the production rollout. Extended
physical-device recovery stress testing remains deferred and is tracked in
`TODO.md`; live findings may still require follow-up recovery hardening.
## 2026-07-09 — Android realtime turns survive background route loss
An on-device foreground/resume failure left a realtime turn showing
+17 -34
View File
@@ -1,56 +1,39 @@
# Hermes-Relay-Plugin v__VERSION__
**Release Date:** July 9, 2026
**Release Date:** July 11, 2026
**Since v1.3.0:** Realtime Agent background work gains queued long requests, quick side-session answers, deterministic exact xAI delivery, stronger resume ownership, and a delivery-health report. The plugin now installs through upstream Hermes' native plugin path, handles modern virtual-environment layouts, targets multiple Android devices, protects credential paths in media delivery, and adds sharper doctor checks.
**Since v1.4.0:** Realtime Agent result delivery is more dependable when a provider closes, stalls, or overlaps a newer response. Completed Hermes work stays authoritative through provider-native delivery where available and a single relay-TTS fallback otherwise.
Pairs with Hermes-Relay-Android v1.4.0 for the matching background-task, model/voice selection, resume, and task-chip behavior. Standard chat and Vanilla Hermes voice remain upstream-owned and do not require this plugin.
Pairs with Hermes-Relay-Android v1.4.1 for the matching background-task, voice-command, and result-delivery behavior. Standard chat and Vanilla Hermes voice remain upstream-owned and do not require this plugin.
## What's changed
### Added
- **Queued background voice work.** Up to three additional long requests can wait behind an active Hermes task and start automatically in order; cancelling the active run also clears its queue.
- **Quick side-session answers.** A short follow-up can be answered while a background run continues, without disturbing the durable task or its eventual delivery.
- **Provider-native exact xAI delivery.** Exact non-structured results use xAI's forced speech event so the selected realtime voice reads the authoritative Hermes answer without another model inference step.
- **Multi-device Android Bridge.** Multiple Android clients can remain connected and tools can target a named device class, alias, or explicit device ID. `/bridge/devices` and `/bridge/select-active` expose current routing.
- **Delivery health report.** `python -m plugin.relay.realtime_agent.report` summarizes recent realtime-voice delivery modes and fallback reasons.
### Changed
- **Compatibility bootstrap covers only true gaps.** Current Hermes owns native session CRUD/messages and skill discovery; the optional hook now limits itself to legacy surfaces with no upstream replacement.
- **aiohttp 3.14.1 or newer.** Plugin/package requirements move to the patched dependency line covering the 2026 aiohttp security advisories.
- **Long gateway turns use liveness, not a short RPC cap.** Prompt submit can wait up to the server's long-turn ceiling while idle-progress watchdogs determine whether a turn has actually stalled.
- **Provider-native delivery carries an explicit mode.** Realtime responses consistently identify forced-summary and fallback delivery so the Android client can present one authoritative result.
- **Exact delivery is more direct.** Non-structured verbatim results can use provider-native exact text while natural summaries retain delivery guidance.
### Fixed
- **Native `hermes plugins install` compatibility.** Runtime imports are package-relative, dashboard loading works under the upstream plugin namespace, and doctor exercises the real import chain.
- **Modern install layouts.** The installer detects classic, uv-managed, and containerized environments and points generated services/shims at the interpreter it actually found.
- **Doctor catches wrong dashboard surfaces and duplicate plugin copies.** Operators get an actionable correction instead of silently loading a stale directory or pointing Manage at a headless API server.
- **Resume ownership is generation-safe.** A stale candidate cannot detach an active phone route; confirmed replacements reject old failure/close/fatal callbacks, and failed opening candidates are never activated after their terminal callback.
- **Background results survive route loss.** Resumable sessions retain unacknowledged input and replay missed output, retry budgets start when a route is lost, and a detached durable run can still deliver by resume or notification.
- **One handoff and one ready event.** Duplicate spoken background acknowledgements and duplicate fresh-session ready telemetry are suppressed.
- **Credential files cannot be served as media.** Resolved paths under auth, token, pairing, SSH, relay-secret, and system-config locations are blocked even when general media delivery is permissive.
- **A completed result survives provider failure.** If tool-result submission or a follow-up provider response fails, the relay speaks the authoritative Hermes answer through its fallback path before reporting the provider error.
- **Delivery confirmation ignores stale work.** A generation token prevents an older confirmation alarm from emitting a duplicate answer after a newer delivery or preemption.
- **Fallback completion is unambiguous.** The fallback path emits one complete result event even when the provider's audio render cannot finish.
## Install / update
```bash
# Native upstream plugin path:
hermes plugins install Codename-11/hermes-relay/plugin --enable
# Native upstream plugin path:
hermes plugins install Codename-11/hermes-relay/plugin --enable
# Classic install / update on a systemd host:
curl -fsSL https://raw.githubusercontent.com/Codename-11/hermes-relay/main/install.sh | bash
# or, if already installed:
hermes-relay-update
```
# Classic install / update on a systemd host:
curl -fsSL https://raw.githubusercontent.com/Codename-11/hermes-relay/main/install.sh | bash
# or, if already installed:
hermes-relay-update
## Verify
```bash
hermes relay doctor
python scripts/check-plugin-version-sync.py --expect __VERSION__
```
hermes relay doctor
python scripts/check-plugin-version-sync.py --expect __VERSION__
---
Tag prefixes: Android releases use `android-v*`, plugin releases use `plugin-v*`, and CLI releases use `cli-v*`.
Tag prefixes: Android releases use android-v*, plugin releases use plugin-v*, and CLI releases use cli-v*.
+21 -31
View File
@@ -1,56 +1,46 @@
# Hermes-Relay-Android v1.4.0
# Hermes-Relay-Android v1.4.1
**Release Date:** July 9, 2026
**Release Date:** July 11, 2026
**Since v1.3.0:** Realtime voice can keep a long task moving while you ask a quick follow-up, queue another long request, and deliver the finished answer in the selected realtime voice. Recovery is substantially stronger across backgrounding and route changes, model choices apply to the next session, and stale listening, thinking, reconnecting, and cancellation states no longer strand the voice screen. This release also adds model-catalog refresh, proactive notification rules, multi-device Bridge targeting, session-cleanup plumbing, and broad chat, startup, and security fixes.
**Since v1.4.0:** Chat now keeps durable work visible and recoverable. Follow background terminal work from the conversation, receive its completion automatically, and reopen the app into the same in-flight answer with its visible progress intact. Voice adds practical spoken controls and mode presets, while streaming chat gets smoother Markdown, table, and image handling.
v1.4.0 is recommended for everyone. Realtime Agent remains experimental and pairs with relay plugin v1.4.0; the no-plugin Standard chat and Vanilla Hermes voice paths remain upstream-compatible.
v1.4.1 is recommended for everyone. Realtime Agent delivery hardening and voice presets pair with relay plugin v1.4.1; Standard chat and Vanilla Hermes voice remain compatible with unmodified upstream Hermes.
---
## Download
**Installing on your phone?** Download **`hermes-relay-1.4.0-sideload-release.apk`** and tap it — that's the direct-install build with the full feature set (installs as `com.axiomlabs.hermesrelay.sideload`). Prefer the conservative build (no Device Control surface)? Get it from [Google Play](https://play.google.com/store/apps/details?id=com.axiomlabs.hermesrelay).
**Installing on your phone?** Download hermes-relay-1.4.1-sideload-release.apk and tap it — that's the direct-install build with the full feature set (installs as com.axiomlabs.hermesrelay.sideload). Prefer the conservative build (no Device Control surface)? Get it from [Google Play](https://play.google.com/store/apps/details?id=com.axiomlabs.hermesrelay).
The other file, `hermes-relay-1.4.0-googlePlay-release.aab`, is an Android App Bundle for uploading to Play Console — it **cannot** be installed by tapping it on a phone.
The other file, hermes-relay-1.4.1-googlePlay-release.aab, is an Android App Bundle for uploading to Play Console — it **cannot** be installed by tapping it on a phone.
Verify integrity with `SHA256SUMS.txt` from the same release. See the [Sideload guide](https://codename-11.github.io/hermes-relay/guide/getting-started.html#sideload-apk) for APK install steps.
Verify integrity with SHA256SUMS.txt from the same release. See the [Sideload guide](https://codename-11.github.io/hermes-relay/guide/getting-started.html#sideload-apk) for APK install steps.
---
## Highlights
### Realtime voice that finishes the job
### Chat that keeps up
- **Keep talking while work runs.** Quick follow-ups can be answered while one Hermes task runs in the background, and another long request can wait in a bounded queue instead of being discarded.
- **Hear the authoritative answer.** Exact xAI delivery uses provider-native forced speech, finished-task answers can be replayed from the task chip, and TTS/text/notification fallbacks keep a result from disappearing when the realtime floor is unavailable.
- **Stronger route recovery.** Recorded turns wait for relay-confirmed resume, unacknowledged audio is replayed without starting a second Hermes run, and long-lived sessions get a fresh bounded retry window when the route actually drops. Retired sockets and sessions cannot overwrite a newer connection.
- **Clean lifecycle state.** Provider transcripts no longer impersonate active microphone capture; Stop settles local placeholders; exit detaches durable work while clearing session-owned UI; rejected, unacknowledged, or terminal cancels cannot leave an undismissable reconnecting task chip.
- **Your model and voice selection sticks.** Realtime Agent model and voice choices are scoped to the active connection/profile, survive restart, and apply when the next session opens.
- **See background work where it belongs.** Standard Chat surfaces active and recent background processes in a compact strip and expandable sheet with elapsed time, output, a targeted Stop action, and local Dismiss.
- **Get the completion without asking again.** When Hermes finishes detached work, its follow-up answer appears in the originating conversation automatically. The server's internal completion marker stays in history but is shown as a compact process notice.
- **Come back to the same answer.** Closing and reopening the app restores the partial reply, live reasoning, lifecycle status, tool and subagent states, background-task state, and any pending approval or clarification. The app reattaches when the server still has a live turn, otherwise it reconciles the finished transcript without repeating your prompt.
### Chat and model management
### Voice you can direct
- **Long turns stay alive.** Gateway submits use the server's long-turn window and idle-progress checks, avoiding premature transport fallback and duplicate turns.
- **Phone context reaches Hermes.** Voice-intent traces, card actions, and supported attachments now use payload channels the upstream server actually consumes; unsupported attachment paths report the gap instead of dropping it silently.
- **Refresh model catalogs on demand.** Chat and Manage can explicitly reload dynamic/custom provider models, while Manage keeps unconfigured providers visible with key-setup guidance.
- **Session cleanup groundwork.** The dashboard client supports export, prune preview/apply, archive, restore, and archived-session filtering for the Manage surface.
- **Use natural spoken controls.** Pause or resume listening, stop speech, cancel background work, repeat a settled result, or start a new Standard voice chat.
- **Choose an interaction preset.** Hands-free, Low latency, Careful tools, and Quiet presets adjust existing voice and long-task behavior without changing your voice identity or routing.
- **Keep delivered answers authoritative.** Realtime delivery is generation-safe and uses one relay-TTS fallback if the provider cannot deliver a completed Hermes result.
### Phone automation
### Clearer conversations
- **Notification triggers.** Opt-in rules can match app notifications and show a safe local "Ask Hermes?" prompt, with recent activity and a global pause switch.
- **Multi-device Bridge targeting.** Relay tools can select a paired phone, foldable, tablet, or explicit device ID instead of assuming one Android client.
### Reliability and security
- **Older Android crash safety.** Collection calls that require Android 15 were removed from lower-API paths, and encrypted-storage dependencies are pinned to the compatible line.
- **Bad server addresses fail safely.** Malformed relay, media, session, voice, and chat URLs surface a normal connection error instead of closing the app.
- **Credential paths stay private.** Relay media delivery resolves symlinks and blocks credential, token, pairing, SSH, and system-config locations.
- **Cleaner voice failures.** Duplicate error surfaces are gone, fallback speech animates the voice UI, routine provider idle expiry opens fresh on the next turn, and fresh sessions emit one ready event.
- **Browse images together.** Adjacent images form a compact gallery that opens at the image you selected.
- **Read while the reply streams.** Markdown settles into its final styling as text arrives, wide tables stay usable, and motion-sensitive indicators respect system accessibility settings.
---
## Upgrade notes
- App-side release on **both** flavors. Realtime Agent background/recovery features require relay plugin **v1.4.0**; Standard chat and Vanilla Hermes voice continue to work against unmodified upstream Hermes.
- `appVersionCode` is **22**.
- Realtime Agent is still an experimental engine. Stable assistant speech remains available through **Hermes Chat + Voice Output**.
- App version: **1.4.1** (versionCode **23**).
- Realtime Agent improvements pair with relay plugin **1.4.1**.
- Standard Chat and Vanilla Hermes voice continue to work against unmodified upstream Hermes.
+132 -119
View File
@@ -6,22 +6,42 @@ For shipped work, see `DEVLOG.md`. For architectural decisions, see `docs/decisi
---
## Active — next up (2026-07-07)
## Active — 1.4.1 release verification (2026-07-10)
Compaction-safe snapshot of where we are; details in the linked sections below.
Implementation plan: `docs/plans/2026-07-09-1.4.1-chat-voice-enhancements.md`.
Android 1.4.0 / versionCode 22 and plugin 1.4.0 were published on 2026-07-09.
The 1.4.1 Chat and Voice waves are code-complete and merged into local `dev` for
device validation. Version bumps, public release artifacts, push, tags, production
deployment, and store upload remain separate owner-controlled steps.
- **RELEASE IN PROGRESS — cut android-v1.4.0 + plugin-v1.4.0 (owner direction 2026-07-09).** Android is **1.4.0 / versionCode 22**, plugin is **1.4.0**, public release notes and store copy are synchronized, focused realtime recovery tests and Android lint are green, both signed release flavors build, the plugin package builds, and the current sideload APK is installed. Extended on-device recovery stress testing and the force-stop persistence check are explicitly deferred rather than release blockers:
1. Push `dev`, wait for its release-facing CI, merge `dev` -> `main` with a merge commit, then tag the shared merge tip as `plugin-v1.4.0` and `android-v1.4.0`. Plugin is a **MINOR** (it carries #165 native-loader + installer-venv, #170 doctor guardrails, #171 multi-device bridge, #178 dedup guard — not the 1.3.1 patch originally queued).
2. Verify both GitHub releases, their checksums/artifacts, and the Android signing summary.
3. Discard the 1.3.0 Play Console draft, inspect the uploaded 1.4.0 production draft, and start rollout deliberately.
- **Voice bugs being worked now** — see "Voice — on-device findings" below for full detail:
1. Background/resume turn stuck on `Listening...` / `Still working...` — **fixed in code; current APK installed; extended live stress test deferred** (2026-07-09).
2. Tool-call status pills/ordering + stuck "Thinking" — fixed in code; final visual ordering re-check remains.
3. Tap/static click between sentences (`RealtimePcmPlayer` boundary) — still needs an on-device audio repro before fixing.
- **Deferred post-release voice validation.** Repeat long-idle prewarm → record → background/foreground → route-change recovery, terminal retry exhaustion, repeated reopen/exit, cancel-without-ack, and force-stop persistence on physical devices. Capture both Android and relay traces for any recurrence; further recovery hardening or UX refinement may be required from those results.
- **Screen-wake-lock — SHIPPED (2026-07-07).** See "Voice — on-device findings" below.
- **Owner / Mizu — GitHub triage.** Close #64 as superseded, plus the queued open-issue comment/close/label batch (see "Open-issue resolution batch" below).
- **Voice exact-mode signoff — PASSED (2026-07-09 e2e).** Both `grok-voice-latest` and the pinned `grok-voice-think-fast-1.0` deferred on model-generated exact delivery, so xAI exact mode now bypasses inference through provider-native `force_message`. The full on-device background path spoke the authoritative answer through xAI with no fallback, and a pure-recall follow-up repeated it from history without a second Hermes route/run. OpenAI's separate out-of-band delivery spike remains on its next-RC roadmap. Full detail + the background-tasks-as-chat UX asks + remaining audio/UI gaps are below.
Before release preparation, keep these owner/device gates explicit:
- Repeat the exact record → background/route loss → foreground reproduction on the
newly installed debug APK; no `Listening...` / `Still working...` row may strand.
- Recheck long-run tool ordering, the screen wake lock, output waveform timing,
final-syllable tail, and the reported PCM tap/static between sentences.
- Exercise the 1.4.1 Chat surfaces: streaming reflow, wide-table overflow, gallery
paging/zoom/sensitive actions, unread tracking, Demo mic gate, and task-card lifecycle.
- Re-run an ordinary Chat background process through the Gateway: the current-chat
process strip/sheet must show running state, live or snapshot output, exact Stop,
recent completion and Dismiss; the synthetic completion must render as a process
notice, its unsolicited assistant follow-up must appear without another prompt,
and both must survive a socket-close/foreground history refresh without crossing
into a different session or profile. Backgrounding with keep-alive disabled must
also let the Gateway socket close normally instead of polling it back open.
- Start a long Standard Chat turn, wait for visible reasoning plus at least one
running tool card, then background/force-stop/reopen the app. The same session
must restore its partial answer, thinking/status line, tool state, and any live
approval card; new deltas must continue without a duplicate prompt, and a turn
that finished while offline must settle from history instead of staying busy.
- Exercise commands and presets on Standard and Realtime Voice, including ordinary
prompts that resemble commands, explicit stop-vs-cancel behavior, Custom detection,
and preservation of route/provider/model/voice/concurrency/barge-in choices.
- Repeat the Tink encrypted-session smoke: pair → force-stop → relaunch; the session
must persist without an encrypted-preferences startup crash.
- Run release preparation separately: 1.4.1 versioning and public release artifacts,
then owner-controlled `dev` → `main` merge, tag, production deployment, and upload.
- Complete the owner/Mizu GitHub triage batch, including closing #64 as superseded.
---
@@ -60,17 +80,15 @@ Theme: stop treating a background run as an ephemeral voice-only side effect —
surface it in chat like any other turn and keep its result. Overlaps the "Voice
background-run v2" chip roadmap below (items 3/4/7) but reframed around
chat/history rather than the voice chip; unify rather than build twice.
- **Titled background tasks.** Give each run a short title/label (first-line- or
model-derived) so it's identifiable in a list and in chat.
- **Chat entry on kickoff + result.** Drop a chat entry when a background task
starts ("Background task: <title> — running") and settle the result into the same
thread when it finishes. Don't leave it voice-only.
- **Results persisted in chat/history.** Show the background result cleanly in chat
history instead of discarding it after it's spoken — especially valuable for
follow-ups ("what did that say again?").
- **Detail view (expand on tap).** Tapping a background-task chat entry expands to
the full run detail like a normal chat message / tool timeline (reuse
`SubagentLane`). Same intent as v2 items 3+4 — build once.
- **First-class Chat task turn — CODE-COMPLETE for 1.4.1; device verification
remains.** Promotion attaches a short objective title and running state to
the existing assistant row; progress, queued count, waiting/delivery, completion,
failure, cancellation, answer text, and expandable tool detail settle that same
identity. The authoritative answer persists in normal session history. The new
in-flight Chat checkpoint preserves client-only task-card metadata while a turn is
still running across a cold app restart. Metadata for an already-completed task is
still absent from the server history schema after the checkpoint is cleared; keep
that terminal-history case as a separate durability decision.
- **Realtime agent retains background-result context in-session — FALLBACK PATH
DONE + SEEDING LIVE-VERIFIED (2026-07-09); NO-RERUN VERIFY PENDING.** On a FALLBACK delivery the broker now
seeds the delivered answer into the provider's history as an assistant turn
@@ -189,10 +207,15 @@ green. Needs relay deploy + APK install + live verify.
audio, history, and completion lifecycle. Structured results and summary modes
remain model-generated; relay TTS remains the validator fallback. The on-device
background path produced a clean `forced_summary_streaming` event and recall
reused the resulting provider history without another Hermes run. Post-audit hardening remains:
provider-death TTS fallback on all three delivery paths, confirm alarm on all
three, barge-in preemption-as-text, blocklist answer-exemption, and
structured-answer prompt routing.
reused the resulting provider history without another Hermes run.
**1.4.1 post-audit hardening is code-complete:** foreground Hermes results now
enter the same validation/confirmation lifecycle, non-structured Exact delivery
passes authoritative text to provider-native forced speech where supported,
structured answers keep instruction-driven routing, an answer equal to a short
acknowledgement is not falsely blocked, and provider tool-result/response-request
failure emits exactly one authoritative fallback before its terminal error. Live
verify foreground delivery and provider-failure fallback. Barge-in preemption as
durable visible text remains open.
- **Audit leftovers (deliberate, small).** (1) DONE-chip respeak always
renders via relay TTS — intentional determinism, but it voice-mismatches
the exact mode's promise; candidate: provider-voiced respeak with TTS
@@ -209,14 +232,6 @@ wrappers, Android `DiagnosticsLog` Voice category) is in good shape — it
carried every live-round forensics session. Three gaps before the release
candidate:
- **Run-dir retention + wav tap gating — DONE (2026-07-08).**
`run_retention_days` (default 14, 0 disables) sweeps JSONL + wav
artifacts at session-log creation; the render wav is a debug-only tap
(`debug_audio_tap`, default off) deleted after PCM streams.
- **Delivery-outcome rollup — DONE (2026-07-08).**
`python -m plugin.relay.realtime_agent.report [--days N] [--json]`
tallies provider-spoken vs fallback deliveries with reasons; new
`forced_summary_delivered` marker makes clean deliveries countable.
- **Buffered flight-recorder writes (minor).** `_log` open/appends per
event on the event loop, including one line per audio chunk. Fine so
far; switch to a buffered writer if voice sessions ever stutter under
@@ -228,16 +243,13 @@ Full findings with sources in
`docs/plans/2026-07-08-openai-realtime-notes.md`. Headline: the OpenAI
provider already exists and is broker-wired
(`plugin/relay/realtime_agent/providers/openai.py`) but has never had a
live round and defaults to a superseded model. Key provider contrasts vs
recorded live round. The default is already updated to `gpt-realtime-2.1`.
Key provider contrasts vs
xAI: hard 60-min wall-clock session cap (not an inactivity timer),
out-of-band responses (`conversation:"none"` + explicit `input`), async
function calls, per-token pricing (2.1 audio $32/$64 per 1M; mini $10/$20)
vs grok's flat $0.05/min.
- **Bump OpenAI realtime default to `gpt-realtime-2.1` — CODE DONE
(2026-07-08).** Default bumped, `2.1-mini` + rollback `2` in the model
options. Remaining: live connect on 2.1 (covered by the live-verify
item below).
- **Live-verify the OpenAI provider end-to-end.** Code-complete but no
recorded live round (all forensics are grok-voice). Run the xAI
on-device battery (pair → voice turn → `hermes_run_task` →
@@ -263,9 +275,6 @@ vs grok's flat $0.05/min.
instructions, retiring `native_pending_delivery_note`. Success bar:
provider history reads "done" (never "still running") after a promoted
run, verified live.
- **Guardrail test: only `hermes_*` tools advertised on OpenAI
realtime.** Assert `session.update` never advertises hosted-MCP or
non-Hermes tools. Success bar: test fails if any such tool appears.
- **(Defer/eval-only) provider `semantic_vad` vs relay-owned floor.**
Better turn-taking naturalness but moves barge-in ownership off
`RealtimeFloor` — re-architecture, not RC scope.
@@ -277,9 +286,9 @@ tool-calling precision) as the new flagship; `grok-voice-fast-1.0` is
deprecated and the `grok-voice-latest` ALIAS NOW RESOLVES TO THINK-FAST.
We default to the alias everywhere (`config.py:106`,
`providers/xai.py:31`), so the live model may have changed under us —
xAI's docs explicitly say to pin versioned models in production. July also
added 21 multilingual voices, speech tags, voice cloning, session
resumption (30-min inactivity history retention), and a
xAI's docs explicitly say to pin versioned models in production. The current
platform documents five built-in expressive voices, 20+ spoken languages,
speech tags, custom voice IDs, session resumption, and a
`turn_detection.idle_timeout_ms` re-engagement knob.
- **Decide pin-vs-alias, then re-baseline the live delivery rounds.** The
@@ -298,10 +307,12 @@ resumption (30-min inactivity history retention), and a
if resumption is real, the idle-close-and-reseed handling can become
reconnect-and-resume. Success bar: fresh empirical timeout/resume
verdicts recorded in the POC doc.
- **Surface the new voices + speech tags.** `provider_options.py` carries
a static grok voice list; refresh or fetch dynamically, and evaluate
speech tags against the enhanced-voice config contract. Success bar:
new voices selectable in Voice Settings against a live relay.
- **xAI voice catalog + speech-tag UX are code-current; live verify only.** Dynamic
discovery uses xAI's paginated `/tts/voices` surface when auth is available; the
unauthenticated fallback matches the documented built-ins (`eve`, `ara`, `rex`,
`sal`, `leo`; verified 2026-07-09). Voice Settings and Voice Output already expose
the enhanced contract's expressive speech-tag toggle. Exercise both surfaces with
a live xAI relay before release.
## Voice — on-device findings (2026-07-08 e2e realtime test)
@@ -330,10 +341,6 @@ test.**
(payload/metadata). Removed everywhere model-visible (get_status/cancel
default to the active run; the client gets ids via events) + explicit
"never say run/session IDs aloud" in all three instruction sites.
- **Model claimed "I'll add that to the queue" — FIXED (relay, instruction).**
No queue exists (v2 item 2 not built). All handoff/busy instructions now
state "there is no task queue — do not offer to queue or claim to have
queued anything." True multi-task chip stacking remains the v2 queue item.
- **Delivery spoke deferral filler instead of the answer — FIXED (relay).**
The forced-summary validator caught run-id speech (that saved the Minnesota
answer via fallback) but not "One moment while I look that up. I'll report
@@ -395,9 +402,9 @@ cancels). Ranked next increments, in value-per-complexity order:
(at-least-once, unread) — same property as promotion; (c) live verify:
during a long background run, ask a quick second question → answered
inline; ask a second long thing → busy answer unchanged.
2. **Task queue** — upgrade the busy answer from refusal to offer ("want me
to queue it?"): small FIFO in the broker session, start-next-on-completion
with a spoken handoff, chip shows "+1 queued". Pairs with (1).
2. **Task queue — SHIPPED + LIVE-VERIFIED (2026-07-08).** FIFO cap 3,
start-next-on-completion, spoken transition, cancel-clears-queue, and the
`+N queued` chip all landed in the A-E batch above.
3. **Chip tap-through to the transcript** — the run executes on a real
gateway session, so full tool calls/outputs already live in that session's
history; make the chip (or the finished turn) open it. Cheapest "see tool
@@ -410,9 +417,9 @@ cancels). Ranked next increments, in value-per-complexity order:
native async function calling, leave the tool call pending and deliver the
real `function_call_output` late instead of interim-ack + synthetic
instruction text. Needs a live xAI parity check first.
6. **Pending-result FIFO** — `pending_background_result` is a single slot
(correct for one run); generalize to an ordered list the day (1)/(2) land
so two results delivered during a detach don't race.
6. **Pending-result FIFO** — `pending_background_result` is a single slot and
remains correct for the shipped serial queue. Generalize it only with N-way
concurrent background runs so multiple completions can race while detached.
7. **Full N-way concurrent background runs — deliberately deferred.** Needs
session-per-run topology (a gateway session serializes turns), which
fragments conversation context, multiplies delivery/floor/failure modes,
@@ -581,13 +588,13 @@ every bubble) + grouping breaks on a >5min gap (`GROUP_GAP_MS`) so a resumed
conversation gets its own beat; long-press haptic on the action menu; streaming dots
gated to pre-first-token. Deferred:
- **Streaming↔final render parity (kill the reflow).** `StreamingMarkdownContent`
renders raw markdown source (`## `, `**bold**`, `- item`) as plain 14sp text for the
whole turn, then swaps to the full renderer at completion — headings still pop
14sp→20sp on finalize (much reduced now that settled headings are small and lists no
longer resize, but not zero). Run the real renderer on the settled prefix and keep
only the trailing unterminated block raw. Riskier (partial-fence flicker) — needs
on-device testing. Highest-effort audit item.
- **Streaming↔final render parity — conservative 1.4.1 slice implemented; live
reflow check remains.** Blank-terminated, unambiguous top-level prose/headings use
the final Markdown renderer during generation while the active tail stays raw.
Lists, quotes, tables, HTML, and fences intentionally remain lightweight until the
final parse because partial CommonMark containers can re-parent earlier blocks.
Verify that the chosen boundary removes the common heading/prose pop without
introducing partial-fence or list flicker.
- **Bubble body 14sp → 15sp/21.** 14sp is the smallest body of the five reference
apps. Bump markdown paragraph/text/list + the two plain `Text` sites
(`MessageBubble.kt` user/system) together; keep ~1.4 leading so the ~272dp measure
@@ -597,9 +604,6 @@ gated to pre-first-token. Deferred:
(every bubble tails). Switching to iMessage-style "tail on the last bubble only"
changes the look — get design intent before flipping. `isLastInGroup` is now
meaningful (grouping breaks on gaps) so it's ready if wanted.
- **Wide tables.** GFM tables use the default renderer on ~272dp (columns crush);
code fences already horizontal-scroll. Add a custom `table` component in
`markdownComponents` with `horizontalScroll` + ~110dp min column + right-edge fade.
- **Assistant bubble width decoupled from user.** Both cap at 300dp though only the
assistant carries markdown/code; let the assistant run wider (~92% of available /
340–360dp cap) so fences wrap/scroll later. Keep user ~300dp.
@@ -610,8 +614,8 @@ gated to pre-first-token. Deferred:
text selection instead of opening Copy/Quote. Pick one owner (drop
`SelectionContainer`, expose Copy via the menu — chat-app norm — or move actions to a
kebab). Needs on-device confirmation of the current conflict first.
- **Jump-to-bottom FAB unread badge** + drop the no-op tap ripple on bubbles
(`combinedClickable onClick={}` still ripples). Telegram pattern.
- **Drop the no-op tap ripple on bubbles.** The 1.4.1 jump-to-bottom unread badge is
code-complete; `combinedClickable(onClick={})` still ripples on a normal bubble tap.
- **Sessions-transport `animateItem` flash.** Stream-complete rebuilds the list with
new ids → every visible bubble replays its enter animation (gateway transport,
stable id, is unaffected). Reuse the streaming bubble's id for the final message.
@@ -760,8 +764,12 @@ Phase 1 (end-to-end spine) shipped on `Codename-11/phone-platform` — `send_mes
- **LOOK INTO (own item, owner-requested 2026-06-29): live `/api/ws` transport for a foregrounded Thread.** Goal: when a Thread is open in the app foreground, give it the *same* live experience as Chat (live `reasoning.delta` + tool-progress) by running the turn over the `/api/ws` dashboard-gateway transport into that `source=phone` session, instead of the notification-grade `proactive.reply` path. Spec the experiment: (1) does `session.resume` + `prompt.submit` on a `source=phone` session over `/api/ws` keep `source=phone` (not silently re-tag `tui`)? (2) does it bypass `PhoneAdapter` / the role_authorized reply loop, and does that matter when the user is the one typing? (3) reconcile the two send paths (foreground→`/api/ws`, background/notification→`proactive.reply`) without double-sends. If it holds, a Thread becomes "background-delivered like a DM, but live like Chat when you open it" — the best of both. Until verified, `proactive.reply` stays the only send path.
- **Docs/user-docs for Threads (lockstep — author with the user-facing slices 4–5).** Dev refs are done (ADR 12 carries the unified-session decision + the two-"gateway" split). Still to write when the surface ships: a plain-language `user-docs/features/threads.md` — what a Thread *is*, **Chat vs Threads** (live foreground work vs. persistent, agent-reachable conversations), the two opt-in gates, that it's relay-only — plus a **brief in-app explainer** (e.g. a one-line hint on the Threads filter empty state or a small info affordance, not a wall of text), and `docs/relay-protocol.md` + relay-server route docs for the wire. Replace the stale user-docs "Coming Soon → Push Notifications" row; keep it distinct from the clipboard inbox and the inbound Notification Companion.
- **More Threads fold-ins (capture now, build with the relevant slice).** (a) **Read-state back to the agent** — tell the gateway you saw a proactive message (Discord-style read receipt) so the agent knows; fold into the `proactive.reply.ack` design (#7). (b) **Cross-surface reply** — because a Thread is just a gateway session, a reply could come from the desktop CLI / dashboard too, not only the phone; near-free once unified, verify the reply routing. (c) **Priority/importance on a proactive message** — let the agent mark urgent vs FYI → notification importance / quiet-hours bypass; small payload field + maps to the notifier channel.
- **Per-thread `chat_id`.** Everything is hardcoded `chat_id="phone"` (one thread) today; the adapter already plumbs `chat_id`, so varying it yields multiple threads (per topic, or the agent opening distinct conversations). Ties into the threaded surface.
- **Message status + delivery state.** Surface sent / delivered / queued / failed per message in the thread (depends on outbound buffering's queued state) so the user knows whether the agent actually reached them.
- **Agent-created per-thread `chat_id`.** User-created named Threads and arbitrary
`chat_id` routing are shipped. Remaining: expose a `send_message`-adjacent
agent affordance that can deliberately open/name a project Thread.
- **Queued message state.** Sending/Delivered/Failed bubbles and relay reply ACKs
are shipped. Add an honest Queued state plus Cancel when the offline outbox
exists; do not infer delivery from socket enqueue alone.
- **Auto-title the phone thread** like other sessions (first confirm whether the gateway already auto-titles platform sessions; wire it through if so).
### Discord/Telegram replacement — capability gaps (to fully retire reaching for them)
@@ -769,15 +777,27 @@ The gateway-platform model is the *correct + sufficient architecture* (the phone
- **Guaranteed background delivery (the biggest gap; no push today).** Delivery is **live-WSS-only** + a 24 h relay buffer; there is **no FCM/UnifiedPush** wake-up. If the app process is dead AND not holding a socket, a message waits for the next reconnect, and the relay buffer is ephemeral (lost on relay restart). Discord/Telegram feel instant because they wake the device via push even when the app is dead. Decide a **push transport**: **UnifiedPush/ntfy** (recommended — self-hostable, no Google dependency, upstream *already* ships an `ntfy` platform, on-brand for self-hosted) vs **FCM** (simplest UX but adds Play Services + a push relay; clashes with self-hosted ethos — at most the `googlePlay` flavor) vs **persistent foreground keep-alive service** holding the relay WSS (zero new infra, like `GatewayKeepAliveService`, but battery cost + Doze-fragile). Likely: UnifiedPush primary + foreground-keepalive fallback.
- **Cron / background-job delivery is BROKEN** (already tracked above): `deliver=phone` standalone path → `Unknown platform: phone`. This is load-bearing for "receiver of crons/background jobs" — fix is required, not optional, for the replacement goal.
- **Multi-thread is wired-for but never varied** (already tracked: per-thread `chat_id`). For real DM/channel parity the agent must *open distinct threads* (vary `chat_id` per topic/job), the app must render a **thread list** (N conversations, not one), and replies route back by `chat_id`+`reply_to` (already plumbed).
- **Durable history / scrollback.** The relay buffer is ephemeral; a real messaging surface needs persisted scrollback. Read the gateway **session store** for the `phone` platform's history (relay-exposed read path) so reopening a thread shows the full conversation, not just buffered-while-away.
- **Agent-initiated multi-thread creation remains.** The app already renders N
`source=phone` sessions, user-created Threads vary `chat_id`, and replies route
by `chat_id` + `reply_to`. The missing parity is letting the agent open/name a
distinct Thread for a topic or job.
- **Durable history / scrollback — SHIPPED.** Threads reopen through the gateway
session store; the relay buffer is only the live/offline-delivery layer, not a
parallel history database.
- **Profile = contact mapping (new idea, fold in).** Multiple Hermes **profiles** (distinct agent personas/configs) could each be a distinct thread *source*/"contact" — DMing different agents. Maps cleanly onto the per-thread `chat_id` + source-attribution work; lets the app feel like a contact list of agents.
- **Per-thread notification controls + deep-link (Discord-parity affordances).** Per-thread notification channels, mute/DND/quiet-hours (Phase 3 partially), and a notification that **deep-links into the exact thread** (tap → land in that conversation) so dipping in/out while multitasking is frictionless.
- **Agent-initiated rich content.** Agent → phone thread with **images/cards** (relay media infra + `InboundAttachmentCard`/`HermesCardBubble` already exist on the chat side — reuse). Inbound (phone → agent) reply media stays deferred (text-first), but outbound rich content is low-cost parity.
- **In-thread "agent is working" indicator.** A typing/working state in the thread while the agent thinks/runs tools (Discord typing-dots parity) — the chat surface already has thinking indicators to reuse.
- **Source/platform attribution + filtering in the drawer (NOW READY — owner-requested 2026-06-29; the gateway/Threads surface has shipped).** `/api/sessions` DOES expose `source` (confirmed live: `tui`, `cli`, `api_server`, `web`, `discord`, `telegram`, `cron`, `webhook`, `phone`). Build: **(a)** a clean **source badge** per session in the drawer — phone → the thread-spool (done); discord / telegram / cron / webhook / web → a small per-platform chip/icon (match hermes-desktop's convention); the app's own `tui`/`api_server` chats get no badge (or a subtle one). **(b)** a **filter** (drawer dropdown) to show/hide sources. **(c)** a **setting** (Chat settings) for the default — **hide the agent's other-gateway/automation sessions (cron / webhook / discord / telegram) by default** so the drawer shows just your chats + Threads, with a toggle to reveal them (the live default `state.db` is full of cron/discord/webhook noise). Persist the visibility prefs. Can't see the official desktop (no clone) — infer its chip styling; match exactly if specifics surface. Standard-path: read-only display of the upstream `source` field. Fold cross-restart **Thread-name persistence** (currently in-memory) into this drawer pass.
- **Beta-gate the Threads featureset (owner direction 2026-06-29).** Mark Threads **Beta** with a clean badge in the UI (the Threads filter chip + the best-path "Threads" capability row) until the enhancements land. Full (non-beta) release is gated on: **live `/api/ws` transport for a foregrounded Thread** (an open Thread streams like Chat — the headline), per-session **unread**, the **`chat_id`-on-`/api/sessions` upstream fix** (so threads route after restart / cross-device), and **outbox/retry**.
- **Source/platform attribution, filtering, and Thread-name persistence — SHIPPED.**
The drawer and Chat settings show source badges and persisted visibility filters;
`ThreadNameStore` persists user Thread names across restart and reapplies them to
session rows. Remaining Threads work is the explicit residual list above
(unread, outbox/retry, exact deep-link, agent-created named Threads, and live
foreground `/api/ws`).
- **Threads Beta badges — SHIPPED.** The Threads filter and best-path capability
row render the shared `BetaChip`. Removing Beta remains gated on live foreground
`/api/ws`, per-session unread, upstream `chat_id` exposure, and outbox/retry.
## Voice — Standard-path parity follow-ups
@@ -829,18 +849,19 @@ The client-side mitigations shipped (see DEVLOG 2026-06-27): the `updateSessions
### Thinking indicator — post-v1.3.0 follow-ups
The animated dot-matrix "thinking" indicator shipped in **android-v1.3.0** (Wave/Pulse/Bounce/Sparkle motions + Auto/accent colors, live preview in Chat settings; static when animations are off). Remaining:
The animated dot-matrix "thinking" indicator shipped in **android-v1.3.0**
(Wave/Pulse/Bounce/Sparkle motions + Auto/accent colors, live preview in Chat
settings). The 1.4.1 path also honors app animation settings, OS animator scale,
and TalkBack touch exploration. Remaining:
- **OS-level reduce-motion / TalkBack** — currently gates only on the app's `animationEnabled` pref. Also honor OS reduce-motion + touch-exploration like `CleanChatMode` does (`rememberCleanMotionState().osAnimations`).
- **Optional: promote to a full avatar style** — the alternative scope (a `DotMatrixAvatar` `AgentAvatar` shown everywhere via `LocalAvailableAvatars`, selected in Appearance). Deferred in favor of the narrower in-bubble indicator.
## Demo mode (2026-06-27) — deferred polish
Shipped offline Demo / Explore mode (see DEVLOG 2026-06-27). Core is in; these are non-blocking polish items, none required for the Play "App access" fix:
- **On-device verify (Studio).** Confirm: "Try the demo" on the onboarding Connect page and the standalone Connect screen lands on Chat showing the canned transcript (Markdown, tool-progress card, weather card, code block); the persistent banner shows and its Connect exits demo into the real wizard; demo runs in airplane mode with no network; Manage/Voice show the demo empty state; Bridge/Terminal show their pair-gate; backing out of demo Chat clears the flag so a real connection still works.
- **On-device verify (Studio).** Confirm: "Try the demo" on the onboarding Connect page and the standalone Connect screen lands on Chat showing the canned transcript (Markdown, tool-progress card, weather card, code block); the persistent banner shows and its Connect exits demo into the real wizard; demo runs in airplane mode with no network; the Chat mic explains locally that Voice needs a connection and never attempts transcription; Manage/Voice show the demo empty state; Bridge/Terminal show their pair-gate; backing out of demo Chat clears the flag so a real connection still works.
- **Demo composer is a silent no-op — DONE 2026-07-08.** `sendMessage` now intercepts while `isDemoMode`: echoes the user bubble and appends `DemoContent.composerReply` ("offline demo, can't answer for real — tap Connect in the banner"), both clientOnly so demo-exit's `clearMessages()` wipes them. Wired via `setDemoModeWiring` (unconditional in RelayApp — the client-gated chat init never runs in demo, so ChatViewModel's own handler is null there). On-device check rides the existing demo verify item above.
- **Live voice mode in demo.** The voice-mode overlay (mic) launched from Chat isn't demo-gated — a tap would attempt a transcribe (fails gracefully, no crash). Add a demo notice / disable the mic in demo. (Voice settings screen already shows the demo empty state.)
- **Light typewriter/stream simulation.** The transcript is statically populated; an optional per-token reveal on first entry would better convey the "streaming" feel. Acceptable as static for v1.
- **Optional richer demo.** Could add a second tool type or an image attachment to the transcript to showcase more surfaces; kept minimal/one-file for now.
@@ -863,7 +884,6 @@ Client-side profile-lock + voice fixes (the items marked above) landed via a pla
- **Per-profile voice on Standard (upstream).** `/api/audio/*` is host-global/text-only; the Standard surface still can't carry a per-request voice. Needs the upstream profile-voice / `/v1/audio/*` PR. Until then the client prefers the relay path; consider surfacing an honest "override needs Relay" state when Standard is the effective surface.
- **Profile lock: ChatScreen glyph + export.** The optional lock glyph on the chat-header avatar was skipped (`ChatScreen.kt` is owned by a concurrent session). Decide whether the per-connection lock belongs in settings export/import (it rides the `profile_selections` DataStore).
- **Unit tests — DONE 2026-06-21 (36/36 pass via `:app:testSideloadDebugUnitTest`).** `ProfileLockStoreTest` (9 — uses an in-memory `DataStore` harness; the file-backed factory hits a Windows write-rename/instance race), `ProfileControllerLockTest` (8, Robolectric), `CoerceAudioRouteTest` (7), `VoiceStatusGatesTest` (12).
- **CHANGELOG.** Add `[Unreleased]` entries (Profile lock → Added; voice override + realtime → Fixed) at build-verify/PR time.
- **On-device verification.** Override applies in 'auto'+relay; realtime survives a &gt;90s background task without stalling and stops over-narrating; Speaking waveform unfolds at first audible frame; profile lock hides pickers + holds on a missing profile; overlay shows the profile icon.
## Hands-free agentic voice backlog
@@ -872,33 +892,25 @@ Goal: make Hermes usable for hands-free work without leaving the operator blind
to tool state, safety prompts, or the current task.
- **Waveform output-start sync** — current input waveform timing feels good, but
- **Waveform output-start sync — SHIPPED; on-device confirmation remains.**
Realtime output now gates on `RealtimePcmPlayer` playback-head movement or
playback-synchronized amplitude through `shouldMarkRealtimeOutputActive`,
matching the basic-TTS path. Confirm visually on-device with the 1.4.1 batch.
the agent-output waveform can unfold and begin movement before audible speech
- **Voice command layer — initial 1.4.1 subset code-complete; live verify and
navigation residuals remain.** Exact final transcripts can stop speech,
explicitly cancel the active background task, pause/resume Continuous mode,
repeat a settled background answer, and start a new Standard chat. Bare `stop`
and `cancel`, partial transcripts, and command-like ordinary prompts stay on the
normal Hermes route. Realtime `new chat` remains gated on a clean websocket
session-rebind boundary; `open overlay` and `return to Hermes` remain future
navigation commands. Verify barge-in Stop, pause during a background run, local
command Chat cleanup, and Continuous rearm on device.
starts. Split "preparing audio" from "speaking audio" in the visual layer, or
gate the unfolded Speaking waveform on the first real playback frame/audio
amplitude. Processing can stay as the folded circular spinner until output is
actually audible.
- **Voice command layer** — reserve local commands that bypass normal agent
routing: "pause", "resume", "stop talking", "cancel", "repeat that", "open
overlay", "return to Hermes", and "new chat". These should work while the
agent is thinking, speaking, or using tools.
- **Spoken tool progress** — when Hermes uses tools, voice mode should speak
short status updates such as "I'm checking the relay logs" or "I found an
error" without waiting for final assistant text. Long tool calls should emit
periodic, low-noise progress updates.
- **Spoken tool progress — baseline shipped; broader hands-free policy remains.**
Realtime background runs already emit milestone speech plus coarse, low-noise
progress with repeat suppression. The 1.4.1 residual is a unified policy across
Voice engines and presets, not another parallel heartbeat implementation.
- **Realtime tool timeline parity** — the voice overlay should render the same
@@ -918,11 +930,12 @@ the current voice task: active objective, last tool result, pending next step,
and whether the agent is waiting on the user.
- **Mode presets** — add presets such as Hands-free, Low latency, Careful tool
mode, and Quiet/visual-only. Hands-free should favor Continuous listening,
spoken tool progress, confirmations, and overlay availability.
- **Mode presets — CODE-COMPLETE for 1.4.1; live apply/Custom-state verification
remains.** Hands-free, Low latency, Careful tools, and Quiet/visual-only compose
existing interaction and relay-promotion controls. They preserve engine, route,
provider, model, voice, credentials, concurrency, and Hands-free's existing
experimental barge-in choice. Relay update is server-first; local Voice/barge-in
values share one DataStore transaction, with relay rollback on local failure.
- **Barge-in hardening** — keep barge-in experimental until echo/self-recording
@@ -1041,7 +1054,6 @@ Follow-ups:
## Attachments (shipped 2026-06-18 — `docs/plans/2026-06-18-attachment-experience.md`)
- **B3 — download progress + cancel.** Inbound fetch is un-cancelable; the previews work scaffolded an indeterminate bar + nullable `onCancel`. Live wiring needs the fetch-path owner (`ChatViewModel`/`Attachment`) to expose determinate progress (Content-Length) + a cancel hook.
- **A6 — multi-image gallery.** N images in one message → grid + swipe-across viewer (Telegram media-group parity).
- **C5 — agent-side sensitivity config gate.** `RELAY_MEDIA_SENSITIVITY_HINTS` (env or per-profile) instructing the agent to annotate sensitive media via the prompt-builder. Transport (relay `X-Media-Sensitive` header + client blur) already ships; the agent isn't asked to set the bit yet.
- **Relay thumbnails (D6).** Server-side thumbnail generation to avoid full-size download for cards/galleries. Needs an image lib (Pillow not currently a dep) — evaluate before adding.
- **D5 — outbound upload progress.** No per-attachment progress during the 60s gateway PDF-render window.
@@ -1049,12 +1061,13 @@ Follow-ups:
## Voice overhaul (shipped 2026-06-18 — `docs/plans/2026-06-18-voice-overhaul.md`)
- **Per-profile voice on Standard (upstream PR).** Upstream `/api/profiles/*` has no voice field and `/api/audio/*` is host-global. Long-term: PR a voice section to the profile config + make `/api/audio/*` honor the active/`?profile=` profile. The relay path already carries per-profile voice; ship that first.
- ~~**Wire connectionId for per-profile voice namespacing.**~~ **Already shipped — stale entry (verified 2026-07-08).** The wiring landed in `0aa1b38` (2026-06-21, the same batch this list belongs to): `RelayApp` has a `LaunchedEffect(activeConnectionId, selectedProfile?.name)` calling `voiceViewModel.setVoicePrefsConnection(activeConnectionId)` *before* `onProfileChanged(...)`, and `applyVoicePrefsScope` pushes `(connectionId, profile)` into `VoicePreferencesRepository.setActiveScope`. Two connections with same-named profiles namespace separately.
- **Realtime-PCM waveform output gating.** The basic-TTS output waveform is now Visualizer-accurate (gated on real playback amplitude), but the realtime path gates `outputAudioActive` on `audioSeen` (first decoded PCM bytes) in `VoiceViewModel.handleRealtimeVoiceEvent`, which can still lead audible output by the `RealtimePcmPlayer` start prebuffer. Gate realtime on actual playback-start (head moved) to match the basic-TTS path.
## Chat clean-mode + pets (shipped 2026-06-18 — `docs/plans/2026-06-18-chat-clean-mode-and-pets.md`)
- **Part-A chat polish (optional bundle).** Per-code-block copy + horizontal scroll, visible copy affordance, mid-stream stall feedback, profile/skill-aware empty-state chips, the ~40-flow recomposition hotspot at the top of `ChatScreen`. (Sphere `contentDescription`/reduced-motion was handled by the clean-mode a11y work.)
- **Part-A chat polish residuals.** Per-code-block copy, horizontal scroll, the
visible copy affordance, and mid-stream stall feedback are shipped. Remaining:
profile/skill-aware empty-state chips and the ~40-flow recomposition hotspot at
the top of `ChatScreen`.
- **Pet hot-load + in-app add/remove (shipped 2026-06-20).** Pets now live-refresh: an `avatarsRefreshTick` keys the avatar `produceState` in `RelayApp`, and Appearance re-scans `pets/` on open and after in-app import/delete — no app restart. Appearance gained "Add a pet" (SAF `.zip` import via `PetImporter`, zip-slip/zip-bomb guarded + validated through `toAvatar`) and an "Installed pets" list with per-pet remove (`PetLoader.deletePet`, confirm dialog, Sphere fallback). Remaining:
- **Sphere-skin parity.** Skins are still process-scoped + `adb push` only — the live tick and the importer cover pets, not skins. Extend the tick to `loadUserSkins` and add a `.json` skin import if hot-loading/adding skins in-app is wanted.
- `**adb push` into `Android/data` hangs on Samsung scoped storage.** Confirmed: pushing a pet pack to `/sdcard/Android/data/<pkg>/files/pets/` stalls (no bytes written) although `adb shell ls` of the dir works. In-app `.zip` import is the supported path; `/sdcard/Download` pushes fine. Consider softening `docs/pet-spec.md` + user-docs to lead with in-app import over adb.
@@ -1,6 +1,12 @@
v1.4.0 — Realtime voice that finishes the job.
v1.4.1 - Chat that keeps up
• Long voice tasks can queue, keep running while you ask quick follow-ups, and deliver answers in the selected realtime voice.
• Voice sessions recover more reliably after background or route changes and clear stale task states.
• Refresh model catalogs on demand; add opt-in notification rules and multi-device Bridge targeting.
• Safer startup, server-address handling, long chat turns, and credential media access.
Chat
* Follow background work from a live process strip; its result appears automatically in the same conversation.
* Reopen while an answer runs: partial text, thinking, tool progress, and approvals return.
Voice
* Speak commands to pause, resume, cancel, repeat a result, or start Standard voice chat.
* Pick Hands-free, Low latency, Careful tools, or Quiet presets.
Polish
* Multi-image galleries plus smoother streaming Markdown and long tables.
+27
View File
@@ -1,5 +1,32 @@
{
"versions": [
{
"version": "1.4.1",
"title": "Chat that keeps up",
"date": "2026-07-11",
"sections": [
{
"header": "Chat that stays with you",
"bullets": [
"Follow background terminal work from a compact process strip and expandable sheet. Its completed answer appears in the same conversation automatically.",
"Close and reopen while a reply runs: partial text, thinking, tool progress, background-task state, and pending approvals return in the same chat without repeating your prompt."
]
},
{
"header": "Voice you can direct",
"bullets": [
"Use spoken commands to pause or resume listening, stop speech, cancel background work, repeat a finished result, or start Standard voice chat.",
"Hands-free, Low latency, Careful tools, and Quiet presets tune existing voice behavior without changing your selected voice or route."
]
},
{
"header": "Clearer conversations",
"bullets": [
"Browse adjacent images as a gallery, read smoother streaming Markdown and wide tables, and see background-process completion as a compact process notice."
]
}
]
},
{
"version": "1.4.0",
"title": "Realtime voice that finishes the job",
+9 -19
View File
@@ -1,22 +1,12 @@
v1.4.0 - Realtime voice that finishes the job
v1.4.1 - Chat that keeps up
Chat
* Follow background work from a live process strip; its result appears automatically in the same conversation.
* Reopen while an answer runs: partial text, thinking, tool progress, and approvals return.
Voice
* Keep talking while long work runs: quick follow-ups can be
answered, another long request can queue, and the finished
answer stays in your selected realtime voice.
* Background and route changes recover more reliably. Stale
listening, thinking, reconnecting, and cancel states clear
instead of trapping the voice screen.
* Pick a Realtime Agent model and voice per connection/profile;
the next session uses it and the choice survives restart.
* Speak commands to pause, resume, cancel, repeat a result, or start Standard voice chat.
* Pick Hands-free, Low latency, Careful tools, or Quiet presets.
More control
* Refresh provider model catalogs from Chat or Manage.
* Opt-in notification rules can offer a local "Ask Hermes?"
action, and Bridge tools can target a specific Android device.
Reliability
* Long chats avoid premature transport fallback, phone context
reaches upstream Hermes on supported paths, malformed server
addresses fail safely, and credential files cannot be served
through relay media.
Polish
* Multi-image galleries plus smoother streaming Markdown and long tables.
@@ -82,9 +82,9 @@ class BargeInPreferencesRepository(
constructor(context: Context) : this(context.relayDataStore)
companion object {
private val KEY_ENABLED = booleanPreferencesKey("barge_in_enabled")
private val KEY_SENSITIVITY = stringPreferencesKey("barge_in_sensitivity")
private val KEY_RESUME_AFTER_INTERRUPTION =
internal val KEY_ENABLED = booleanPreferencesKey("barge_in_enabled")
internal val KEY_SENSITIVITY = stringPreferencesKey("barge_in_sensitivity")
internal val KEY_RESUME_AFTER_INTERRUPTION =
booleanPreferencesKey("barge_in_resume_after_interruption")
}
@@ -120,8 +120,42 @@ data class ChatMessage(
* no status affix.
*/
val deliveryStatus: MessageDeliveryStatus? = null,
/**
* Client-side lifecycle for a promoted/durable Hermes run that belongs to
* this assistant turn. The same message owns the state from promotion
* through delivery so Chat never needs a separate system notice and final
* reply for one task. On the normal post-turn history reconcile this field
* is carried forward with the rest of the client-only enrichment whenever
* the live message can be matched to its server row.
*/
val backgroundTask: BackgroundTaskState? = null,
)
/** One Chat-visible identity for a promoted/durable realtime Hermes run. */
data class BackgroundTaskState(
/** Relay run id when supplied; otherwise a stable id derived from the message. */
val id: String,
/** Short objective derived from the associated user turn. */
val title: String,
/** ADR 33 tier: `promoted` or `durable`. */
val tier: String = "promoted",
val phase: BackgroundTaskPhase = BackgroundTaskPhase.RUNNING,
/** Latest meaningful progress line, deliberately not a raw event trace. */
val statusLine: String? = null,
val completedToolCount: Int = 0,
val queuedCount: Int = 0,
val startedAt: Long = System.currentTimeMillis(),
)
enum class BackgroundTaskPhase {
RUNNING,
WAITING,
DELIVERING,
COMPLETE,
FAILED,
CANCELLED,
}
/**
* Structured details about a phone-local voice intent that was dispatched
* in-process via [com.hermesandroid.relay.network.relay.BridgeCommandHandler.handleLocalCommand].
@@ -0,0 +1,162 @@
package com.hermesandroid.relay.data
import android.content.Context
import androidx.datastore.core.DataStore
import androidx.datastore.preferences.core.Preferences
import androidx.datastore.preferences.core.edit
import androidx.datastore.preferences.core.stringPreferencesKey
import kotlinx.coroutines.flow.first
import kotlinx.serialization.Serializable
import kotlinx.serialization.encodeToString
import kotlinx.serialization.json.Json
/**
* Durable, client-owned snapshot of one in-flight chat turn.
*
* Hermes history is authoritative once a turn finishes, but it cannot recreate
* transient UI that existed before persistence (live reasoning, a running tool,
* an interactive ask, or the latest lifecycle line). This checkpoint bridges
* that gap across Activity recreation and process death. It deliberately stores
* no entered secret/approval response; only the server-issued ask is retained.
*/
@Serializable
data class ChatTurnCheckpoint(
val schemaVersion: Int = CURRENT_SCHEMA,
val contextKey: String,
val sessionId: String,
val liveSessionId: String? = null,
val transport: String,
val user: ChatTurnUserCheckpoint,
val assistant: ChatTurnAssistantCheckpoint,
val turnStatus: String? = null,
val priorUserMessageCount: Int,
val baselineAssistantCount: Int,
val pendingAsk: ChatTurnAskCheckpoint? = null,
val startedAt: Long,
val updatedAt: Long,
) {
companion object {
const val CURRENT_SCHEMA = 1
const val MAX_AGE_MS = 24L * 60L * 60L * 1_000L
}
}
@Serializable
data class ChatTurnUserCheckpoint(
val id: String,
val content: String,
val timestamp: Long,
)
@Serializable
data class ChatTurnAssistantCheckpoint(
val id: String,
val content: String = "",
val timestamp: Long,
val isStreaming: Boolean = true,
val thinkingContent: String = "",
val isThinkingStreaming: Boolean = false,
val inputTokens: Int? = null,
val outputTokens: Int? = null,
val totalTokens: Int? = null,
val estimatedCost: Double? = null,
val agentName: String? = null,
val badges: List<String> = emptyList(),
val cards: List<HermesCard> = emptyList(),
val cardDispatches: List<HermesCardDispatch> = emptyList(),
val toolCalls: List<ChatTurnToolCheckpoint> = emptyList(),
val backgroundTask: ChatTurnBackgroundTaskCheckpoint? = null,
)
@Serializable
data class ChatTurnToolCheckpoint(
val id: String? = null,
val name: String,
val result: String? = null,
val success: Boolean? = null,
val isComplete: Boolean = false,
val error: String? = null,
val runId: String? = null,
val provenance: String? = null,
val startedAt: Long,
val completedAt: Long? = null,
val isGenerating: Boolean = false,
val taskIndex: Int? = null,
val taskLabel: String? = null,
)
@Serializable
data class ChatTurnBackgroundTaskCheckpoint(
val id: String,
val title: String,
val tier: String,
val phase: String,
val statusLine: String? = null,
val completedToolCount: Int = 0,
val queuedCount: Int = 0,
val startedAt: Long,
)
@Serializable
data class ChatTurnAskCheckpoint(
val kind: String,
val requestId: String? = null,
val text: String,
val choices: List<String>? = null,
val envVar: String? = null,
val timeoutSeconds: Int,
val messageId: String,
val cardKey: String,
/** Original receive time, used to preserve an ask's expiry after reopen. */
val receivedAt: Long,
)
interface ChatTurnCheckpointStore {
suspend fun read(): ChatTurnCheckpoint?
suspend fun write(checkpoint: ChatTurnCheckpoint)
suspend fun clear()
}
class DataStoreChatTurnCheckpointStore(
private val dataStore: DataStore<Preferences>,
private val now: () -> Long = System::currentTimeMillis,
) : ChatTurnCheckpointStore {
constructor(context: Context) : this(context.applicationContext.relayDataStore)
private val json = Json {
ignoreUnknownKeys = true
encodeDefaults = true
isLenient = true
}
override suspend fun read(): ChatTurnCheckpoint? {
val raw = runCatching { dataStore.data.first()[KEY_CHECKPOINT] }.getOrNull()
?: return null
val checkpoint = runCatching { json.decodeFromString<ChatTurnCheckpoint>(raw) }.getOrNull()
if (checkpoint == null ||
checkpoint.schemaVersion != ChatTurnCheckpoint.CURRENT_SCHEMA ||
now() - checkpoint.updatedAt > ChatTurnCheckpoint.MAX_AGE_MS
) {
// Cleanup is best-effort. In particular, Windows can briefly keep
// the just-read preferences file open and reject DataStore's atomic
// temp-file rename; an invalid checkpoint must still read as null.
runCatching { clear() }
return null
}
return checkpoint
}
override suspend fun write(checkpoint: ChatTurnCheckpoint) {
dataStore.edit { preferences ->
preferences[KEY_CHECKPOINT] = json.encodeToString(checkpoint)
}
}
override suspend fun clear() {
dataStore.edit { preferences -> preferences.remove(KEY_CHECKPOINT) }
}
private companion object {
val KEY_CHECKPOINT = stringPreferencesKey("chat_inflight_turn_checkpoint_v1")
}
}
@@ -0,0 +1,69 @@
package com.hermesandroid.relay.data
/**
* A process event that upstream Hermes injected into transcript history as a
* synthetic user message.
*
* Hermes intentionally persists these events with role=user so the agent can
* react to them without breaking message-role alternation. UI code should use
* [ChatMessage.hermesProcessNotificationOrNull] to present them as process
* notices without changing their canonical role or content.
*/
data class HermesProcessNotification(
val processId: String,
val headline: String,
val detail: String?,
)
/**
* Recognizes the exact envelope emitted by upstream
* `tools.process_registry.format_process_notification` for background-process
* completion and watch events.
*
* The parser deliberately excludes other `[IMPORTANT: ...]` messages. Those
* can carry unrelated agent instructions and must continue through the normal
* transcript renderer.
*/
object HermesProcessNotificationParser {
private const val ENVELOPE_PREFIX = "[IMPORTANT: Background process "
private const val HEADLINE_PREFIX = "Background process "
fun parse(content: String): HermesProcessNotification? {
val normalized = content.trim()
if (!normalized.startsWith(ENVELOPE_PREFIX) || !normalized.endsWith(']')) {
return null
}
val body = normalized
.removePrefix("[IMPORTANT: ")
.dropLast(1)
val headline = body.substringBefore('\n').trim()
if (!headline.startsWith(HEADLINE_PREFIX)) return null
val identityAndStatus = headline.removePrefix(HEADLINE_PREFIX)
val processId = identityAndStatus.substringBefore(' ')
val status = identityAndStatus.substringAfter(' ', missingDelimiterValue = "")
if (processId.isBlank() || status.isBlank()) return null
val detail = body
.substringAfter('\n', missingDelimiterValue = "")
.trim()
.ifBlank { null }
return HermesProcessNotification(
processId = processId,
headline = headline,
detail = detail,
)
}
}
/**
* Returns the upstream process-notification presentation model only for the
* canonical synthetic user-row shape. The original [ChatMessage.role] remains
* [MessageRole.USER].
*/
fun ChatMessage.hermesProcessNotificationOrNull(): HermesProcessNotification? =
takeIf { it.role == MessageRole.USER }
?.content
?.let(HermesProcessNotificationParser::parse)
@@ -0,0 +1,191 @@
package com.hermesandroid.relay.data
/**
* One-tap bundles over voice settings that already exist in the app and relay.
*
* Presets intentionally do not own voice identity or routing: engine, audio
* route, provider, model, voice, enhanced-voice overrides, and background-run
* concurrency all remain exactly as the user configured them. A preset only
* coordinates interaction ergonomics, barge-in, Realtime trace/session
* behavior, and the existing ADR 33 background-delivery controls.
*/
enum class VoiceModePreset(
val displayName: String,
val shortLabel: String,
val description: String,
internal val localSettings: VoicePresetLocalSettings,
internal val bargeInUpdate: VoicePresetBargeInUpdate,
val promotionUpdate: VoicePresetPromotionUpdate,
) {
HandsFree(
displayName = "Hands-free",
shortLabel = "Hands-free",
description =
"Continuous listening, exact answers, detailed trace, and low-noise " +
"spoken progress after 15 seconds. Your barge-in choice is preserved.",
localSettings = VoicePresetLocalSettings(
interactionMode = "continuous",
silenceThresholdMs = 1250L,
realtimeTraceDetails = true,
realtimePersistentSession = true,
),
// Barge-in remains an explicit experimental opt-in until echo and
// self-recording hardening is complete. Never enable it via a preset.
bargeInUpdate = VoicePresetBargeInUpdate(),
promotionUpdate = VoicePresetPromotionUpdate(
enabled = true,
promoteAfterMs = 6000,
backgroundDefaultMode = "promote",
spokenHandoff = true,
progressSpokenAfterMs = 15000,
progressRepeatMs = 90000,
resultDelivery = "speak_verbatim",
),
),
LowLatency(
displayName = "Low latency",
shortLabel = "Fast",
description =
"Tap capture, the shortest supported silence window, a persistent " +
"session, and a fast visual handoff for long work.",
localSettings = VoicePresetLocalSettings(
interactionMode = "tap",
silenceThresholdMs = 750L,
realtimeTraceDetails = false,
realtimePersistentSession = true,
),
bargeInUpdate = VoicePresetBargeInUpdate(enabled = false),
promotionUpdate = VoicePresetPromotionUpdate(
enabled = true,
promoteAfterMs = 2500,
backgroundDefaultMode = "promote",
spokenHandoff = false,
progressSpokenAfterMs = 0,
resultDelivery = "speak_when_idle",
),
),
CarefulTools(
displayName = "Careful tools",
shortLabel = "Careful",
description =
"Hold-to-talk, uninterrupted foreground tool runs, a detailed trace, and exact result delivery.",
localSettings = VoicePresetLocalSettings(
interactionMode = "hold",
silenceThresholdMs = 1750L,
realtimeTraceDetails = true,
realtimePersistentSession = true,
),
bargeInUpdate = VoicePresetBargeInUpdate(enabled = false),
promotionUpdate = VoicePresetPromotionUpdate(
enabled = false,
backgroundDefaultMode = "foreground",
spokenHandoff = false,
progressSpokenAfterMs = 0,
resultDelivery = "speak_verbatim",
),
),
QuietVisualOnly(
displayName = "Quiet / visual-only",
shortLabel = "Quiet",
description =
"Manual capture with visual long-task handoffs and results. Normal short voice replies still speak.",
localSettings = VoicePresetLocalSettings(
interactionMode = "tap",
silenceThresholdMs = 1250L,
realtimeTraceDetails = true,
realtimePersistentSession = true,
),
bargeInUpdate = VoicePresetBargeInUpdate(enabled = false),
promotionUpdate = VoicePresetPromotionUpdate(
enabled = true,
promoteAfterMs = 6000,
backgroundDefaultMode = "promote",
spokenHandoff = false,
progressSpokenAfterMs = 0,
resultDelivery = "visual_only",
),
);
/** Apply only fields owned by this preset; every other value is preserved. */
fun applyTo(current: VoiceModePresetState): VoiceModePresetState =
current.copy(
voiceSettings = current.voiceSettings.copy(
interactionMode = localSettings.interactionMode,
silenceThresholdMs = localSettings.silenceThresholdMs,
realtimeTraceDetails = localSettings.realtimeTraceDetails,
realtimePersistentSession = localSettings.realtimePersistentSession,
),
bargeInPreferences = current.bargeInPreferences.copy(
enabled = bargeInUpdate.enabled ?: current.bargeInPreferences.enabled,
sensitivity =
bargeInUpdate.sensitivity ?: current.bargeInPreferences.sensitivity,
resumeAfterInterruption = bargeInUpdate.resumeAfterInterruption
?: current.bargeInPreferences.resumeAfterInterruption,
),
promotion = current.promotion?.let(promotionUpdate::applyTo),
)
/** A preset is active only when every field it owns still matches. */
fun matches(current: VoiceModePresetState): Boolean =
current.promotion != null && applyTo(current) == current
}
/** Snapshot used by the pure preset reducer and active-preset detector. */
data class VoiceModePresetState(
val voiceSettings: VoiceSettings,
val bargeInPreferences: BargeInPreferences,
val promotion: VoicePresetPromotionSettings?,
)
/** Relay promotion values mirrored without introducing a data -> network dependency. */
data class VoicePresetPromotionSettings(
val enabled: Boolean = true,
val promoteAfterMs: Int = 6000,
val backgroundDefaultMode: String = "promote",
val spokenHandoff: Boolean = true,
val progressSpokenAfterMs: Int = 0,
val progressRepeatMs: Int = 90000,
val resultDelivery: String = "speak_verbatim",
val maxBackgroundRuns: Int = 1,
)
/** Nullable fields map directly to RelayVoiceClient's partial PATCH contract. */
data class VoicePresetPromotionUpdate(
val enabled: Boolean? = null,
val promoteAfterMs: Int? = null,
val backgroundDefaultMode: String? = null,
val spokenHandoff: Boolean? = null,
val progressSpokenAfterMs: Int? = null,
val progressRepeatMs: Int? = null,
val resultDelivery: String? = null,
val maxBackgroundRuns: Int? = null,
) {
internal fun applyTo(current: VoicePresetPromotionSettings): VoicePresetPromotionSettings =
current.copy(
enabled = enabled ?: current.enabled,
promoteAfterMs = promoteAfterMs ?: current.promoteAfterMs,
backgroundDefaultMode = backgroundDefaultMode ?: current.backgroundDefaultMode,
spokenHandoff = spokenHandoff ?: current.spokenHandoff,
progressSpokenAfterMs = progressSpokenAfterMs ?: current.progressSpokenAfterMs,
progressRepeatMs = progressRepeatMs ?: current.progressRepeatMs,
resultDelivery = resultDelivery ?: current.resultDelivery,
maxBackgroundRuns = maxBackgroundRuns ?: current.maxBackgroundRuns,
)
}
internal data class VoicePresetLocalSettings(
val interactionMode: String,
val silenceThresholdMs: Long,
val realtimeTraceDetails: Boolean,
val realtimePersistentSession: Boolean,
)
internal data class VoicePresetBargeInUpdate(
val enabled: Boolean? = null,
val sensitivity: BargeInSensitivity? = null,
val resumeAfterInterruption: Boolean? = null,
)
/** Null means the current manual values are Custom. */
fun detectVoiceModePreset(current: VoiceModePresetState): VoiceModePreset? =
VoiceModePreset.entries.firstOrNull { it.matches(current) }
@@ -374,4 +374,31 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
suspend fun setRealtimePersistentSession(enabled: Boolean) {
dataStore.edit { it[KEY_REALTIME_PERSISTENT_SESSION] = enabled }
}
/**
* Atomically apply the phone-side portion of [preset]. Only fields owned by
* the preset are written, so route/provider/model/voice overrides and other
* preferences remain untouched. Barge-in shares this DataStore and is
* updated in the same transaction so observers never see a half-applied
* local preset.
*/
suspend fun applyModePreset(preset: VoiceModePreset) {
val local = preset.localSettings
val bargeIn = preset.bargeInUpdate
dataStore.edit { prefs ->
prefs[KEY_INTERACTION_MODE] = local.interactionMode
prefs[KEY_SILENCE_THRESHOLD_MS] = local.silenceThresholdMs.coerceAtLeast(500L)
prefs[KEY_REALTIME_TRACE_DETAILS] = local.realtimeTraceDetails
prefs[KEY_REALTIME_PERSISTENT_SESSION] = local.realtimePersistentSession
bargeIn.enabled?.let {
prefs[BargeInPreferencesRepository.KEY_ENABLED] = it
}
bargeIn.sensitivity?.let {
prefs[BargeInPreferencesRepository.KEY_SENSITIVITY] = it.name
}
bargeIn.resumeAfterInterruption?.let {
prefs[BargeInPreferencesRepository.KEY_RESUME_AFTER_INTERRUPTION] = it
}
}
}
}
@@ -19,6 +19,7 @@ import kotlinx.coroutines.withTimeoutOrNull
import kotlinx.serialization.Serializable
import kotlinx.serialization.SerialName
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.addJsonObject
import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.buildJsonArray
@@ -2896,10 +2897,12 @@ class RelayVoiceClient(
val resultPreviewValue = (obj["result_preview"] as? JsonPrimitive)?.contentOrNull
?: (obj["result"] as? JsonPrimitive)?.contentOrNull
val reasonValue = (obj["reason"] as? JsonPrimitive)?.contentOrNull
val errorValue = (obj["error"] as? JsonPrimitive)?.contentOrNull
RealtimeVoiceEvent(
type = (obj["type"] as? JsonPrimitive)?.content ?: "unknown",
source = (obj["source"] as? JsonPrimitive)?.contentOrNull,
message = (obj["message"] as? JsonPrimitive)?.contentOrNull
?: errorValue
?: reasonValue,
reason = reasonValue,
statusKey = (obj["status_key"] as? JsonPrimitive)?.contentOrNull,
@@ -2924,7 +2927,7 @@ class RelayVoiceClient(
toolName = toolNameValue,
toolCallId = toolCallIdValue,
resultPreview = resultPreviewValue,
success = (obj["success"] as? JsonPrimitive)?.contentOrNull?.toBooleanStrictOrNull(),
success = realtimeEventSuccess(obj),
audioBase64 = (obj["audio_base64"] as? JsonPrimitive)?.contentOrNull,
byteCount = (obj["byte_count"] as? JsonPrimitive)?.intOrNull,
sampleRate = (obj["sample_rate"] as? JsonPrimitive)?.intOrNull,
@@ -2937,7 +2940,8 @@ class RelayVoiceClient(
tier = (obj["tier"] as? JsonPrimitive)?.contentOrNull,
floor = (obj["floor"] as? JsonPrimitive)?.contentOrNull,
activeToolName = (obj["active_tool_name"] as? JsonPrimitive)?.contentOrNull,
completedToolCount = (obj["completed_tool_count"] as? JsonPrimitive)?.intOrNull,
completedToolCount = (obj["completed_tool_count"] as? JsonPrimitive)?.intOrNull
?: (obj["tool_count"] as? JsonPrimitive)?.intOrNull,
elapsedMs = (obj["elapsed_ms"] as? JsonPrimitive)?.longOrNull,
queuedCount = (obj["queued_count"] as? JsonPrimitive)?.intOrNull,
delivery = (obj["delivery"] as? JsonPrimitive)?.contentOrNull,
@@ -3330,6 +3334,10 @@ data class RealtimeVoiceEvent(
get() = type == "voice.audio.delta" || type == "voice.output_audio.delta"
}
internal fun realtimeEventSuccess(obj: JsonObject): Boolean? =
(obj["success"] as? JsonPrimitive)?.contentOrNull?.toBooleanStrictOrNull()
?: (obj["ok"] as? JsonPrimitive)?.contentOrNull?.toBooleanStrictOrNull()
data class VoiceHandoffEvent(
val label: String,
val detail: String? = null,
@@ -2,8 +2,11 @@ package com.hermesandroid.relay.network.upstream
import android.util.Log
import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.data.BackgroundTaskPhase
import com.hermesandroid.relay.data.BackgroundTaskState
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.ChatSession
import com.hermesandroid.relay.data.ChatTurnCheckpoint
import com.hermesandroid.relay.data.HermesCard
import com.hermesandroid.relay.data.MessageDeliveryStatus
import com.hermesandroid.relay.data.MessageRole
@@ -242,9 +245,8 @@ class ChatHandler {
// User-chosen Thread names (sessionId → name), authoritative over the
// server's auto-title — applied in [updateSessions] so the gateway's async
// auto-titler can't clobber the name. Fed by ChatViewModel. In-memory for
// now (survives list refreshes within a session); cross-restart persistence
// is a follow-up (see TODO).
// auto-titler can't clobber the name. ChatViewModel hydrates this map from
// ThreadNameStore, so names survive both list refreshes and app restarts.
private val userThreadNames = mutableMapOf<String, String>()
/** Record a user-chosen name for one Thread session + re-apply it now. */
@@ -366,6 +368,31 @@ class ChatHandler {
}
}
/** Attach the first Chat-visible state for a promoted/durable Hermes run. */
fun setBackgroundTask(messageId: String, task: BackgroundTaskState) {
_messages.update { list ->
list.map { message ->
if (message.id == messageId) message.copy(backgroundTask = task) else message
}
}
}
/** Update an existing task in place; no-op when the message/task is absent. */
fun updateBackgroundTask(
messageId: String,
transform: (BackgroundTaskState) -> BackgroundTaskState,
) {
_messages.update { list ->
list.map { message ->
if (message.id == messageId && message.backgroundTask != null) {
message.copy(backgroundTask = transform(message.backgroundTask))
} else {
message
}
}
}
}
/**
* Append a SYSTEM-role notice bubble (e.g. a gateway interactive ask the
* phone can't answer). SYSTEM role keeps it out of the voice TTS observer
@@ -863,6 +890,136 @@ class ChatHandler {
}
}
/**
* Rehydrate the last client-owned state of an unfinished turn.
*
* The caller loads server history first. That means the user row may already
* be present while the assistant row is not yet durable; positional matching
* avoids duplicating short repeated prompts. Rich assistant-only state is
* then restored so thinking and tool cards do not reset to an empty spinner.
*/
fun restoreInFlightTurn(
checkpoint: ChatTurnCheckpoint,
upstreamAssistantText: String? = null,
) {
val user = checkpoint.user
val assistant = checkpoint.assistant
val upstreamText = upstreamAssistantText.orEmpty()
val currentAssistant = _messages.value.lastOrNull { it.id == assistant.id }
val restoredContent = listOf(
assistant.content,
upstreamText,
currentAssistant?.content.orEmpty(),
).maxByOrNull { it.length }.orEmpty()
val checkpointTools = assistant.toolCalls.map { tool ->
ToolCall(
id = tool.id,
name = tool.name,
args = null,
result = tool.result,
success = tool.success,
isComplete = tool.isComplete,
error = tool.error,
runId = tool.runId,
provenance = tool.provenance,
startedAt = tool.startedAt,
completedAt = tool.completedAt,
isGenerating = tool.isGenerating,
taskIndex = tool.taskIndex,
taskLabel = tool.taskLabel,
)
}
val currentTools = currentAssistant?.toolCalls.orEmpty()
val restoredTools = buildList {
checkpointTools.forEach { checkpointTool ->
val live = currentTools.firstOrNull {
(it.id != null && it.id == checkpointTool.id) ||
(it.id == null && checkpointTool.id == null &&
it.name == checkpointTool.name &&
it.taskIndex == checkpointTool.taskIndex)
}
add(live ?: checkpointTool)
}
currentTools.filterTo(this) { live ->
checkpointTools.none { checkpointTool ->
(live.id != null && live.id == checkpointTool.id) ||
(live.id == null && checkpointTool.id == null &&
live.name == checkpointTool.name &&
live.taskIndex == checkpointTool.taskIndex)
}
}
}
val restoredBackgroundTask = assistant.backgroundTask?.let { task ->
BackgroundTaskState(
id = task.id,
title = task.title,
tier = task.tier,
phase = runCatching { BackgroundTaskPhase.valueOf(task.phase) }
.getOrDefault(BackgroundTaskPhase.RUNNING),
statusLine = task.statusLine,
completedToolCount = task.completedToolCount,
queuedCount = task.queuedCount,
startedAt = task.startedAt,
)
}
val restoredAssistant = ChatMessage(
id = assistant.id,
role = MessageRole.ASSISTANT,
content = restoredContent,
timestamp = assistant.timestamp,
isStreaming = true,
toolCalls = restoredTools,
thinkingContent = listOf(
assistant.thinkingContent,
currentAssistant?.thinkingContent.orEmpty(),
).maxByOrNull { it.length }.orEmpty(),
isThinkingStreaming = currentAssistant?.isThinkingStreaming
?: assistant.isThinkingStreaming,
inputTokens = currentAssistant?.inputTokens ?: assistant.inputTokens,
outputTokens = currentAssistant?.outputTokens ?: assistant.outputTokens,
totalTokens = currentAssistant?.totalTokens ?: assistant.totalTokens,
estimatedCost = currentAssistant?.estimatedCost ?: assistant.estimatedCost,
agentName = currentAssistant?.agentName ?: assistant.agentName ?: activeAgentName,
badges = (assistant.badges + currentAssistant?.badges.orEmpty()).distinct(),
cards = currentAssistant?.cards?.takeIf { it.isNotEmpty() } ?: assistant.cards,
cardDispatches = currentAssistant?.cardDispatches?.takeIf { it.isNotEmpty() }
?: assistant.cardDispatches,
backgroundTask = currentAssistant?.backgroundTask ?: restoredBackgroundTask,
)
activeAgentName = restoredAssistant.agentName ?: activeAgentName
_messages.update { current ->
val withoutOldAssistant = current.filterNot { it.id == assistant.id }
val users = withoutOldAssistant.filter { it.role == MessageRole.USER }
val positionalUser = users.getOrNull(checkpoint.priorUserMessageCount)
val hasUser = withoutOldAssistant.any { it.id == user.id } ||
positionalUser?.content?.trim() == user.content.trim()
val withUser = if (hasUser) {
withoutOldAssistant
} else {
withoutOldAssistant + ChatMessage(
id = user.id,
role = MessageRole.USER,
content = user.content,
timestamp = user.timestamp,
)
}
val insertBeforeAsk = withUser.indexOfFirst {
it.clientOnly && it.id.startsWith("ask-")
}
val restored = if (insertBeforeAsk >= 0) {
withUser.toMutableList().apply { add(insertBeforeAsk, restoredAssistant) }
} else {
withUser + restoredAssistant
}
restored.let { list ->
if (list.size > MAX_MESSAGES) list.drop(list.size - MAX_MESSAGES) else list
}
}
_isStreaming.value = true
_turnStatus.value = checkpoint.turnStatus ?: "Reconnecting to the active turn…"
}
fun clearMessages() {
_messages.value = emptyList()
// Drop any pending line buffers / dedupe state so a fresh session
@@ -2815,6 +2972,10 @@ class ChatHandler {
fun setLastSentMessage(text: String) {
_lastSentMessage.value = text
}
fun clearLastSentMessage() {
_lastSentMessage.value = null
}
}
/**
@@ -25,6 +25,7 @@ import kotlinx.serialization.json.booleanOrNull
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.contentOrNull
import kotlinx.serialization.json.intOrNull
import kotlinx.serialization.json.longOrNull
import kotlinx.serialization.json.put
import okhttp3.OkHttpClient
import okhttp3.Request
@@ -32,7 +33,9 @@ import okhttp3.Response
import okhttp3.WebSocket
import okhttp3.WebSocketListener
import java.util.concurrent.ConcurrentHashMap
import java.util.concurrent.CountDownLatch
import java.util.concurrent.TimeUnit
import java.util.concurrent.atomic.AtomicBoolean
import java.util.concurrent.atomic.AtomicLong
/**
@@ -49,8 +52,9 @@ import java.util.concurrent.atomic.AtomicLong
* in a background thread that keeps emitting on the id it was STARTED with,
* regardless of WS state. So a mid-turn socket drop is recovered by
* reconnecting the socket and KEEPING the in-flight session id (see
* [attemptMidTurnRejoin]) — NOT by `session.resume`, which mints a brand-new
* id + a fresh agent rebuilt from DB and would orphan the still-running turn.
* [attemptMidTurnRejoin]). Current upstream Hermes can also rebind a detached
* live session through `session.activate` / `session.resume`; the direct socket
* rejoin remains compatible with older gateways and avoids an extra RPC.
* No background reconnect loops; a fresh send reconnects on demand.
*
* Auth: every connect attempt mints a FRESH single-use ws-ticket (30s TTL)
@@ -151,6 +155,8 @@ class GatewayChatClient(
private const val CONNECT_FAILURE_COOLDOWN_MS = 5_000L
private const val RATE_LIMIT_COOLDOWN_MS = 300_000L
private const val CONNECT_ATTEMPTS = 2
private const val INBOUND_BIND_TIMEOUT_MS = 2_000L
private const val CANCELLED_TURN_SUBMIT_WAIT_MS = 2_000L
/** Distinct socket-loss (flap) events per turn we'll try to recover from. */
private const val MAX_TURN_REJOINS = 4
@@ -213,6 +219,14 @@ class GatewayChatClient(
private val _connectionState = MutableStateFlow(GatewayConnectionState.Idle)
val connectionState: StateFlow<GatewayConnectionState> = _connectionState.asStateFlow()
/**
* Per-socket feature probe for upstream's session-scoped `process.*` RPCs.
* Method-not-found marks the current socket unsupported; reconnecting resets
* this to [GatewayProcessCapability.Unknown] so a server upgrade is noticed.
*/
private val _processCapability = MutableStateFlow(GatewayProcessCapability.Unknown)
val processCapability: StateFlow<GatewayProcessCapability> = _processCapability.asStateFlow()
/**
* Active personality the gateway is applying, as a config value ("none" when
* the overlay is cleared, otherwise the personality name). Tracks the
@@ -286,6 +300,8 @@ class GatewayChatClient(
private var readySignal: CompletableDeferred<Unit>? = null
private val rpcId = AtomicLong(1)
/** Invalidates an older async prewarm when a newer session selection wins. */
private val prewarmRequestGeneration = AtomicLong(0)
private val pendingRpcs = ConcurrentHashMap<Long, CompletableDeferred<JsonObject>>()
/** Live (per-connection) session id ←→ the stored DB id it was resumed/created from. */
@@ -295,6 +311,10 @@ class GatewayChatClient(
@Volatile
private var storedSessionId: String? = null
/** Profile namespace that owns [liveSessionId]; stored IDs are not globally unique. */
@Volatile
private var liveSessionProfile: String? = null
/**
* Supplies the profile to bind each `session.create` / `session.resume` to —
* the upstream `tui_gateway` opens that profile's HERMES_HOME/db and builds
@@ -331,6 +351,49 @@ class GatewayChatClient(
@Volatile
private var activeTurn: GatewayTurn? = null
/**
* Upstream may emit the interrupted turn's tail and terminal event after
* `session.interrupt` returns. Keep a short exact-session tombstone so that
* tail cannot be mistaken for an unsolicited completion or complete a
* newly submitted turn. A same-session send briefly waits for this drain
* before `prompt.submit`; once a turn actually started, its tombstone stays
* until the required terminal event even when that submit wait elapses.
*/
@Volatile
private var cancelledTurnDrain: CancelledTurnDrain? = null
private data class CancelledTurnDrain(
val storedSessionId: String,
val liveSessionId: String,
val submitWaitUntilMs: Long,
val terminalRequired: Boolean,
)
/**
* Creates UI callbacks when the server starts a turn that has no matching
* [sendTurn] call (for example a background-process completion). The
* provider is consulted only for an explicit `message.start` whose live
* session id exactly matches this client's active session.
*/
@Volatile
private var unsolicitedTurnProvider: ((storedSessionId: String) -> GatewayInboundTurnRegistration?)? = null
/** Recover persisted events that may have completed while the socket was closed. */
@Volatile
private var coldPrewarmSessionReadyListener: ((storedSessionId: String) -> Unit)? = null
/** Exact-session completion observed without a bound live mapper. */
@Volatile
private var unmatchedTurnCompleteListener:
((storedSessionId: String, expectedAssistantText: String?) -> Unit)? = null
/**
* Connection-level process listener. Unlike [GatewayTurnCallbacks], this is
* consulted even when there is no locally initiated [activeTurn].
*/
@Volatile
private var processEventListener: ((GatewayProcessEvent) -> Unit)? = null
/**
* Which upload RPC name this socket understands — set after the first
* successful upload so the legacy fallback is probed at most once per
@@ -411,7 +474,10 @@ class GatewayChatClient(
// alive AND the requested session already live). A "cold" turn re-pays
// ticket/ws/session — exactly the asymmetry vs always-connected desktop.
val socketWarm = webSocket != null && readySignal?.isCompleted == true
val sessionWarm = liveSessionId != null && storedSessionId == sessionId && sessionId != null
val sessionWarm = liveSessionId != null &&
storedSessionId == sessionId &&
sessionId != null &&
liveSessionProfile == currentSessionProfile()
turn.tracer.warm(socketWarm && sessionWarm)
scope.launch {
try {
@@ -428,6 +494,7 @@ class GatewayChatClient(
}
}
if (turn.cancelled) return@launch
if (!awaitCancelledTurnDrain(turn, storedSessionId)) return@launch
activeTurn = turn
turn.armWatchdog()
val submitted = rpc(
@@ -457,7 +524,7 @@ class GatewayChatClient(
)
return@launch
}
activeTurn = null
if (activeTurn === turn) activeTurn = null
turn.disarmWatchdog()
throw GatewayPreflightException(
submitted.exceptionOrNull()?.message ?: "prompt.submit failed",
@@ -470,7 +537,7 @@ class GatewayChatClient(
// read-the-absence exercise.
Log.i(TAG, "Gateway turn submitted (session=$storedSessionId)")
} catch (e: Exception) {
activeTurn = null
if (activeTurn === turn) activeTurn = null
if (!turn.cancelled) {
Log.w(TAG, "Gateway preflight failed: ${e.message}")
turn.tracer.done("preflight-fail")
@@ -483,8 +550,11 @@ class GatewayChatClient(
/** Drop the remembered session so the next send creates a fresh one. */
fun clearSession() {
prewarmRequestGeneration.incrementAndGet()
liveSessionId = null
storedSessionId = null
liveSessionProfile = null
cancelledTurnDrain = null
}
/**
@@ -496,6 +566,10 @@ class GatewayChatClient(
*/
fun hasActiveTurn(): Boolean = activeTurn?.ended == false
/** Live id to persist beside a durable stored id while a turn is active. */
fun currentLiveSessionId(storedId: String): String? =
liveSessionId?.takeIf { storedSessionId == storedId }
/**
* Point this client at a new dashboard route (e.g. LAN→Tailscale after a
* sustained network change). If a turn is in flight, the current socket is
@@ -528,6 +602,30 @@ class GatewayChatClient(
if (!enabled && !AppForegroundTracker.isForeground.value) scheduleBackgroundClose()
}
/** Process polling must not silently undo the normal background socket close. */
fun isBackgroundProcessPollingAllowed(): Boolean =
AppForegroundTracker.isForeground.value || keepAliveInBackground
fun setUnsolicitedTurnProvider(
provider: ((storedSessionId: String) -> GatewayInboundTurnRegistration?)?,
) {
unsolicitedTurnProvider = provider
}
fun setColdPrewarmSessionReadyListener(listener: ((storedSessionId: String) -> Unit)?) {
coldPrewarmSessionReadyListener = listener
}
fun setUnmatchedTurnCompleteListener(
listener: ((storedSessionId: String, expectedAssistantText: String?) -> Unit)?,
) {
unmatchedTurnCompleteListener = listener
}
fun setProcessEventListener(listener: ((GatewayProcessEvent) -> Unit)?) {
processEventListener = listener
}
/**
* Establish the socket (and resume an existing session) ahead of the
* user's first send, so a warm turn reaches first token in tens of ms
@@ -557,15 +655,150 @@ class GatewayChatClient(
* session, so this path is the common case, not the edge case).
*/
suspend fun prewarmAwait(storedSessionId: String?): Boolean {
val requestGeneration = prewarmRequestGeneration.incrementAndGet()
val requestedProfile = currentSessionProfile()
val wasLiveForRequestedSession = storedSessionId != null &&
liveSessionId != null &&
this.storedSessionId == storedSessionId &&
liveSessionProfile == requestedProfile
try {
connectMutex.withLock {
ensureConnected()
if (storedSessionId != null) resumeForPrewarm(storedSessionId)
if (storedSessionId != null) {
resumeForPrewarm(storedSessionId, requestedProfile, requestGeneration)
}
}
} catch (e: Exception) {
Log.d(TAG, "Gateway prewarm skipped: ${e.message}")
}
return liveSessionId != null
val sessionReady = storedSessionId != null &&
liveSessionId != null &&
this.storedSessionId == storedSessionId &&
liveSessionProfile == requestedProfile
if (!wasLiveForRequestedSession && sessionReady) {
callbackDispatcher {
coldPrewarmSessionReadyListener?.invoke(storedSessionId)
}
}
return sessionReady
}
/**
* Reattach callbacks to a turn that survived the Android UI/process.
*
* New Hermes gateways expose `session.activate`, which attaches the new
* WebSocket transport to the exact live id saved in the client checkpoint.
* If that id has already been reaped (or the method is unavailable), fall
* back to `session.resume` by durable session id. Its `running` + `inflight`
* fields decide whether a live mapper is installed or history should settle
* the turn instead.
*/
suspend fun recoverTurn(
storedId: String,
preferredLiveId: String?,
callbacks: GatewayTurnCallbacks,
): Result<GatewaySessionRecovery> = runCatching {
require(storedId.isNotBlank()) { "stored session id required" }
val requestedProfile = currentSessionProfile()
connectMutex.withLock {
val existing = activeTurn
if (existing != null && !existing.ended) {
throw GatewayRpcException("a gateway turn is already attached")
}
ensureConnected()
var response: JsonObject? = null
var boundTurn: GatewayTurn? = null
if (!preferredLiveId.isNullOrBlank()) {
// Bind before session.activate: upstream swaps the live session's
// transport during the RPC, so an immediate next delta must not
// fall through the active-turn gate while the ack is in flight.
liveSessionId = preferredLiveId
storedSessionId = storedId
liveSessionProfile = requestedProfile
boundTurn = GatewayTurn(
callbacks = dispatchOn(callbacks),
dedupeAdjacentMessageStarts = true,
).also { turn ->
turn.markRecoveredStarted()
activeTurn = turn
}
val activated = rpc(
"session.activate",
buildJsonObject {
put("session_id", preferredLiveId)
},
)
response = activated.getOrNull()
if (response == null) {
if (activeTurn === boundTurn) activeTurn = null
boundTurn.detach()
boundTurn = null
Log.d(
TAG,
"Exact live-session activation unavailable; resuming durable session " +
"(${activated.exceptionOrNull()?.message})",
)
}
}
if (response == null) {
response = rpc(
"session.resume",
buildJsonObject {
put("session_id", storedId)
put("cols", DEFAULT_COLS)
requestedProfile?.let { put("profile", it) }
},
).getOrElse { error -> throw error }
}
val recoveredLiveId = response.stringField("session_id")
?: throw GatewayRpcException("session recovery returned no session_id")
liveSessionId = recoveredLiveId
storedSessionId = storedId
liveSessionProfile = requestedProfile
updateCancelledDrainLiveSession(storedId, recoveredLiveId)
(response["info"] as? JsonObject)?.let { applySessionInfo(it) }
val inflight = (response["inflight"] as? JsonObject)?.let { value ->
GatewayInflightTurn(
user = value.stringField("user").orEmpty(),
assistant = value.stringField("assistant").orEmpty(),
streaming = value.booleanField("streaming") == true,
)
}
val running = response.booleanField("running") == true || inflight?.streaming == true
if (running) {
if (boundTurn == null || boundTurn.ended) {
boundTurn = GatewayTurn(
callbacks = dispatchOn(callbacks),
dedupeAdjacentMessageStarts = true,
).also { turn ->
turn.markRecoveredStarted()
activeTurn = turn
}
}
boundTurn.armWatchdog()
} else {
if (boundTurn != null) {
if (activeTurn === boundTurn) activeTurn = null
boundTurn.detach()
}
boundTurn = null
}
GatewaySessionRecovery(
storedSessionId = storedId,
liveSessionId = recoveredLiveId,
running = running,
status = response.stringField("status"),
inflight = inflight,
handle = boundTurn?.takeUnless { it.ended },
)
}
}
/**
@@ -669,6 +902,61 @@ class GatewayChatClient(
.onSuccess { commandsCatalogCache = it }
}
/**
* Fetch the current chat session's running and recently-finished background
* processes. Callers never provide a session id: this wrapper resolves and
* sends the exact LIVE gateway id, not the stored history id exposed to UI.
*
* A remembered stored session is resumed after a socket reconnect. A brand-
* new chat has no server process ownership yet and therefore returns an empty
* snapshot without creating an otherwise-empty session.
*/
suspend fun listProcesses(): Result<List<GatewayProcess>> {
if (_processCapability.value == GatewayProcessCapability.Unsupported) {
return Result.failure(processFeatureUnsupported())
}
if (liveSessionId == null && storedSessionId == null) {
return Result.success(emptyList())
}
val sid = ensureLiveProcessSession().getOrElse { return Result.failure(it) }
val result = rpc(
"process.list",
buildJsonObject { put("session_id", sid) },
)
if (result.isFailure) {
markProcessUnsupportedIfNeeded(result.exceptionOrNull())
return Result.failure(result.exceptionOrNull() ?: GatewayRpcException("process.list failed"))
}
_processCapability.value = GatewayProcessCapability.Supported
return Result.success(
((result.getOrThrow()["processes"] as? JsonArray).orEmpty()).mapNotNull(::parseGatewayProcess),
)
}
/** Stop one process owned by the current live gateway session. */
suspend fun killProcess(processId: String): Result<Unit> {
if (processId.isBlank()) {
return Result.failure(GatewayRpcException("process id required"))
}
if (_processCapability.value == GatewayProcessCapability.Unsupported) {
return Result.failure(processFeatureUnsupported())
}
val sid = ensureLiveProcessSession().getOrElse { return Result.failure(it) }
val result = rpc(
"process.kill",
buildJsonObject {
put("session_id", sid)
put("process_id", processId)
},
)
if (result.isFailure) {
markProcessUnsupportedIfNeeded(result.exceptionOrNull())
return Result.failure(result.exceptionOrNull() ?: GatewayRpcException("process.kill failed"))
}
_processCapability.value = GatewayProcessCapability.Supported
return Result.success(Unit)
}
/**
* Run a full slash command line (`slash.exec {session_id, command}`) on
* the live session. Returns the raw result object; failures carry the
@@ -911,6 +1199,11 @@ class GatewayChatClient(
fun shutdown() {
activeTurn?.cancel()
activeTurn = null
cancelledTurnDrain = null
unsolicitedTurnProvider = null
coldPrewarmSessionReadyListener = null
unmatchedTurnCompleteListener = null
processEventListener = null
closeSocket("client shutdown")
backgroundCloseJob?.cancel()
// Stop the foreground collector — a replaced client must not keep
@@ -951,6 +1244,7 @@ class GatewayChatClient(
private suspend fun connectOnce() {
val connectStart = System.nanoTime()
_processCapability.value = GatewayProcessCapability.Unknown
_connectionState.value = GatewayConnectionState.MintingTicket
val ticket = dashboardClient.requestWsTicket().getOrElse { e ->
throw GatewayConnectAttemptException("ws-ticket mint failed: ${e.message}")
@@ -989,28 +1283,39 @@ class GatewayChatClient(
* (no [GatewayTurn] context). Failure is silent — the real send's
* [ensureSession] will resume-or-create properly.
*/
private suspend fun resumeForPrewarm(storedId: String) {
// Never resume while a turn is in flight: a resume mints a NEW live
// session id, and the running turn's events (still tagged with the
// OLD id) would then be filtered out as "foreign" — orphaning the
// turn and letting a stale reconcile repaint an earlier reply. This
// is the screen-return (prewarm) variant of the same hazard the
// mid-turn rejoin avoids by NOT resuming.
private suspend fun resumeForPrewarm(
storedId: String,
requestedProfile: String?,
requestGeneration: Long,
) {
// Never change session binding while this client already owns a live
// mapper. Current upstream may reuse the same live session, while older
// builds mint a new id; either way the existing mapper owns recovery.
if (activeTurn != null) return
if (liveSessionId != null && storedSessionId == storedId) return
if (
liveSessionId != null &&
storedSessionId == storedId &&
liveSessionProfile == requestedProfile
) return
val resumed = rpc(
"session.resume",
buildJsonObject {
put("session_id", storedId)
put("cols", DEFAULT_COLS)
currentSessionProfile()?.let { put("profile", it) }
requestedProfile?.let { put("profile", it) }
},
)
val result = resumed.getOrNull()
val live = result?.stringField("session_id")
if (live != null) {
if (
live != null &&
activeTurn == null &&
prewarmRequestGeneration.get() == requestGeneration
) {
liveSessionId = live
storedSessionId = storedId
liveSessionProfile = requestedProfile
updateCancelledDrainLiveSession(storedId, live)
// Paint the session's real model/provider/effort/etc NOW from the
// resume result's embedded `info` (same shape session.info carries),
// so a reopened session shows its ACTUAL model immediately instead of
@@ -1055,13 +1360,84 @@ class GatewayChatClient(
}
}
/** Resolve a process RPC against the exact live id, resuming after reconnect when possible. */
private suspend fun ensureLiveProcessSession(): Result<String> {
val requestedProfile = currentSessionProfile()
liveSessionId?.takeIf { liveSessionProfile == requestedProfile }
?.let { return Result.success(it) }
val rememberedStoredId = storedSessionId
?: return Result.failure(GatewayRpcException("no live session"))
val requestGeneration = prewarmRequestGeneration.incrementAndGet()
return try {
connectMutex.withLock {
ensureConnected()
if (liveSessionId == null || liveSessionProfile != requestedProfile) {
resumeForPrewarm(rememberedStoredId, requestedProfile, requestGeneration)
}
}
val resumedLiveId = liveSessionId
if (
resumedLiveId != null &&
storedSessionId == rememberedStoredId &&
liveSessionProfile == requestedProfile
) {
Result.success(resumedLiveId)
} else {
Result.failure(GatewayRpcException("could not resume live session"))
}
} catch (e: Exception) {
Result.failure(e)
}
}
private fun parseGatewayProcess(element: kotlinx.serialization.json.JsonElement): GatewayProcess? {
val process = element as? JsonObject ?: return null
val id = process.stringField("session_id")?.takeIf { it.isNotBlank() } ?: return null
return GatewayProcess(
id = id,
command = process.stringField("command").orEmpty(),
cwd = process.stringField("cwd"),
pid = (process["pid"] as? JsonPrimitive)?.longOrNull,
startedAt = process.stringField("started_at"),
uptimeSeconds = (process["uptime_seconds"] as? JsonPrimitive)?.longOrNull ?: 0L,
status = process.stringField("status") ?: "unknown",
outputPreview = process.stringField("output_preview")?.takeIf { it.isNotEmpty() },
outputTail = process.stringField("output_tail")?.takeIf { it.isNotEmpty() },
exitCode = (process["exit_code"] as? JsonPrimitive)?.intOrNull,
detached = (process["detached"] as? JsonPrimitive)?.booleanOrNull ?: false,
notifyOnComplete = (process["notify_on_complete"] as? JsonPrimitive)?.booleanOrNull ?: false,
sessionScoped = (process["session_scoped"] as? JsonPrimitive)?.booleanOrNull ?: false,
watchPatterns = (process["watch_patterns"] as? JsonArray).orEmpty()
.mapNotNull { (it as? JsonPrimitive)?.contentOrNull },
watchHit = (process["watch_hit"] as? JsonPrimitive)?.booleanOrNull ?: false,
)
}
private fun markProcessUnsupportedIfNeeded(error: Throwable?) {
if (error.isMethodNotFound()) {
_processCapability.value = GatewayProcessCapability.Unsupported
}
}
private fun processFeatureUnsupported(): GatewayRpcException =
GatewayRpcException("background process RPCs are not supported by this gateway", JSONRPC_METHOD_NOT_FOUND)
/** Must hold [connectMutex]. Resolves [liveSessionId] for the requested stored id. */
private suspend fun ensureSession(
requestedStoredId: String?,
newSessionTitle: String?,
turn: GatewayTurn,
) {
if (liveSessionId != null && storedSessionId == requestedStoredId && requestedStoredId != null) {
val requestedProfile = currentSessionProfile()
if (requestedStoredId != null && requestedStoredId != storedSessionId) {
cancelledTurnDrain = null
}
if (
liveSessionId != null &&
storedSessionId == requestedStoredId &&
requestedStoredId != null &&
liveSessionProfile == requestedProfile
) {
return
}
@@ -1071,7 +1447,7 @@ class GatewayChatClient(
buildJsonObject {
put("session_id", requestedStoredId)
put("cols", DEFAULT_COLS)
currentSessionProfile()?.let { put("profile", it) }
requestedProfile?.let { put("profile", it) }
},
)
val result = resumed.getOrNull()
@@ -1079,6 +1455,8 @@ class GatewayChatClient(
if (live != null) {
liveSessionId = live
storedSessionId = requestedStoredId
liveSessionProfile = requestedProfile
updateCancelledDrainLiveSession(requestedStoredId, live)
(result["info"] as? JsonObject)?.let { applySessionInfo(it) }
return
}
@@ -1094,7 +1472,7 @@ class GatewayChatClient(
buildJsonObject {
put("cols", DEFAULT_COLS)
if (!newSessionTitle.isNullOrBlank()) put("title", newSessionTitle)
currentSessionProfile()?.let { put("profile", it) }
requestedProfile?.let { put("profile", it) }
// Bind the in-chat overrides to the new session as its
// per-session overrides. Upstream tui_gateway session.create
// reads `model`/`provider` (→ model_override), `reasoning_effort`
@@ -1121,6 +1499,8 @@ class GatewayChatClient(
val stored = created.stringField("stored_session_id") ?: live
liveSessionId = live
storedSessionId = stored
liveSessionProfile = requestedProfile
if (cancelledTurnDrain?.storedSessionId != stored) cancelledTurnDrain = null
turn.callbacks.onSessionId(stored)
}
@@ -1201,7 +1581,7 @@ class GatewayChatClient(
// delta types log length only, everything else logs a payload
// excerpt so on-device diagnosis doesn't read absences.
when (type) {
"message.delta", "reasoning.delta", "thinking.delta" ->
"message.delta", "reasoning.delta", "thinking.delta", "agent.terminal.output" ->
Log.d(TAG, "GW ← $type (${payload?.toString()?.length ?: 0} chars)")
else ->
Log.d(TAG, "GW ← $type | ${payload?.toString()?.take(300) ?: "{}"}")
@@ -1226,18 +1606,110 @@ class GatewayChatClient(
payload?.let { applySessionInfo(it) }
}
val turn = activeTurn ?: return
// Foreign-session events (another client's chat on the same gateway) are not ours.
if (eventSessionId != null && liveSessionId != null && eventSessionId != liveSessionId) {
return
}
dispatchProcessEvent(type, payload, eventSessionId)
if (consumeCancelledTurnEvent(type, eventSessionId)) return
var turn = activeTurn
if (turn == null && type == "message.start") {
// Unsolicited turns are accepted only with an explicit exact live-
// session match. This gateway stream is process-wide; treating a
// missing id as ours would leak another desktop tab's response.
val liveId = liveSessionId
val storedId = storedSessionId
if (!eventSessionId.isNullOrBlank() &&
eventSessionId == liveId &&
!storedId.isNullOrBlank()
) {
val registration = unsolicitedTurnProvider?.invoke(storedId)
if (registration != null) {
val inboundTurn = GatewayTurn(
callbacks = dispatchOn(registration.callbacks),
dedupeAdjacentMessageStarts = true,
)
if (bindInboundTurn(registration, inboundTurn)) {
turn = inboundTurn
activeTurn = inboundTurn
inboundTurn.armWatchdog()
Log.i(TAG, "Accepted unsolicited gateway turn for session=$storedId")
} else {
Log.i(TAG, "Deferred unsolicited gateway turn for session=$storedId")
}
}
}
}
if (turn == null) {
// A foreground SSE/realtime turn may already own Chat, or the
// socket may have reconnected after message.start. The exact
// terminal event is still authoritative and persisted upstream;
// recover it through history once the UI becomes idle.
val storedId = storedSessionId
if (type == "message.complete" &&
!eventSessionId.isNullOrBlank() &&
eventSessionId == liveSessionId &&
!storedId.isNullOrBlank()
) {
val expectedText = payload?.stringField("text")
callbackDispatcher {
unmatchedTurnCompleteListener?.invoke(storedId, expectedText)
}
}
return
}
turn.onEvent(type, payload)
if (turn.ended) {
activeTurn = null
if (activeTurn === turn) activeTurn = null
if (AppForegroundTracker.isForeground.value.not()) scheduleBackgroundClose()
}
}
/**
* Deliver session-scoped process events before the active-turn gate. The
* gateway socket is process-wide, so an exact non-blank live id match is
* required; missing/foreign ids must never leak another window's process.
*/
private fun dispatchProcessEvent(type: String, payload: JsonObject?, eventSessionId: String?) {
val liveId = liveSessionId ?: return
if (eventSessionId.isNullOrBlank() || eventSessionId != liveId) return
val event = when (type) {
"tool.complete" -> when (payload?.stringField("name")) {
"terminal", "process" -> GatewayProcessEvent.Invalidated(
GatewayProcessEvent.Trigger.TOOL_COMPLETE,
)
else -> null
}
"status.update" -> if (payload?.stringField("kind") == "process") {
GatewayProcessEvent.Invalidated(GatewayProcessEvent.Trigger.STATUS_UPDATE)
} else {
null
}
// Some upstream tool paths complete a background terminal launch
// without emitting tool.start/tool.complete to this UI session.
// The assistant still closes the initiating turn normally, so use
// that exact-session boundary as a cheap authoritative discovery
// fallback. process.list then starts the running-only poller.
"message.complete" -> GatewayProcessEvent.Invalidated(
GatewayProcessEvent.Trigger.MESSAGE_COMPLETE,
)
"agent.terminal.output" -> {
val processId = payload?.stringField("process_id")
val chunk = payload?.stringField("chunk")
if (!processId.isNullOrBlank() && chunk != null) {
GatewayProcessEvent.Output(processId, chunk)
} else {
null
}
}
"terminal.close" -> payload?.stringField("process_id")
?.takeIf { it.isNotBlank() }
?.let { GatewayProcessEvent.TerminalClosed(it) }
else -> null
} ?: return
callbackDispatcher { processEventListener?.invoke(event) }
}
private fun onSocketDown(reason: String) {
Log.i(TAG, "Gateway socket down ($reason)")
// Capture the in-flight session id BEFORE clearing it — the mid-turn
@@ -1249,6 +1721,7 @@ class GatewayChatClient(
liveSessionId = null
attachMethodForSocket = null
commandsCatalogCache = null
_processCapability.value = GatewayProcessCapability.Unknown
_connectionState.value = GatewayConnectionState.Idle
pendingRpcs.values.forEach {
it.completeExceptionally(GatewayRpcException("gateway connection lost"))
@@ -1256,7 +1729,7 @@ class GatewayChatClient(
pendingRpcs.clear()
val turn = activeTurn ?: return
if (turn.ended) {
activeTurn = null
if (activeTurn === turn) activeTurn = null
return
}
// A rejoin already owns recovery — a connect attempt failing inside
@@ -1276,7 +1749,7 @@ class GatewayChatClient(
}
}
} else {
activeTurn = null
if (activeTurn === turn) activeTurn = null
turn.failFromTransport("Connection to the gateway was lost")
}
}
@@ -1288,13 +1761,10 @@ class GatewayChatClient(
* Recover an in-flight turn after a mid-turn socket loss by reconnecting
* the SOCKET ONLY and keeping [preservedLiveId] as the live session id.
*
* Why not `session.resume`: upstream resume mints a brand-new session id
* and rebuilds a fresh agent from persisted DB history — it does NOT
* reattach to the running turn's thread, which keeps emitting on the OLD
* id over the shared gateway stream. Resuming would point our event filter
* at the wrong id and orphan the turn (the server finishes it and the
* answer is dropped — confirmed on-device). A bare reconnect lets the tail
* — including the final `message.complete` — keep matching this turn.
* A bare reconnect lets the tail — including the final
* `message.complete` — keep matching without another RPC. Current upstream
* can rebind live sessions too, but older builds cannot, so this remains the
* lowest-common-denominator same-client recovery path.
*
* Retries with backoff for up to [midTurnRejoinWindowMs] so a multi-second
* radio blip doesn't abandon the turn. Events emitted while the socket was
@@ -1354,6 +1824,7 @@ class GatewayChatClient(
liveSessionId = null
attachMethodForSocket = null
commandsCatalogCache = null
_processCapability.value = GatewayProcessCapability.Unknown
_connectionState.value = GatewayConnectionState.Idle
}
@@ -1499,8 +1970,9 @@ class GatewayChatClient(
private inner class GatewayTurn(
val callbacks: GatewayTurnCallbacks,
dedupeAdjacentMessageStarts: Boolean = false,
) : ActiveTurnHandle {
private val mapper = GatewayEventMapper(callbacks)
private val mapper = GatewayEventMapper(callbacks, dedupeAdjacentMessageStarts)
/** t0 = construction ≈ sendTurn entry (the moment the user sent). */
val tracer = TurnLatencyTracer("gateway")
@@ -1564,6 +2036,11 @@ class GatewayChatClient(
watchdog = null
}
/** Recovery attaches after the original prompt.submit, so it is already started. */
fun markRecoveredStarted() {
started = true
}
/** Transport-level failure after submit — surface as a stream error once. */
fun failFromTransport(message: String) {
disarmWatchdog()
@@ -1578,10 +2055,21 @@ class GatewayChatClient(
cancelled = true
disarmWatchdog()
tracer.done("cancelled")
if (activeTurn === this) activeTurn = null
if (activeTurn === this) {
armCancelledTurnDrain(terminalRequired = started)
activeTurn = null
}
interruptServerSide()
}
override fun detach() {
if (ended) return
cancelled = true
disarmWatchdog()
tracer.done("detached")
if (activeTurn === this) activeTurn = null
}
private fun interruptServerSide() {
val sid = liveSessionId ?: return
scope.launch {
@@ -1592,9 +2080,104 @@ class GatewayChatClient(
}
}
private fun armCancelledTurnDrain(terminalRequired: Boolean) {
val storedId = storedSessionId ?: return
val liveId = liveSessionId ?: return
cancelledTurnDrain = CancelledTurnDrain(
storedSessionId = storedId,
liveSessionId = liveId,
submitWaitUntilMs = System.currentTimeMillis() + CANCELLED_TURN_SUBMIT_WAIT_MS,
terminalRequired = terminalRequired,
)
}
private fun updateCancelledDrainLiveSession(storedId: String, liveId: String) {
val drain = cancelledTurnDrain ?: return
if (drain.storedSessionId == storedId) {
cancelledTurnDrain = drain.copy(liveSessionId = liveId)
}
}
/** Ignore the canceled turn's exact-session tail through its terminal event. */
private fun consumeCancelledTurnEvent(type: String, eventSessionId: String?): Boolean {
val drain = cancelledTurnDrain ?: return false
if (!drain.terminalRequired && System.currentTimeMillis() >= drain.submitWaitUntilMs) {
if (cancelledTurnDrain === drain) cancelledTurnDrain = null
return false
}
if (eventSessionId != drain.liveSessionId) return false
if (type == "message.complete" || type == "error") {
if (cancelledTurnDrain === drain) cancelledTurnDrain = null
}
Log.d(TAG, "Ignored canceled gateway turn event: $type")
return true
}
/**
* Preserve same-session event ordering after Stop. The interrupt terminal
* normally drains immediately. The wait is bounded so a slow cooperative
* interrupt does not indefinitely delay the next submit, but a started
* turn's tombstone remains until its terminal event and continues draining
* the old tail in front of the server-serialized next prompt.
*/
private suspend fun awaitCancelledTurnDrain(
turn: GatewayTurn,
targetStoredSessionId: String?,
): Boolean {
while (!turn.cancelled) {
val drain = cancelledTurnDrain ?: return true
if (drain.storedSessionId != targetStoredSessionId) return true
if (System.currentTimeMillis() >= drain.submitWaitUntilMs) {
if (!drain.terminalRequired && cancelledTurnDrain === drain) {
cancelledTurnDrain = null
}
return true
}
delay(25L)
}
return false
}
/**
* Serialize inbound ownership onto the callback/main dispatcher before any
* mapper callback is posted. A local SSE turn can otherwise start between
* the socket event and UI binding and lose its cancellable handle.
*/
private fun bindInboundTurn(
registration: GatewayInboundTurnRegistration,
turn: GatewayTurn,
): Boolean {
val pending = AtomicBoolean(true)
val completed = CountDownLatch(1)
var accepted = false
try {
callbackDispatcher {
try {
if (pending.compareAndSet(true, false)) {
accepted = registration.onHandle(turn)
}
} finally {
completed.countDown()
}
}
} catch (e: Exception) {
Log.w(TAG, "Inbound turn dispatcher rejected callback", e)
return false
}
if (!completed.await(INBOUND_BIND_TIMEOUT_MS, TimeUnit.MILLISECONDS)) {
// If the callback has not started, expire it so a late main-thread
// delivery cannot bind a handle the client already discarded. If
// it has started, wait for that short ownership check to finish.
if (pending.compareAndSet(true, false)) return false
completed.await()
}
return accepted
}
/** Wrap callbacks so every invocation lands on the callback dispatcher (main thread). */
private fun dispatchOn(callbacks: GatewayTurnCallbacks) = GatewayTurnCallbacks(
onSessionId = { v -> callbackDispatcher { callbacks.onSessionId(v) } },
onStart = { callbackDispatcher { callbacks.onStart() } },
onTextDelta = { v -> callbackDispatcher { callbacks.onTextDelta(v) } },
onThinkingDelta = { v -> callbackDispatcher { callbacks.onThinkingDelta(v) } },
onToolCallStart = { a, b -> callbackDispatcher { callbacks.onToolCallStart(a, b) } },
@@ -1664,3 +2247,6 @@ private fun Throwable?.isMethodNotFound(): Boolean {
private fun JsonObject.stringField(key: String): String? =
(get(key) as? JsonPrimitive)?.contentOrNull
private fun JsonObject.booleanField(key: String): Boolean? =
(get(key) as? JsonPrimitive)?.booleanOrNull
@@ -21,13 +21,17 @@ import kotlinx.serialization.json.intOrNull
* why dispatch is a manual `when (type)` over [JsonObject] rather than a
* sealed polymorphic hierarchy (which throws on unknown discriminators).
*/
class GatewayEventMapper(private val callbacks: GatewayTurnCallbacks) {
class GatewayEventMapper(
private val callbacks: GatewayTurnCallbacks,
private val dedupeAdjacentMessageStarts: Boolean = false,
) {
/** True once `message.complete` or `error` has been seen — the turn is over. */
var turnEnded: Boolean = false
private set
private var sawMessageStart = false
private var previousEventType: String? = null
private var sawTextDelta = false
private var sawThinkingDelta = false
private var syntheticToolCounter = 0
@@ -77,11 +81,18 @@ class GatewayEventMapper(private val callbacks: GatewayTurnCallbacks) {
}
"message.start" -> {
// The upstream background-completion poller currently emits
// message.start immediately before _run_prompt_submit(), which
// emits the same start again. Treat an adjacent pair as one
// boundary; a later start after any other event still closes
// the previous assistant message as before.
if (dedupeAdjacentMessageStarts && previousEventType == "message.start") return
// Gateway has no server-side message id (placeholder UUID
// stays). A second start inside one turn means a new
// assistant message began — close out the previous one.
if (sawMessageStart) callbacks.onTurnComplete()
sawMessageStart = true
callbacks.onStart()
}
"tool.generating" -> {
@@ -228,6 +239,7 @@ class GatewayEventMapper(private val callbacks: GatewayTurnCallbacks) {
// alike: ignore.
else -> Unit
}
previousEventType = type
}
private fun syntheticToolId(name: String): String {
@@ -81,8 +81,33 @@ fun resolveStreamingEndpointPreference(
*/
fun interface ActiveTurnHandle {
fun cancel()
/**
* Release this client's callbacks without interrupting server-side work.
* Gateway turns override this for process/UI teardown; transports that
* cannot be reattached retain their existing cancel behavior.
*/
fun detach() = cancel()
}
/** Partial text checkpoint returned by current upstream Hermes on live resume. */
data class GatewayInflightTurn(
val user: String,
val assistant: String,
val streaming: Boolean,
)
/** Result of reattaching Android to an existing durable Gateway session. */
data class GatewaySessionRecovery(
val storedSessionId: String,
val liveSessionId: String,
val running: Boolean,
val status: String?,
val inflight: GatewayInflightTurn?,
/** Non-null only when subsequent turn events are bound to [GatewayTurnCallbacks]. */
val handle: ActiveTurnHandle?,
)
/**
* One server-side interactive ask. The agent thread upstream is BLOCKED
* until the matching respond RPC arrives, the ask times out (resolves to ""
@@ -135,6 +160,67 @@ data class GatewaySubagentEvent(
enum class Phase { START, THINKING, TOOL, PROGRESS, COMPLETE }
}
/**
* One session-owned background process returned by the upstream gateway's
* `process.list` RPC. The registry calls its process id `session_id`; Android
* exposes it as [id] so it cannot be confused with either the stored chat id or
* the gateway's live, per-connection session id.
*
* [outputPreview] is the registry's short preview, while [outputTail] is the
* gateway's larger (currently 4,000-character) snapshot used to recover output
* missed while the WebSocket was unavailable. Unknown/new fields are ignored
* by the parser so this remains compatible with older and newer gateways.
*/
data class GatewayProcess(
val id: String,
val command: String,
val cwd: String? = null,
val pid: Long? = null,
val startedAt: String? = null,
val uptimeSeconds: Long = 0L,
val status: String,
val outputPreview: String? = null,
val outputTail: String? = null,
val exitCode: Int? = null,
val detached: Boolean = false,
val notifyOnComplete: Boolean = false,
val sessionScoped: Boolean = false,
val watchPatterns: List<String> = emptyList(),
val watchHit: Boolean = false,
) {
val isRunning: Boolean get() = status.equals("running", ignoreCase = true)
}
/** Whether this gateway socket supports the session-scoped process RPCs. */
enum class GatewayProcessCapability {
/** Not probed on this socket yet (or no socket is currently connected). */
Unknown,
/** A `process.list` / `process.kill` call succeeded. */
Supported,
/** The gateway returned JSON-RPC method-not-found for the process surface. */
Unsupported,
}
/**
* Connection-level background-process events. These are deliberately separate
* from [GatewayTurnCallbacks]: output and completion notifications can arrive
* while no app-initiated turn is active.
*/
sealed interface GatewayProcessEvent {
enum class Trigger { TOOL_COMPLETE, STATUS_UPDATE, MESSAGE_COMPLETE }
/** The process snapshot may have changed and should be refreshed. */
data class Invalidated(val trigger: Trigger) : GatewayProcessEvent
/** Live output from `agent.terminal.output`. */
data class Output(val processId: String, val chunk: String) : GatewayProcessEvent
/** The agent requested that its read-only terminal view be closed. */
data class TerminalClosed(val processId: String) : GatewayProcessEvent
}
/**
* One provider from the gateway `model.options` RPC — the curated, authenticated
* provider/model list the upstream desktop + TUI model picker uses (NOT the
@@ -208,6 +294,8 @@ data class GatewayReasoningSettings(
class GatewayTurnCallbacks(
/** Stored (DB) session id — fired on session create/rotate so the drawer + persistence stay correct. */
val onSessionId: (String) -> Unit,
/** A gateway `message.start` opened an assistant response for this turn. */
val onStart: () -> Unit,
val onTextDelta: (String) -> Unit,
val onThinkingDelta: (String) -> Unit,
val onToolCallStart: (toolCallId: String, toolName: String) -> Unit,
@@ -238,3 +326,18 @@ class GatewayTurnCallbacks(
*/
val onStatusUpdate: (kind: String?, text: String) -> Unit = { _, _ -> },
)
/**
* UI registration for one server-initiated gateway turn.
*
* Background-process completion is converted upstream into a normal assistant
* turn on the originating session. It has no matching client [GatewayChatClient.sendTurn]
* call, so the client asks the active conversation for callbacks when the first
* `message.start` arrives. [onHandle] binds the resulting cancellable turn into
* the same Stop/steer lifecycle as a locally submitted turn.
*/
class GatewayInboundTurnRegistration(
val callbacks: GatewayTurnCallbacks,
/** Main-thread admission. False leaves the server turn unbound for history recovery. */
val onHandle: (ActiveTurnHandle) -> Boolean,
)
@@ -188,8 +188,8 @@ private fun Modifier.topFadeEdge(fade: Dp = 28.dp): Modifier = this
)
}
/** Resolved motion/accessibility posture for clean mode. */
private data class CleanMotionState(
/** Shared OS motion/accessibility posture for animated chat affordances. */
internal data class AccessibleMotionState(
/** OS animator scale is non-zero (i.e. system animations are ON). */
val osAnimations: Boolean,
/** TalkBack-style touch exploration is active — faded text is unreadable
@@ -198,7 +198,7 @@ private data class CleanMotionState(
)
@Composable
private fun rememberCleanMotionState(): CleanMotionState {
internal fun rememberAccessibleMotionState(): AccessibleMotionState {
val context = LocalContext.current
// ANIMATOR_DURATION_SCALE == 0 is the platform "remove animations" / many
// OEM "reduce motion" toggles. Read once on entry; a mid-mode toggle is
@@ -225,7 +225,10 @@ private fun rememberCleanMotionState(): CleanMotionState {
a11y?.addTouchExplorationStateChangeListener(listener)
onDispose { a11y?.removeTouchExplorationStateChangeListener(listener) }
}
return CleanMotionState(osAnimations = osAnimations, touchExploration = touchExploration)
return AccessibleMotionState(
osAnimations = osAnimations,
touchExploration = touchExploration,
)
}
/**
@@ -511,7 +514,7 @@ fun CleanChatMode(
onExit: () -> Unit,
modifier: Modifier = Modifier,
) {
val motion = rememberCleanMotionState()
val motion = rememberAccessibleMotionState()
val sphereAnimated = animationEnabled && motion.osAnimations
// Faded text is unreadable to touch exploration, so the text path goes
// static (readable + announced) whenever TalkBack is exploring.
@@ -0,0 +1,353 @@
package com.hermesandroid.relay.ui.components
import android.graphics.BitmapFactory
import android.net.Uri
import androidx.compose.foundation.ExperimentalFoundationApi
import androidx.compose.foundation.Image
import androidx.compose.foundation.background
import androidx.compose.foundation.combinedClickable
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.aspectRatio
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.widthIn
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.filled.BrokenImage
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.Icon
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Text
import androidx.compose.runtime.Composable
import androidx.compose.runtime.LaunchedEffect
import androidx.compose.runtime.collectAsState
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateMapOf
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.rememberCoroutineScope
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.clip
import androidx.compose.ui.graphics.ImageBitmap
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.graphics.asImageBitmap
import androidx.compose.ui.layout.ContentScale
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.platform.testTag
import androidx.compose.ui.unit.Dp
import androidx.compose.ui.unit.dp
import coil3.compose.AsyncImagePainter
import coil3.compose.SubcomposeAsyncImage
import coil3.compose.SubcomposeAsyncImageContent
import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.data.AttachmentRenderMode
import com.hermesandroid.relay.data.AttachmentState
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
/**
* One item in a message's attachment render order. Loaded images are grouped
* into [Gallery] only when there are at least two; every other attachment
* keeps its original index so retry/manual-fetch callbacks still target the
* exact [com.hermesandroid.relay.data.ChatMessage.attachments] entry.
*/
internal sealed interface AttachmentLayoutItem {
data class Single(val attachmentIndex: Int) : AttachmentLayoutItem
data class Gallery(val attachmentIndices: List<Int>) : AttachmentLayoutItem
}
/**
* Build the attachment render plan without reordering non-image cards. The
* gallery occupies the first eligible image's slot and absorbs the remaining
* loaded images, including images separated by a PDF/file card.
*/
internal fun attachmentLayoutItems(attachments: List<Attachment>): List<AttachmentLayoutItem> {
return buildList {
var index = 0
while (index < attachments.size) {
if (!attachments[index].isGalleryImage()) {
add(AttachmentLayoutItem.Single(index))
index++
continue
}
val run = buildList {
var cursor = index
while (cursor < attachments.size && attachments[cursor].isGalleryImage()) {
add(cursor)
cursor++
}
}
if (run.size >= 2) add(AttachmentLayoutItem.Gallery(run))
else add(AttachmentLayoutItem.Single(index))
index += run.size
}
}
}
private fun Attachment.isGalleryImage(): Boolean =
state == AttachmentState.LOADED && renderMode == AttachmentRenderMode.IMAGE
/** Two-column, non-lazy rows for a gallery nested inside the chat LazyColumn. */
internal fun galleryRows(itemCount: Int): List<List<Int>> =
(0 until itemCount.coerceAtLeast(0)).chunked(GALLERY_COLUMNS)
internal fun galleryPreviewIndices(itemCount: Int): List<Int> =
(0 until itemCount.coerceAtLeast(0)).take(GALLERY_PREVIEW_LIMIT)
/**
* Telegram-style media group for two or more loaded image attachments.
*
* The chat bubble shows a compact two-column grid. Tapping a tile opens the
* full-screen horizontal pager at that image; per-image blur reveal, long-
* press actions, and one-tap Save remain available instead of regressing the
* single-image attachment behavior.
*/
@OptIn(ExperimentalFoundationApi::class)
@Composable
fun AttachmentGallery(
attachments: List<Attachment>,
modifier: Modifier = Modifier,
maxWidth: Dp = 280.dp,
) {
if (attachments.size < 2) return
val context = LocalContext.current
val scope = rememberCoroutineScope()
val blurMode = LocalMediaBlurMode.current
val revealed = remember { mutableStateMapOf<String, Boolean>() }
var viewerStartIndex by remember { mutableStateOf<Int?>(null) }
viewerStartIndex?.let { startIndex ->
AttachmentGalleryViewer(
attachments = attachments,
initialIndex = startIndex.coerceIn(attachments.indices),
initiallyRevealedKeys = revealed
.filterValues { it }
.keys,
onDismiss = { viewerStartIndex = null },
)
}
Column(
modifier = modifier
.widthIn(max = maxWidth)
.fillMaxWidth()
.semantics { contentDescription = "${attachments.size} image gallery" },
verticalArrangement = Arrangement.spacedBy(GALLERY_GAP),
) {
val previewIndices = galleryPreviewIndices(attachments.size)
galleryRows(previewIndices.size).forEach { row ->
Row(
modifier = Modifier.fillMaxWidth(),
horizontalArrangement = Arrangement.spacedBy(GALLERY_GAP),
) {
row.forEach { previewIndex ->
val galleryIndex = previewIndices[previewIndex]
val attachment = attachments[galleryIndex]
val attachmentKey = galleryAttachmentKey(attachment, galleryIndex)
val blurred = revealed[attachmentKey] != true &&
shouldBlurImage(blurMode, attachment.sensitive)
var menuExpanded by remember(attachment, galleryIndex) { mutableStateOf(false) }
Box(
modifier = Modifier
.weight(1f)
// An odd final tile spans both columns without
// becoming a full-width square taller than the grid.
.aspectRatio(if (row.size == 1) 2f else 1f),
) {
BlurredMedia(
blurred = blurred,
onReveal = { revealed[attachmentKey] = true },
modifier = Modifier.fillMaxSize(),
) {
GalleryImageTile(
attachment = attachment,
position = galleryIndex,
count = attachments.size,
modifier = Modifier
.fillMaxSize()
.testTag("attachment-gallery-tile-$galleryIndex")
.clip(RoundedCornerShape(GALLERY_CORNER))
.combinedClickable(
onClick = { viewerStartIndex = galleryIndex },
onLongClick = { menuExpanded = true },
),
)
}
if (!blurred) {
SaveOverlayButton(
onClick = {
scope.launch { saveAttachment(context, attachment) }
},
modifier = Modifier
.align(Alignment.TopEnd)
.padding(4.dp),
)
}
AttachmentActionsMenu(
expanded = menuExpanded,
onDismiss = { menuExpanded = false },
context = context,
scope = scope,
attachment = attachment,
)
val hiddenCount = attachments.size - GALLERY_PREVIEW_LIMIT
if (
hiddenCount > 0 &&
previewIndex == GALLERY_PREVIEW_LIMIT - 1
) {
Box(
modifier = Modifier
.align(Alignment.BottomEnd)
.padding(7.dp)
.clip(RoundedCornerShape(50))
.background(Color.Black.copy(alpha = 0.68f))
.padding(horizontal = 9.dp, vertical = 4.dp),
) {
Text(
text = "+$hiddenCount",
style = MaterialTheme.typography.labelMedium,
color = Color.White,
)
}
}
}
}
}
}
}
}
@Composable
private fun GalleryImageTile(
attachment: Attachment,
position: Int,
count: Int,
modifier: Modifier,
) {
val description = listOfNotNull(
attachment.fileName?.takeIf { it.isNotBlank() },
"image ${position + 1} of $count",
).joinToString(", ")
val cachedUri = attachment.cachedUri?.takeIf { it.isNotBlank() }
if (cachedUri != null) {
SubcomposeAsyncImage(
model = Uri.parse(cachedUri),
contentDescription = description,
contentScale = ContentScale.Crop,
modifier = modifier,
) {
val state by painter.state.collectAsState()
when (state) {
is AsyncImagePainter.State.Success -> SubcomposeAsyncImageContent()
is AsyncImagePainter.State.Loading -> GalleryImagePlaceholder(modifier = Modifier.fillMaxSize())
else -> GalleryImageFailure(description, Modifier.fillMaxSize())
}
}
return
}
var bitmap by remember(attachment.content) { mutableStateOf<ImageBitmap?>(null) }
var failed by remember(attachment.content) { mutableStateOf(false) }
LaunchedEffect(attachment.content) {
val decoded = withContext(Dispatchers.IO) {
runCatching {
val bytes = android.util.Base64.decode(
attachment.content,
android.util.Base64.DEFAULT,
)
decodeGalleryBitmap(bytes)?.asImageBitmap()
}.getOrNull()
}
if (decoded != null) bitmap = decoded else failed = true
}
when {
bitmap != null -> Image(
bitmap = bitmap!!,
contentDescription = description,
contentScale = ContentScale.Crop,
modifier = modifier,
)
failed -> GalleryImageFailure(description, modifier)
else -> GalleryImagePlaceholder(modifier)
}
}
@Composable
private fun GalleryImagePlaceholder(modifier: Modifier) {
Box(
modifier = modifier.background(MaterialTheme.colorScheme.surfaceVariant),
contentAlignment = Alignment.Center,
) {
CircularProgressIndicator(modifier = Modifier.size(22.dp), strokeWidth = 2.dp)
}
}
@Composable
private fun GalleryImageFailure(description: String, modifier: Modifier) {
Box(
modifier = modifier.background(MaterialTheme.colorScheme.surfaceVariant),
contentAlignment = Alignment.Center,
) {
Icon(
imageVector = Icons.Filled.BrokenImage,
contentDescription = "Couldn't load $description",
tint = MaterialTheme.colorScheme.onSurfaceVariant,
modifier = Modifier.size(28.dp),
)
}
}
/** Decode a bounded thumbnail rather than retaining every full-size image. */
private fun decodeGalleryBitmap(bytes: ByteArray): android.graphics.Bitmap? {
if (bytes.isEmpty()) return null
val bounds = BitmapFactory.Options().apply { inJustDecodeBounds = true }
BitmapFactory.decodeByteArray(bytes, 0, bytes.size, bounds)
if (bounds.outWidth <= 0 || bounds.outHeight <= 0) return null
var sample = 1
while (
bounds.outWidth / sample > GALLERY_DECODE_TARGET_PX ||
bounds.outHeight / sample > GALLERY_DECODE_TARGET_PX
) {
sample *= 2
}
val options = BitmapFactory.Options().apply { inSampleSize = sample }
return BitmapFactory.decodeByteArray(bytes, 0, bytes.size, options)
}
private const val GALLERY_COLUMNS = 2
private const val GALLERY_PREVIEW_LIMIT = 4
private const val GALLERY_DECODE_TARGET_PX = 512
private val GALLERY_GAP = 3.dp
private val GALLERY_CORNER = 8.dp
internal fun galleryAttachmentKey(attachment: Attachment, index: Int): String =
attachment.relayToken?.takeIf { it.isNotBlank() }
?: attachment.cachedUri?.takeIf { it.isNotBlank() }
?: buildString {
append(attachment.fileName.orEmpty())
append('|')
append(attachment.contentType)
append('|')
append(attachment.content.hashCode())
append('|')
append(index)
}
@@ -15,7 +15,8 @@ import androidx.compose.foundation.Image
import androidx.compose.foundation.background
import androidx.compose.foundation.clickable
import androidx.compose.foundation.gestures.detectTapGestures
import androidx.compose.foundation.gestures.detectTransformGestures
import androidx.compose.foundation.gestures.rememberTransformableState
import androidx.compose.foundation.gestures.transformable
import androidx.compose.foundation.horizontalScroll
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
@@ -34,6 +35,8 @@ import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.windowInsetsPadding
import androidx.compose.foundation.lazy.LazyColumn
import androidx.compose.foundation.lazy.items
import androidx.compose.foundation.pager.HorizontalPager
import androidx.compose.foundation.pager.rememberPagerState
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.text.selection.SelectionContainer
@@ -59,6 +62,7 @@ import androidx.compose.runtime.Composable
import androidx.compose.runtime.DisposableEffect
import androidx.compose.runtime.LaunchedEffect
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateMapOf
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.rememberCoroutineScope
@@ -77,8 +81,10 @@ import androidx.compose.ui.input.pointer.pointerInput
import androidx.compose.ui.layout.ContentScale
import androidx.compose.ui.layout.onSizeChanged
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.platform.testTag
import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.unit.dp
import androidx.compose.ui.unit.IntSize
import androidx.compose.ui.viewinterop.AndroidView
import androidx.compose.ui.window.Dialog
import androidx.compose.ui.window.DialogProperties
@@ -98,6 +104,7 @@ import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import kotlinx.coroutines.withContext
import java.io.File
import kotlin.math.abs
import kotlin.math.sqrt
// ---------------------------------------------------------------------------
@@ -191,13 +198,62 @@ fun BlurredMedia(
fun Modifier.zoomable(maxScale: Float = 6f): Modifier {
var scale by remember { mutableStateOf(1f) }
var offset by remember { mutableStateOf(Offset.Zero) }
return this
.pointerInput(Unit) {
detectTransformGestures { _, pan, zoom, _ ->
scale = (scale * zoom).coerceIn(1f, maxScale)
offset = if (scale > 1f) offset + pan else Offset.Zero
}
var viewportSize by remember { mutableStateOf(IntSize.Zero) }
fun maxOffset(forScale: Float): Offset = Offset(
x = ((forScale - 1f) * viewportSize.width / 2f).coerceAtLeast(0f),
y = ((forScale - 1f) * viewportSize.height / 2f).coerceAtLeast(0f),
)
fun clampOffset(candidate: Offset, forScale: Float): Offset {
val max = maxOffset(forScale)
return Offset(
x = candidate.x.coerceIn(-max.x, max.x),
y = candidate.y.coerceIn(-max.y, max.y),
)
}
val transformState = rememberTransformableState { _, zoomChange, panChange, _ ->
val nextScale = (scale * zoomChange).coerceIn(1f, maxScale)
offset = if (nextScale > 1f) {
clampOffset(offset + panChange, nextScale)
} else {
Offset.Zero
}
scale = nextScale
}
return this
.onSizeChanged {
viewportSize = it
offset = clampOffset(offset, scale)
}
// Let a one-finger drag bubble to HorizontalPager at 1×. Once the
// image is zoomed, the image owns panning; pinch zoom always works.
.transformable(
state = transformState,
canPan = { pan ->
if (scale <= 1f) {
false
} else {
val max = maxOffset(scale)
val canMoveHorizontally = when {
pan.x > 0f -> offset.x < max.x
pan.x < 0f -> offset.x > -max.x
else -> false
}
val canMoveVertically = when {
pan.y > 0f -> offset.y < max.y
pan.y < 0f -> offset.y > -max.y
else -> false
}
if (abs(pan.x) >= abs(pan.y)) {
canMoveHorizontally
} else {
canMoveVertically
}
}
},
)
.pointerInput(Unit) {
detectTapGestures(
onDoubleTap = {
@@ -345,6 +401,7 @@ fun AttachmentViewer(
MediaViewerToolbar(
title = title,
busy = busy,
actionsEnabled = !blurred,
onShare = onShare,
onSave = onSave,
onOpenExternal = onOpenExternal,
@@ -355,11 +412,204 @@ fun AttachmentViewer(
}
}
/**
* Full-screen viewer for an image attachment group. The pager starts at the
* tapped tile, swipes horizontally at 1×, and keeps the existing per-image
* zoom, blur, Save, Share, and Open-externally behavior.
*
* [initiallyRevealedKeys] carries reveal state from the grid so a sensitive
* image that was already uncovered is not unexpectedly hidden again on open.
*/
@Composable
internal fun AttachmentGalleryViewer(
attachments: List<Attachment>,
initialIndex: Int,
onDismiss: () -> Unit,
initiallyRevealedKeys: Set<String> = emptySet(),
modifier: Modifier = Modifier,
) {
if (attachments.isEmpty()) return
if (attachments.size == 1) {
AttachmentViewer(
attachment = attachments.first(),
onDismiss = onDismiss,
modifier = modifier,
initiallyRevealed = galleryAttachmentKey(attachments.first(), 0) in
initiallyRevealedKeys,
)
return
}
Dialog(
onDismissRequest = onDismiss,
properties = DialogProperties(usePlatformDefaultWidth = false),
) {
val context = LocalContext.current
AllowDeviceRotation()
val scope = rememberCoroutineScope()
var busy by remember { mutableStateOf(false) }
val revealed = remember { mutableStateMapOf<String, Boolean>() }
LaunchedEffect(initiallyRevealedKeys) {
initiallyRevealedKeys.forEach { revealed[it] = true }
}
val pagerState = rememberPagerState(
initialPage = initialIndex.coerceIn(attachments.indices),
pageCount = { attachments.size },
)
val currentIndex = pagerState.currentPage.coerceIn(attachments.indices)
val attachment = attachments[currentIndex]
val currentKey = galleryAttachmentKey(attachment, currentIndex)
val blurMode = LocalMediaBlurMode.current
val currentBlurred = revealed[currentKey] != true &&
shouldBlurImage(blurMode, attachment.sensitive)
val title = attachment.fileName
?: attachment.contentType.substringBefore(';').ifBlank { "Image" }
val toolbarTitle = "$title · ${currentIndex + 1} of ${attachments.size}"
// Capture the currently visible attachment in each click lambda. A
// swipe while IO is running must not redirect Save/Share to a new page.
fun runWithBytes(action: suspend (Attachment, ByteArray) -> Unit) {
if (currentBlurred || busy) return
val target = attachment
scope.launch {
busy = true
try {
val bytes = attachmentBytes(context, target)
if (bytes == null) {
viewerToast(context, "Couldn't read this image")
return@launch
}
action(target, bytes)
} catch (error: Exception) {
viewerToast(
context,
error.message?.takeIf { it.isNotBlank() }
?: "Couldn't complete that image action",
)
} finally {
busy = false
}
}
}
val onShare = {
runWithBytes { target, bytes ->
val uri = MediaSaver.stageForShare(
context,
bytes,
target.fileName,
target.contentType,
)
MediaSaver.share(context, uri, target.contentType)
}
}
val onSave = {
runWithBytes { target, bytes ->
when (val result = MediaSaver.saveImage(
context,
bytes,
target.fileName,
target.contentType,
)) {
is MediaSaver.SaveResult.Saved ->
viewerToast(context, "Saved to ${result.location}")
MediaSaver.SaveResult.UseShareInstead -> {
val uri = MediaSaver.stageForShare(
context,
bytes,
target.fileName,
target.contentType,
)
MediaSaver.share(context, uri, target.contentType)
}
is MediaSaver.SaveResult.Failed ->
viewerToast(context, "Save failed: ${result.message}")
}
}
}
val onOpenExternal: () -> Unit = openExternal@{
if (currentBlurred || busy) return@openExternal
val target = attachment
val cached = target.cachedUri
if (!cached.isNullOrBlank()) {
runCatching {
MediaSaver.open(context, Uri.parse(cached), target.contentType)
}.onFailure {
viewerToast(context, "Couldn't open this image")
}
} else {
runWithBytes { item, bytes ->
val uri = MediaSaver.stageForShare(
context,
bytes,
item.fileName,
item.contentType,
)
MediaSaver.open(context, uri, item.contentType)
}
}
}
Box(
modifier = modifier
.fillMaxSize()
.background(Color.Black.copy(alpha = 0.96f)),
) {
HorizontalPager(
state = pagerState,
beyondViewportPageCount = 0,
pageSpacing = 12.dp,
modifier = Modifier
.fillMaxSize()
.testTag("attachment-gallery-pager"),
) { page ->
val pageAttachment = attachments[page]
val pageKey = galleryAttachmentKey(pageAttachment, page)
val blurred = revealed[pageKey] != true && shouldBlurImage(
blurMode,
pageAttachment.sensitive,
)
ImageBody(
attachment = pageAttachment,
blurred = blurred,
onReveal = { revealed[pageKey] = true },
)
}
MediaViewerToolbar(
title = toolbarTitle,
busy = busy,
actionsEnabled = !currentBlurred,
onShare = onShare,
onSave = onSave,
onOpenExternal = onOpenExternal,
onClose = onDismiss,
modifier = Modifier.align(Alignment.TopCenter),
)
Text(
text = "${currentIndex + 1} / ${attachments.size}",
style = MaterialTheme.typography.labelMedium,
color = Color.White,
modifier = Modifier
.align(Alignment.BottomCenter)
.windowInsetsPadding(WindowInsets.safeDrawing)
.padding(bottom = 12.dp)
.clip(RoundedCornerShape(50))
.background(Color.Black.copy(alpha = 0.55f))
.padding(horizontal = 12.dp, vertical = 6.dp),
)
}
}
}
/** The single shared control bar used across every attachment type. */
@Composable
private fun MediaViewerToolbar(
title: String,
busy: Boolean,
actionsEnabled: Boolean = true,
onShare: () -> Unit,
onSave: () -> Unit,
onOpenExternal: () -> Unit,
@@ -393,13 +643,17 @@ private fun MediaViewerToolbar(
modifier = Modifier.size(18.dp).padding(end = 4.dp),
)
}
IconButton(onClick = onOpenExternal, colors = tint) {
IconButton(
onClick = onOpenExternal,
enabled = actionsEnabled && !busy,
colors = tint,
) {
Icon(Icons.Filled.OpenInNew, contentDescription = "Open externally")
}
IconButton(onClick = onShare, colors = tint) {
IconButton(onClick = onShare, enabled = actionsEnabled && !busy, colors = tint) {
Icon(Icons.Filled.Share, contentDescription = "Share")
}
IconButton(onClick = onSave, colors = tint) {
IconButton(onClick = onSave, enabled = actionsEnabled && !busy, colors = tint) {
Icon(Icons.Filled.Download, contentDescription = "Save")
}
}
@@ -0,0 +1,224 @@
package com.hermesandroid.relay.ui.components
import androidx.compose.foundation.clickable
import androidx.compose.foundation.background
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.width
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.filled.Check
import androidx.compose.material.icons.filled.Close
import androidx.compose.material.icons.filled.ExpandLess
import androidx.compose.material.icons.filled.ExpandMore
import androidx.compose.material.icons.filled.HourglassTop
import androidx.compose.material3.Card
import androidx.compose.material3.CardDefaults
import androidx.compose.material3.HorizontalDivider
import androidx.compose.material3.Icon
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Text
import androidx.compose.runtime.Composable
import androidx.compose.runtime.LaunchedEffect
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.saveable.rememberSaveable
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.graphics.vector.ImageVector
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.text.style.TextOverflow
import androidx.compose.ui.unit.dp
import com.hermesandroid.relay.data.BackgroundTaskPhase
import com.hermesandroid.relay.data.BackgroundTaskState
import com.hermesandroid.relay.data.ToolCall
import com.hermesandroid.relay.ui.theme.relayMetadataStyle
/**
* The Chat-side identity for one promoted/durable Hermes run. It stays in the
* owning assistant turn while [BackgroundTaskState.phase] advances, rather
* than creating a running system notice and a second completion row.
*
* Tool activity is deliberately subordinate: the compact timeline expands
* inside this card and reuses [CompactToolCall]/[SubagentLane], so background
* work reads like the same task at every stage instead of a mini dashboard.
*/
@Composable
fun BackgroundTaskCard(
task: BackgroundTaskState,
toolCalls: List<ToolCall>,
showTimeline: Boolean,
modifier: Modifier = Modifier,
) {
val terminal = task.phase in terminalBackgroundTaskPhases
val timelineCalls = if (showTimeline) toolCalls else emptyList()
val hasTimeline = timelineCalls.isNotEmpty()
var expanded by rememberSaveable(task.id) { mutableStateOf(hasTimeline && !terminal) }
LaunchedEffect(terminal, hasTimeline) {
if (!hasTimeline || terminal) expanded = false
}
val phaseLabel = backgroundTaskPhaseLabel(task.phase)
val meta = backgroundTaskMeta(task, timelineCalls)
val icon: ImageVector
val iconTint = when (task.phase) {
BackgroundTaskPhase.COMPLETE -> {
icon = Icons.Filled.Check
MaterialTheme.colorScheme.primary
}
BackgroundTaskPhase.FAILED, BackgroundTaskPhase.CANCELLED -> {
icon = Icons.Filled.Close
MaterialTheme.colorScheme.error
}
else -> {
icon = Icons.Filled.HourglassTop
MaterialTheme.colorScheme.tertiary
}
}
Card(
modifier = modifier
.fillMaxWidth()
.semantics {
contentDescription = buildString {
append("Background task, ")
append(task.title)
append(", ")
append(phaseLabel.lowercase())
task.statusLine?.takeIf { it.isNotBlank() }?.let {
append(", ")
append(it)
}
if (meta.isNotBlank()) {
append(", ")
append(meta)
}
}
},
colors = CardDefaults.cardColors(
containerColor = MaterialTheme.colorScheme.surfaceVariant.copy(alpha = 0.58f),
),
) {
Column {
Row(
modifier = Modifier
.fillMaxWidth()
.clickable(enabled = hasTimeline) { expanded = !expanded }
.padding(horizontal = 12.dp, vertical = 10.dp),
verticalAlignment = Alignment.CenterVertically,
) {
Icon(
imageVector = icon,
contentDescription = null,
tint = iconTint,
modifier = Modifier.size(16.dp),
)
Spacer(modifier = Modifier.width(8.dp))
Column(modifier = Modifier.weight(1f)) {
Text(
text = task.title,
style = MaterialTheme.typography.labelMedium,
maxLines = 1,
overflow = TextOverflow.Ellipsis,
)
task.statusLine?.takeIf { it.isNotBlank() }?.let { status ->
Spacer(modifier = Modifier.height(2.dp))
Text(
text = status,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
maxLines = 2,
overflow = TextOverflow.Ellipsis,
)
}
}
Spacer(modifier = Modifier.width(8.dp))
Column(horizontalAlignment = Alignment.End) {
Text(
text = phaseLabel,
style = relayMetadataStyle(),
color = iconTint,
)
if (meta.isNotBlank()) {
Text(
text = meta,
style = relayMetadataStyle(),
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
if (hasTimeline) {
Spacer(modifier = Modifier.width(4.dp))
Icon(
imageVector = if (expanded) Icons.Filled.ExpandLess else Icons.Filled.ExpandMore,
contentDescription = if (expanded) "Collapse task timeline" else "Expand task timeline",
tint = MaterialTheme.colorScheme.onSurfaceVariant,
modifier = Modifier.size(16.dp),
)
}
}
if (!terminal) {
// A fixed accent rail communicates active state without adding
// another indeterminate animation to an already-live transcript.
Box(
modifier = Modifier
.fillMaxWidth()
.height(2.dp)
.background(MaterialTheme.colorScheme.tertiary.copy(alpha = 0.7f)),
)
}
if (expanded) {
HorizontalDivider(color = MaterialTheme.colorScheme.outlineVariant.copy(alpha = 0.55f))
Column(
modifier = Modifier.padding(horizontal = 10.dp, vertical = 8.dp),
verticalArrangement = Arrangement.spacedBy(4.dp),
) {
val lanes = timelineCalls.groupBy { it.taskIndex }
lanes[null].orEmpty().forEach { call ->
CompactToolCall(toolCall = call)
}
lanes.keys.filterNotNull().sorted().forEach { taskIndex ->
SubagentLane(
taskIndex = taskIndex,
calls = lanes.getValue(taskIndex),
)
}
}
}
}
}
}
internal fun backgroundTaskPhaseLabel(phase: BackgroundTaskPhase): String = when (phase) {
BackgroundTaskPhase.RUNNING -> "Working"
BackgroundTaskPhase.WAITING -> "Needs input"
BackgroundTaskPhase.DELIVERING -> "Delivering"
BackgroundTaskPhase.COMPLETE -> "Complete"
BackgroundTaskPhase.FAILED -> "Failed"
BackgroundTaskPhase.CANCELLED -> "Cancelled"
}
internal fun backgroundTaskMeta(task: BackgroundTaskState, toolCalls: List<ToolCall>): String {
val completed = maxOf(task.completedToolCount, toolCalls.count { it.isComplete })
return buildList {
if (completed > 0) add("$completed step${if (completed == 1) "" else "s"}")
if (task.queuedCount > 0) add("+${task.queuedCount} queued")
}.joinToString(" · ")
}
private val terminalBackgroundTaskPhases = setOf(
BackgroundTaskPhase.COMPLETE,
BackgroundTaskPhase.FAILED,
BackgroundTaskPhase.CANCELLED,
)
@@ -109,6 +109,12 @@ data class ThinkingIndicatorConfig(
/** Chat-root provided streaming-indicator config; see [ThinkingIndicatorConfig]. */
val LocalThinkingIndicator = compositionLocalOf { ThinkingIndicatorConfig() }
internal fun shouldAnimateDotMatrix(
appAnimationsEnabled: Boolean,
osAnimationsEnabled: Boolean,
touchExplorationEnabled: Boolean,
): Boolean = appAnimationsEnabled && osAnimationsEnabled && !touchExplorationEnabled
/**
* A compact dot-matrix "thinking" animation — a small grid of dots evoking a
* dot-matrix / LED display (the dot-anime-react concept reimplemented natively
@@ -142,10 +148,15 @@ fun DotMatrixIndicator(
fps: Int = 30,
animated: Boolean = true,
) {
val motion = rememberAccessibleMotionState()
val phase = rememberAmbientPhase(
periodMillis = pattern.periodMillis,
fps = fps,
running = animated,
running = shouldAnimateDotMatrix(
appAnimationsEnabled = animated,
osAnimationsEnabled = motion.osAnimations,
touchExplorationEnabled = motion.touchExploration,
),
)
val gridWidth = columnSpacing * (columns - 1)
val gridHeight = rowSpacing * (rows - 1)
@@ -0,0 +1,437 @@
package com.hermesandroid.relay.ui.components
import androidx.compose.animation.AnimatedVisibility
import androidx.compose.animation.animateContentSize
import androidx.compose.foundation.background
import androidx.compose.foundation.clickable
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.heightIn
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.lazy.LazyColumn
import androidx.compose.foundation.lazy.items
import androidx.compose.foundation.shape.CircleShape
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.text.selection.SelectionContainer
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.filled.CheckCircle
import androidx.compose.material.icons.filled.ErrorOutline
import androidx.compose.material.icons.filled.ExpandLess
import androidx.compose.material.icons.filled.ExpandMore
import androidx.compose.material.icons.filled.Refresh
import androidx.compose.material.icons.filled.Stop
import androidx.compose.material.icons.filled.Terminal
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.ExperimentalMaterial3Api
import androidx.compose.material3.HorizontalDivider
import androidx.compose.material3.Icon
import androidx.compose.material3.IconButton
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.ModalBottomSheet
import androidx.compose.material3.Surface
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.material3.rememberModalBottomSheetState
import androidx.compose.runtime.Composable
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.clip
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.semantics.stateDescription
import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.text.style.TextOverflow
import androidx.compose.ui.unit.dp
import com.hermesandroid.relay.network.upstream.GatewayProcess
/**
* Composer-adjacent summary of upstream Hermes processes for the active chat.
* The process registry is session-scoped, so this deliberately does not live
* in the global session/navigation drawer.
*/
@Composable
fun GatewayBackgroundProcessStrip(
processes: List<GatewayProcess>,
loading: Boolean,
onClick: () -> Unit,
modifier: Modifier = Modifier,
) {
// Initial/switch refreshes are silent. The strip appears only after the
// session actually owns a process, avoiding a transient "Checking" row on
// every ordinary chat open.
if (processes.isEmpty()) return
val running = processes.count { it.isRunning }
val failed = processes.count { !it.isRunning && (it.exitCode ?: 0) != 0 }
val displayedCount = if (running > 0) running else processes.size
val status = when {
running > 0 -> "$running running"
failed > 0 -> "$failed failed"
else -> "Complete"
}
Surface(
modifier = modifier
.fillMaxWidth()
.padding(horizontal = 16.dp, vertical = 3.dp)
.heightIn(min = 48.dp)
.semantics {
contentDescription =
"Background processes, $status. Open current chat activity."
stateDescription = status
}
.clickable(
onClickLabel = "Open background processes",
onClick = onClick,
),
shape = RoundedCornerShape(14.dp),
color = MaterialTheme.colorScheme.surfaceVariant.copy(alpha = 0.72f),
tonalElevation = 1.dp,
) {
Row(
modifier = Modifier.padding(horizontal = 12.dp, vertical = 9.dp),
verticalAlignment = Alignment.CenterVertically,
) {
if (running > 0 || loading) {
CircularProgressIndicator(modifier = Modifier.size(16.dp), strokeWidth = 2.dp)
} else {
Icon(
imageVector = if (failed > 0) Icons.Filled.ErrorOutline else Icons.Filled.CheckCircle,
contentDescription = null,
modifier = Modifier.size(17.dp),
tint = if (failed > 0) {
MaterialTheme.colorScheme.error
} else {
MaterialTheme.colorScheme.primary
},
)
}
Spacer(Modifier.width(9.dp))
Text(
text = "Background · $displayedCount",
style = MaterialTheme.typography.labelLarge,
modifier = Modifier.weight(1f),
)
Text(
text = status,
style = MaterialTheme.typography.labelMedium,
color = if (failed > 0 && running == 0) {
MaterialTheme.colorScheme.error
} else {
MaterialTheme.colorScheme.onSurfaceVariant
},
)
Icon(
imageVector = Icons.Filled.ExpandLess,
contentDescription = null,
modifier = Modifier
.padding(start = 6.dp)
.size(17.dp),
tint = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
}
/** Mobile analogue of Hermes Desktop's composer process stack + terminal viewer. */
@OptIn(ExperimentalMaterial3Api::class)
@Composable
fun GatewayBackgroundProcessSheet(
processes: List<GatewayProcess>,
loading: Boolean,
stoppingProcessIds: Set<String>,
onRefresh: () -> Unit,
onStop: (String) -> Unit,
onDismissProcess: (String) -> Unit,
onDismiss: () -> Unit,
) {
val sheetState = rememberModalBottomSheetState(skipPartiallyExpanded = false)
val running = processes.filter { it.isRunning }
val recent = processes.filterNot { it.isRunning }
ModalBottomSheet(
onDismissRequest = onDismiss,
sheetState = sheetState,
) {
Column(
modifier = Modifier
.fillMaxWidth()
.padding(bottom = 24.dp),
) {
Row(
modifier = Modifier
.fillMaxWidth()
.padding(start = 20.dp, end = 8.dp, bottom = 4.dp),
verticalAlignment = Alignment.CenterVertically,
) {
Column(modifier = Modifier.weight(1f)) {
Text("Background processes", style = MaterialTheme.typography.titleLarge)
Text(
"Current chat · live output and recent results",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
IconButton(onClick = onRefresh, enabled = !loading) {
if (loading) {
CircularProgressIndicator(modifier = Modifier.size(20.dp), strokeWidth = 2.dp)
} else {
Icon(Icons.Filled.Refresh, contentDescription = "Refresh processes")
}
}
}
if (processes.isEmpty() && !loading) {
Column(
modifier = Modifier
.fillMaxWidth()
.padding(horizontal = 24.dp, vertical = 36.dp),
horizontalAlignment = Alignment.CenterHorizontally,
) {
Icon(
Icons.Filled.Terminal,
contentDescription = null,
modifier = Modifier.size(30.dp),
tint = MaterialTheme.colorScheme.onSurfaceVariant,
)
Text(
"No background processes in this chat",
modifier = Modifier.padding(top = 12.dp),
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
} else {
LazyColumn(
modifier = Modifier
.fillMaxWidth()
.heightIn(max = 560.dp),
) {
if (running.isNotEmpty()) {
item { ProcessSectionLabel("Running", running.size) }
items(running, key = { it.id }) { process ->
GatewayProcessRow(
process = process,
stopping = process.id in stoppingProcessIds,
onStop = { onStop(process.id) },
onDismiss = null,
)
}
}
if (running.isNotEmpty() && recent.isNotEmpty()) {
item { HorizontalDivider(modifier = Modifier.padding(vertical = 6.dp)) }
}
if (recent.isNotEmpty()) {
item { ProcessSectionLabel("Recent", recent.size) }
items(recent, key = { it.id }) { process ->
GatewayProcessRow(
process = process,
stopping = false,
onStop = null,
onDismiss = { onDismissProcess(process.id) },
)
}
}
}
}
}
}
}
@Composable
private fun ProcessSectionLabel(label: String, count: Int) {
Text(
text = "$label · $count",
modifier = Modifier.padding(horizontal = 20.dp, vertical = 8.dp),
style = MaterialTheme.typography.labelMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
@Composable
private fun GatewayProcessRow(
process: GatewayProcess,
stopping: Boolean,
onStop: (() -> Unit)?,
onDismiss: (() -> Unit)?,
) {
var expanded by remember(process.id) { mutableStateOf(false) }
val failed = !process.isRunning && (process.exitCode ?: 0) != 0
val output = sanitizeTerminalText(
process.outputTail.orEmpty().ifBlank { process.outputPreview.orEmpty() },
).trimEnd()
val command = sanitizeTerminalText(process.command)
.lineSequence()
.firstOrNull()
?.trim()
.orEmpty()
.ifBlank {
"Background process"
}
Column(
modifier = Modifier
.fillMaxWidth()
.animateContentSize()
.clickable(
enabled = output.isNotBlank(),
onClickLabel = if (expanded) "Collapse process output" else "Expand process output",
) { expanded = !expanded }
.padding(horizontal = 20.dp, vertical = 10.dp),
) {
Row(verticalAlignment = Alignment.CenterVertically) {
ProcessStateIcon(process = process, failed = failed, stopping = stopping)
Column(
modifier = Modifier
.weight(1f)
.padding(horizontal = 10.dp),
) {
Text(
text = command,
style = MaterialTheme.typography.bodyMedium,
maxLines = 2,
overflow = TextOverflow.Ellipsis,
)
Text(
text = processMetadata(process, failed),
style = MaterialTheme.typography.labelSmall,
color = if (failed) {
MaterialTheme.colorScheme.error
} else {
MaterialTheme.colorScheme.onSurfaceVariant
},
maxLines = 1,
overflow = TextOverflow.Ellipsis,
)
}
if (onStop != null) {
TextButton(onClick = onStop, enabled = !stopping) {
Icon(
Icons.Filled.Stop,
contentDescription = null,
modifier = Modifier.size(16.dp),
)
Spacer(Modifier.width(4.dp))
Text(if (stopping) "Stopping" else "Stop")
}
} else if (onDismiss != null) {
TextButton(onClick = onDismiss) { Text("Dismiss") }
}
if (output.isNotBlank()) {
Icon(
imageVector = if (expanded) Icons.Filled.ExpandLess else Icons.Filled.ExpandMore,
contentDescription = if (expanded) "Collapse output" else "Expand output",
modifier = Modifier.size(20.dp),
tint = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
AnimatedVisibility(visible = expanded && output.isNotBlank()) {
Box(
modifier = Modifier
.fillMaxWidth()
.padding(top = 10.dp)
.clip(RoundedCornerShape(10.dp))
.background(MaterialTheme.colorScheme.surfaceContainerHighest)
.padding(12.dp),
) {
SelectionContainer {
Text(
text = output,
style = MaterialTheme.typography.bodySmall.copy(fontFamily = FontFamily.Monospace),
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
}
}
}
@Composable
private fun ProcessStateIcon(process: GatewayProcess, failed: Boolean, stopping: Boolean) {
val tint: Color = when {
failed -> MaterialTheme.colorScheme.error
process.isRunning || stopping -> MaterialTheme.colorScheme.primary
else -> MaterialTheme.colorScheme.tertiary
}
Surface(
modifier = Modifier.size(30.dp),
shape = CircleShape,
color = tint.copy(alpha = 0.12f),
) {
Box(contentAlignment = Alignment.Center) {
when {
process.isRunning || stopping -> CircularProgressIndicator(
modifier = Modifier.size(17.dp),
strokeWidth = 2.dp,
color = tint,
)
failed -> Icon(
Icons.Filled.ErrorOutline,
contentDescription = "Failed",
modifier = Modifier.size(18.dp),
tint = tint,
)
else -> Icon(
Icons.Filled.CheckCircle,
contentDescription = "Completed",
modifier = Modifier.size(18.dp),
tint = tint,
)
}
}
}
}
private fun processMetadata(process: GatewayProcess, failed: Boolean): String {
val state = when {
process.isRunning -> "Running"
failed -> "Failed${process.exitCode?.let { " · exit $it" }.orEmpty()}"
else -> "Completed${process.exitCode?.let { " · exit $it" }.orEmpty()}"
}
return "$state · ${formatElapsed(process.uptimeSeconds)}" +
if (process.detached) " · recovered" else ""
}
private val ansiTerminalEscape = Regex(
"\u001B(?:\\].*?(?:\u0007|\u001B\\\\)|\\[[0-?]*[ -/]*[@-~]|[ -/]*[@-~])",
RegexOption.DOT_MATCHES_ALL,
)
private val unterminatedOsc = Regex("\u001B\\][^\\n]*")
/** Plain-text mobile output viewer: remove terminal control/ANSI while keeping layout text. */
internal fun sanitizeTerminalText(raw: String): String {
val withoutAnsi = unterminatedOsc.replace(
ansiTerminalEscape.replace(raw, ""),
"",
)
return withoutAnsi
.replace("\r\n", "\n")
.replace('\r', '\n')
.filter { char ->
char == '\n' || char == '\t' || (char.code >= 0x20 && char.code != 0x7F)
}
}
internal fun formatElapsed(totalSeconds: Long): String {
val seconds = totalSeconds.coerceAtLeast(0)
val hours = seconds / 3_600
val minutes = (seconds % 3_600) / 60
val remainder = seconds % 60
return when {
hours > 0 -> "${hours}h ${minutes}m"
minutes > 0 -> "${minutes}m ${remainder}s"
else -> "${remainder}s"
}
}
@@ -491,7 +491,7 @@ private fun FileCardRender(
* the default tap previews in-app.
*/
@Composable
private fun AttachmentActionsMenu(
fun AttachmentActionsMenu(
expanded: Boolean,
onDismiss: () -> Unit,
context: Context,
@@ -527,7 +527,7 @@ private fun AttachmentActionsMenu(
/** Small circular download button overlaid on inline images (B2). */
@Composable
private fun SaveOverlayButton(onClick: () -> Unit, modifier: Modifier = Modifier) {
fun SaveOverlayButton(onClick: () -> Unit, modifier: Modifier = Modifier) {
Surface(
shape = CircleShape,
color = Color.Black.copy(alpha = 0.45f),
@@ -556,7 +556,7 @@ private suspend fun shareAttachment(context: Context, attachment: Attachment) {
MediaSaver.share(context, uri, attachment.contentType)
}
private suspend fun saveAttachment(context: Context, attachment: Attachment) {
suspend fun saveAttachment(context: Context, attachment: Attachment) {
val bytes = attachmentBytes(context, attachment)
if (bytes == null) {
attachmentToast(context, "Couldn't read this file")
@@ -1,13 +1,18 @@
package com.hermesandroid.relay.ui.components
import androidx.compose.foundation.background
import androidx.compose.foundation.horizontalScroll
import com.hermesandroid.relay.ui.theme.LocalBrand
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.BoxWithConstraints
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.fillMaxHeight
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.requiredWidth
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.material.icons.Icons
@@ -26,8 +31,13 @@ import androidx.compose.runtime.remember
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.clip
import androidx.compose.ui.graphics.Brush
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.platform.LocalClipboardManager
import androidx.compose.ui.semantics.CollectionInfo
import androidx.compose.ui.semantics.collectionInfo
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.text.AnnotatedString
import androidx.compose.ui.text.SpanStyle
import androidx.compose.ui.text.TextLinkStyles
@@ -38,18 +48,33 @@ import androidx.compose.ui.text.style.TextDecoration
import androidx.compose.ui.text.style.TextOverflow
import androidx.compose.ui.unit.dp
import androidx.compose.ui.unit.sp
import kotlinx.coroutines.delay
import androidx.compose.ui.unit.times
import com.mikepenz.markdown.compose.components.markdownComponents
import com.mikepenz.markdown.compose.components.MarkdownComponentModel
import com.mikepenz.markdown.compose.LocalMarkdownColors
import com.mikepenz.markdown.compose.LocalMarkdownDimens
import com.mikepenz.markdown.compose.elements.MarkdownDivider
import com.mikepenz.markdown.compose.elements.MarkdownHighlightedCodeBlock
import com.mikepenz.markdown.compose.elements.MarkdownHighlightedCodeFence
import com.mikepenz.markdown.compose.elements.MarkdownTable
import com.mikepenz.markdown.compose.elements.MarkdownTableHeader
import com.mikepenz.markdown.compose.elements.MarkdownTableRow
import com.mikepenz.markdown.compose.extendedspans.ExtendedSpans
import com.mikepenz.markdown.compose.extendedspans.RoundedCornerSpanPainter
import com.mikepenz.markdown.m3.Markdown
import com.mikepenz.markdown.m3.markdownColor
import com.mikepenz.markdown.m3.markdownTypography
import com.mikepenz.markdown.model.markdownDimens
import com.mikepenz.markdown.model.markdownExtendedSpans
import com.hermesandroid.relay.ui.theme.LocalBrand
import dev.snipme.highlights.Highlights
import dev.snipme.highlights.model.SyntaxThemes
import kotlinx.coroutines.delay
import org.intellij.markdown.ast.findChildOfType
import org.intellij.markdown.flavours.gfm.GFMElementTypes.HEADER
import org.intellij.markdown.flavours.gfm.GFMElementTypes.ROW
import org.intellij.markdown.flavours.gfm.GFMTokenTypes.CELL
import org.intellij.markdown.flavours.gfm.GFMTokenTypes.TABLE_SEPARATOR
@Composable
fun MarkdownContent(
@@ -61,7 +86,6 @@ fun MarkdownContent(
val highlightsBuilder = remember(isDarkTheme) {
Highlights.Builder().theme(SyntaxThemes.atom(darkMode = isDarkTheme))
}
Markdown(
content = content,
modifier = modifier,
@@ -133,6 +157,13 @@ fun MarkdownContent(
),
),
),
// Tables get a phone-friendly minimum measure. The stock renderer uses
// one-line cells; our table component below keeps the same AST/inline
// annotator path but permits wrapping and exposes horizontal overflow.
dimens = markdownDimens(
tableCellWidth = 110.dp,
tableCellPadding = 12.dp,
),
components = markdownComponents(
codeBlock = {
MarkdownHighlightedCodeBlock(
@@ -149,7 +180,8 @@ fun MarkdownContent(
highlightsBuilder = highlightsBuilder,
showHeader = true
)
}
},
table = { WideMarkdownTable(it) },
),
extendedSpans = markdownExtendedSpans {
remember { ExtendedSpans(RoundedCornerSpanPainter()) }
@@ -157,20 +189,141 @@ fun MarkdownContent(
)
}
/**
* GFM table renderer tuned for a narrow chat bubble.
*
* Every column keeps the configured 110dp minimum and cells wrap instead of
* truncating to one line. Tables wider than the bubble scroll horizontally;
* the trailing fade is deliberately subtle and disappears once the reader has
* reached the final column.
*/
@Composable
private fun WideMarkdownTable(model: MarkdownComponentModel) {
val columnsCount = remember(model.node) {
model.node.findChildOfType(HEADER)?.children?.count { it.type == CELL } ?: 0
}
if (columnsCount == 0) {
MarkdownTable(
content = model.content,
node = model.node,
style = model.typography.table,
)
return
}
val rowsCount = remember(model.node) {
model.node.children.count { it.type == ROW } + 1
}
val tableCellWidth = LocalMarkdownDimens.current.tableCellWidth
val tableWidth = columnsCount * tableCellWidth
val tableCornerSize = LocalMarkdownDimens.current.tableCornerSize
val tableBackground = LocalMarkdownColors.current.tableBackground
val scrollState = rememberScrollState()
BoxWithConstraints(
modifier = Modifier
.fillMaxWidth()
.clip(RoundedCornerShape(tableCornerSize))
.background(tableBackground)
.semantics {
collectionInfo = CollectionInfo(
rowCount = rowsCount,
columnCount = columnsCount,
)
},
) {
val scrollable = maxWidth < tableWidth
Box {
Column(
modifier = if (scrollable) {
Modifier
.horizontalScroll(scrollState)
.requiredWidth(tableWidth)
} else {
Modifier.fillMaxWidth()
},
) {
var rowIndex = 1
model.node.children.forEach { child ->
when (child.type) {
HEADER -> MarkdownTableHeader(
content = model.content,
header = child,
tableWidth = tableWidth,
style = model.typography.table,
verticalAlignment = Alignment.Top,
maxLines = Int.MAX_VALUE,
overflow = TextOverflow.Clip,
)
ROW -> {
MarkdownTableRow(
content = model.content,
header = child,
tableWidth = tableWidth,
style = model.typography.table,
rowIndex = rowIndex,
verticalAlignment = Alignment.Top,
maxLines = Int.MAX_VALUE,
overflow = TextOverflow.Clip,
)
rowIndex++
}
TABLE_SEPARATOR -> MarkdownDivider()
}
}
}
if (scrollable && scrollState.canScrollForward) {
Box(
modifier = Modifier.matchParentSize(),
) {
Box(
modifier = Modifier
.align(Alignment.CenterEnd)
.fillMaxHeight()
.width(28.dp)
.background(
Brush.horizontalGradient(
colors = listOf(Color.Transparent, tableBackground),
),
),
)
}
}
}
}
}
@Composable
fun StreamingMarkdownContent(
content: String,
textColor: Color,
modifier: Modifier = Modifier
isStreaming: Boolean = true,
modifier: Modifier = Modifier,
) {
val blocks = remember(content) { parseStreamingMarkdownBlocks(content) }
val blocks = remember(content, isStreaming) {
if (isStreaming) {
parseStreamingMarkdownBlocks(content)
} else {
listOf(StreamingMarkdownBlock.Markdown(content))
}
}
Column(
modifier = modifier,
verticalArrangement = Arrangement.spacedBy(6.dp),
// Match the final renderer's block spacer so moving the active tail
// into the settled Markdown prefix does not add a second layout jump.
verticalArrangement = Arrangement.spacedBy(2.dp),
) {
blocks.forEach { block ->
when (block) {
is StreamingMarkdownBlock.Markdown -> MarkdownContent(
content = block.content,
textColor = textColor,
)
is StreamingMarkdownBlock.Text -> Text(
text = block.text,
style = MaterialTheme.typography.bodyMedium,
@@ -267,23 +420,96 @@ private fun CodeCopyButton(code: String) {
}
}
private sealed interface StreamingMarkdownBlock {
internal sealed interface StreamingMarkdownBlock {
data class Markdown(val content: String) : StreamingMarkdownBlock
data class Text(val text: String) : StreamingMarkdownBlock
data class Code(val language: String, val code: String) : StreamingMarkdownBlock
}
private fun parseStreamingMarkdownBlocks(content: String): List<StreamingMarkdownBlock> {
/**
* Splits an in-flight response into stable Markdown and one structurally
* incomplete tail.
*
* Only conservative, blank-terminated top-level prose/heading blocks promote
* to the real renderer. Lists, quotes, tables, indented blocks, HTML, and code
* remain on the lightweight streaming surface until the message settles; those
* containers can legally absorb later lines, so promoting them early causes a
* visible re-parenting jump when the final CommonMark tree is parsed.
*/
internal fun parseStreamingMarkdownBlocks(content: String): List<StreamingMarkdownBlock> {
if (content.isBlank()) return emptyList()
val normalized = content
.replace("\r\n", "\n")
.replace('\r', '\n')
val blocks = mutableListOf<StreamingMarkdownBlock>()
val settledEnd = findStableMarkdownBoundary(normalized).coerceAtLeast(0)
if (settledEnd > 0) {
normalized.substring(0, settledEnd).trimEnd().let { stable ->
if (stable.isNotBlank()) blocks += StreamingMarkdownBlock.Markdown(stable)
}
}
val activeTail = normalized.substring(settledEnd)
if (activeTail.isNotBlank()) {
blocks += parseActiveStreamingTail(activeTail)
}
return blocks
}
/** Last unambiguous blank-line boundary in a contiguous simple-markdown prefix. */
private fun findStableMarkdownBoundary(content: String): Int {
var offset = 0
var blockStart = 0
var activeFence: StreamingFence? = null
var lastStableBoundary = 0
while (offset < content.length) {
val newline = content.indexOf('\n', offset)
val lineEnd = if (newline >= 0) newline else content.length
val line = content.substring(offset, lineEnd)
activeFence = when (val current = activeFence) {
null -> streamingFence(line)
else -> if (isClosingFence(line, current)) null else current
}
// Whitespace-only lines can be meaningful indentation inside a list.
// Require a truly empty delimiter and stop at the first ambiguous
// container so every promoted prefix remains structurally final.
if (activeFence == null && newline >= 0 && line.isEmpty()) {
val candidate = content.substring(blockStart, offset).trimEnd()
if (candidate.isNotBlank() && !isConservativeStableBlock(candidate)) break
lastStableBoundary = newline + 1
blockStart = lastStableBoundary
}
if (newline < 0) break
offset = newline + 1
}
return lastStableBoundary
}
private fun isConservativeStableBlock(block: String): Boolean = block
.lineSequence()
.filter { it.isNotEmpty() }
.none { line ->
line.firstOrNull()?.isWhitespace() == true ||
AMBIGUOUS_STREAMING_BLOCK.matches(line)
}
private fun parseActiveStreamingTail(content: String): List<StreamingMarkdownBlock> {
val blocks = mutableListOf<StreamingMarkdownBlock>()
val paragraph = StringBuilder()
val code = StringBuilder()
var inFence = false
var activeFence = ""
var activeFence: StreamingFence? = null
var language = ""
fun flushParagraph() {
val text = paragraph.toString().trimEnd()
val text = paragraph.toString().trim('\n').trimEnd()
if (text.isNotBlank()) {
blocks += StreamingMarkdownBlock.Text(text)
}
@@ -298,35 +524,30 @@ private fun parseStreamingMarkdownBlocks(content: String): List<StreamingMarkdow
code.clear()
}
val lines = content
.replace("\r\n", "\n")
.replace('\r', '\n')
.split('\n')
val lines = content.split('\n')
lines.forEachIndexed { index, line ->
val lineWithBreak = if (index == lines.lastIndex) line else "$line\n"
if (!inFence) {
val fence = streamingFenceMarker(line)
if (activeFence == null) {
val fence = streamingFence(line)
if (fence != null) {
flushParagraph()
inFence = true
activeFence = fence
language = streamingFenceLanguage(line, fence)
} else {
paragraph.append(lineWithBreak)
}
} else if (streamingFenceMarker(line) == activeFence) {
} else if (isClosingFence(line, activeFence)) {
flushCode()
inFence = false
activeFence = ""
activeFence = null
language = ""
} else {
code.append(lineWithBreak)
}
}
if (inFence) {
if (activeFence != null) {
flushCode()
} else {
flushParagraph()
@@ -335,18 +556,32 @@ private fun parseStreamingMarkdownBlocks(content: String): List<StreamingMarkdow
return blocks
}
private fun streamingFenceMarker(line: String): String? {
private data class StreamingFence(
val marker: Char,
val length: Int,
)
private fun streamingFence(line: String): StreamingFence? {
val trimmed = line.trimStart()
return when {
trimmed.startsWith("```") -> "```"
trimmed.startsWith("~~~") -> "~~~"
else -> null
}
val marker = trimmed.firstOrNull()?.takeIf { it == '`' || it == '~' } ?: return null
val length = trimmed.takeWhile { it == marker }.length
return length.takeIf { it >= 3 }?.let { StreamingFence(marker, it) }
}
private fun streamingFenceLanguage(line: String, marker: String): String {
val tail = line.trimStart().removePrefix(marker).trim()
private fun isClosingFence(line: String, activeFence: StreamingFence): Boolean {
val trimmed = line.trimStart()
if (trimmed.firstOrNull() != activeFence.marker) return false
val markerLength = trimmed.takeWhile { it == activeFence.marker }.length
return markerLength >= activeFence.length && trimmed.drop(markerLength).isBlank()
}
private fun streamingFenceLanguage(line: String, fence: StreamingFence): String {
val tail = line.trimStart().drop(fence.length).trim()
return tail
.takeWhile { !it.isWhitespace() && it != '`' && it != '~' }
.takeWhile { !it.isWhitespace() && it != fence.marker }
.take(32)
}
private val AMBIGUOUS_STREAMING_BLOCK = Regex(
"""^(?:[-+*]\s|\d{1,9}[.)]\s|>|```|~~~|<|\||(?:-{3,}|={3,})\s*$).*""",
)
@@ -395,21 +395,16 @@ fun MessageBubble(
color = textColor
)
} else {
// Use a stable lightweight renderer while streaming.
// The full parser/highlighter can rebuild block shapes
// on every partial fence/token and visibly flicker.
// Settled blocks keep the real Markdown renderer while
// only the structurally incomplete tail stays raw. The
// same composable settles the final tail so retained
// parser state survives the streaming -> final handoff.
if (markdownBody.isNotEmpty()) {
if (message.isStreaming) {
StreamingMarkdownContent(
content = markdownBody,
textColor = textColor
)
} else {
MarkdownContent(
content = markdownBody,
textColor = textColor
)
}
StreamingMarkdownContent(
content = markdownBody,
textColor = textColor,
isStreaming = message.isStreaming,
)
}
}
}
@@ -453,22 +448,34 @@ fun MessageBubble(
}
}
// Attachments — dispatched through the unified InboundAttachmentCard
// so outbound and inbound attachments share the same render pipeline.
// Outbound attachments (user-authored) always have state=LOADED so
// they route straight to the LOADED branch; inbound attachments (via
// MEDIA markers) cycle through LOADING → LOADED / FAILED as the
// background fetch progresses.
// Attachments — two or more loaded images collapse into one
// grid + swipe-across gallery. Every other item stays on the
// unified InboundAttachmentCard path, and layout items retain
// their original ChatMessage.attachments indices so retry /
// manual-fetch callbacks cannot drift after grouping.
if (message.attachments.isNotEmpty()) {
Spacer(modifier = Modifier.height(4.dp))
message.attachments.forEachIndexed { index, attachment ->
InboundAttachmentCard(
attachment = attachment,
onRetry = { onAttachmentRetry(message.id, index) },
onManualFetch = { onAttachmentManualFetch(message.id, index) },
maxWidth = maxBubbleWidth - 24.dp,
modifier = Modifier.padding(vertical = 2.dp)
)
val attachmentItems = remember(message.attachments) {
attachmentLayoutItems(message.attachments)
}
attachmentItems.forEach { item ->
when (item) {
is AttachmentLayoutItem.Gallery -> AttachmentGallery(
attachments = item.attachmentIndices.map(message.attachments::get),
maxWidth = maxBubbleWidth - 24.dp,
modifier = Modifier.padding(vertical = 2.dp),
)
is AttachmentLayoutItem.Single -> {
val index = item.attachmentIndex
InboundAttachmentCard(
attachment = message.attachments[index],
onRetry = { onAttachmentRetry(message.id, index) },
onManualFetch = { onAttachmentManualFetch(message.id, index) },
maxWidth = maxBubbleWidth - 24.dp,
modifier = Modifier.padding(vertical = 2.dp),
)
}
}
}
}
@@ -0,0 +1,146 @@
package com.hermesandroid.relay.ui.components
import androidx.compose.animation.AnimatedVisibility
import androidx.compose.foundation.background
import androidx.compose.foundation.clickable
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.heightIn
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.layout.widthIn
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.text.selection.SelectionContainer
import androidx.compose.foundation.verticalScroll
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.filled.Code
import androidx.compose.material.icons.filled.ExpandLess
import androidx.compose.material.icons.filled.ExpandMore
import androidx.compose.material3.Icon
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Text
import androidx.compose.runtime.Composable
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.saveable.rememberSaveable
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.clip
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.semantics.stateDescription
import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.unit.dp
import com.hermesandroid.relay.data.HermesProcessNotification
import com.hermesandroid.relay.ui.theme.relayMetadataStyle
/**
* Compact transcript treatment for the user-role process events that upstream
* Hermes injects when background work completes or matches a watch pattern.
*
* This component is intentionally quieter than a user bubble: the row is
* machine-authored process state, while the following assistant message is the
* conversational response. Command and output detail remain selectable behind
* progressive disclosure.
*/
@Composable
fun SyntheticProcessNotificationNotice(
notification: HermesProcessNotification,
modifier: Modifier = Modifier,
) {
val detail = notification.detail?.let(::sanitizeTerminalText)
val hasDetail = !detail.isNullOrBlank()
var expanded by rememberSaveable(notification.processId, notification.headline) {
mutableStateOf(false)
}
Column(
modifier = modifier.fillMaxWidth(),
horizontalAlignment = Alignment.CenterHorizontally,
) {
Column(
modifier = Modifier
.widthIn(max = 560.dp)
.padding(horizontal = 8.dp, vertical = 2.dp),
) {
Row(
modifier = Modifier
.fillMaxWidth()
.heightIn(min = 48.dp)
.clickable(
enabled = hasDetail,
onClickLabel = if (expanded) "Collapse process output" else "Expand process output",
) { expanded = !expanded }
.semantics {
contentDescription = buildString {
append("Background process notice. ")
append(notification.headline)
if (hasDetail) {
append(if (expanded) ". Output expanded" else ". Output collapsed")
}
}
if (hasDetail) {
stateDescription = if (expanded) "Output expanded" else "Output collapsed"
}
}
.padding(vertical = 4.dp),
verticalAlignment = Alignment.CenterVertically,
) {
Icon(
imageVector = Icons.Filled.Code,
contentDescription = null,
tint = MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.62f),
modifier = Modifier.size(14.dp),
)
Spacer(modifier = Modifier.width(6.dp))
Text(
text = notification.headline,
style = relayMetadataStyle(),
color = MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.72f),
modifier = Modifier.weight(1f),
)
if (hasDetail) {
Spacer(modifier = Modifier.width(4.dp))
Text(
text = "output",
style = relayMetadataStyle(),
color = MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.58f),
)
Icon(
imageVector = if (expanded) Icons.Filled.ExpandLess else Icons.Filled.ExpandMore,
contentDescription = null,
tint = MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.58f),
modifier = Modifier.size(16.dp),
)
}
}
AnimatedVisibility(visible = expanded && hasDetail) {
SelectionContainer {
Box(
modifier = Modifier
.fillMaxWidth()
.heightIn(max = 192.dp)
.clip(RoundedCornerShape(8.dp))
.background(MaterialTheme.colorScheme.surfaceVariant.copy(alpha = 0.5f))
.verticalScroll(rememberScrollState())
.padding(horizontal = 10.dp, vertical = 8.dp),
) {
Text(
text = detail.orEmpty(),
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.8f),
fontFamily = FontFamily.Monospace,
)
}
}
}
}
}
}
@@ -50,6 +50,8 @@ import androidx.compose.material.icons.filled.Share
import androidx.compose.material.icons.filled.Tune
import androidx.compose.material3.AssistChip
import androidx.compose.material3.AssistChipDefaults
import androidx.compose.material3.Badge
import androidx.compose.material3.BadgedBox
import androidx.compose.material3.Button
import androidx.compose.material3.CardDefaults
import androidx.compose.material3.CircularProgressIndicator
@@ -99,6 +101,9 @@ import androidx.compose.ui.platform.LocalClipboard
import androidx.compose.ui.platform.LocalFocusManager
import androidx.compose.ui.platform.LocalHapticFeedback
import androidx.compose.ui.res.painterResource
import androidx.compose.ui.semantics.clearAndSetSemantics
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.unit.Dp
import androidx.compose.ui.unit.dp
import androidx.compose.ui.zIndex
@@ -145,7 +150,9 @@ import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.Connection
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.data.hermesProcessNotificationOrNull
import com.hermesandroid.relay.ui.components.AgentInfoSheet
import com.hermesandroid.relay.ui.components.BackgroundTaskCard
import com.hermesandroid.relay.ui.components.LocalRelayServerImageResolver
import com.hermesandroid.relay.ui.components.RelayServerImageResolver
import com.hermesandroid.relay.ui.components.ChatInputBar
@@ -160,10 +167,13 @@ import com.hermesandroid.relay.ui.components.ConnectionStatusBadge
import com.hermesandroid.relay.ui.components.CommandRow
import com.hermesandroid.relay.ui.components.CompactToolCall
import com.hermesandroid.relay.ui.components.ContextMeterBar
import com.hermesandroid.relay.ui.components.GatewayBackgroundProcessSheet
import com.hermesandroid.relay.ui.components.GatewayBackgroundProcessStrip
import com.hermesandroid.relay.ui.components.InjectedContextSheet
import com.hermesandroid.relay.ui.components.InlineAutocomplete
import com.hermesandroid.relay.ui.components.loadedContentTransform
import com.hermesandroid.relay.ui.components.MessageBubble
import com.hermesandroid.relay.ui.components.SyntheticProcessNotificationNotice
import com.hermesandroid.relay.ui.components.avatar.AvatarRenderState
import androidx.compose.ui.layout.ContentScale
import coil3.compose.AsyncImage
@@ -406,6 +416,7 @@ fun ChatScreen(
onNavigateToProfileInspector: (String) -> Unit = {},
) {
val voiceUiState by voiceViewModel.uiState.collectAsState()
val isDemoMode by connectionViewModel.isDemoMode.collectAsState()
var voiceCompactMode by remember { mutableStateOf(false) }
val chatAlpha by animateFloatAsState(
targetValue = if (voiceUiState.voiceMode && !voiceCompactMode) 0.4f else 1f,
@@ -436,9 +447,11 @@ fun ChatScreen(
) { granted ->
if (granted) {
micPermissionDenied = false
if (pendingVoiceEnter) {
if (pendingVoiceEnter && !isDemoMode) {
pendingVoiceEnter = false
voiceViewModel.enterVoiceMode()
} else {
pendingVoiceEnter = false
}
} else {
pendingVoiceEnter = false
@@ -490,6 +503,9 @@ fun ChatScreen(
val sessions by chatViewModel.sessions.collectAsState()
val serverAutoTitles by chatViewModel.serverAutoTitles.collectAsState()
val currentSessionId by chatViewModel.currentSessionId.collectAsState()
val backgroundProcesses by chatViewModel.backgroundProcesses.collectAsState()
val backgroundProcessesLoading by chatViewModel.backgroundProcessesLoading.collectAsState()
val stoppingProcessIds by chatViewModel.stoppingProcessIds.collectAsState()
val isLoadingHistory by chatViewModel.isLoadingHistory.collectAsState()
val isLoadingSessions by chatViewModel.isLoadingSessions.collectAsState()
val selectedPersonality by chatViewModel.selectedPersonality.collectAsState()
@@ -554,15 +570,15 @@ fun ChatScreen(
connectionViewModel.resolveStreamingEndpoint(streamingEndpointPref) == "gateway"
}
// Pre-warm the gateway (connect + resume the current session) whenever the
// chat surface is visible, the app is foregrounded, and the gateway is the
// resolved transport — so the first send is warm (tens of ms to first
// token) instead of paying the cold connect + session.resume on the send
// path. Best-effort / idempotent; re-fires on return-to-foreground.
// Recover any durable in-flight chat checkpoint whenever Chat returns to
// the foreground. On Gateway this also pre-warms/re-attaches the socket;
// sessions-SSE falls back to bounded persisted-history reconciliation.
val appForeground by com.hermesandroid.relay.util.AppForegroundTracker.isForeground.collectAsState()
LaunchedEffect(isGatewayTransport, appForeground, chatReady) {
if (isGatewayTransport && appForeground && chatReady) {
if (appForeground && chatReady) {
chatViewModel.prewarmGateway()
}
if (isGatewayTransport && appForeground && chatReady) {
chatViewModel.refreshModelOptions()
chatViewModel.refreshReasoningSettings()
}
@@ -646,6 +662,13 @@ fun ChatScreen(
var showCommandPalette by remember { mutableStateOf(false) }
var showModelSheet by remember { mutableStateOf(false) }
var showAgentInfo by remember { mutableStateOf(false) }
var showBackgroundProcesses by remember { mutableStateOf(false) }
// A process inventory is scoped to one gateway session. Never leave a
// sheet opened onto a different chat after a drawer/profile switch.
LaunchedEffect(currentSessionId, selectedProfile?.name, activeConnection?.id) {
showBackgroundProcesses = false
}
// Server command dispatch can ask the composer to prefill (e.g. /undo).
LaunchedEffect(chatViewModel) {
@@ -961,8 +984,26 @@ fun ChatScreen(
// streaming auto-scroll effect respects this — it will not yank the
// user back to the latest token while they are reading history.
// Reset to false the moment the user returns to the bottom.
var userScrolledAway by remember { mutableStateOf(false) }
var userScrolledAway by remember(currentSessionId) { mutableStateOf(false) }
var programmaticBottomScroll by remember { mutableStateOf(false) }
val currentUnreadSnapshot = remember(messages) { messages.toUnreadSnapshot() }
var lastReadSnapshot by remember(currentSessionId) {
mutableStateOf(currentUnreadSnapshot)
}
LaunchedEffect(currentSessionId, currentUnreadSnapshot, userScrolledAway) {
if (!userScrolledAway) lastReadSnapshot = currentUnreadSnapshot
}
val unreadMessageCount = remember(
currentUnreadSnapshot,
lastReadSnapshot,
userScrolledAway,
) {
if (userScrolledAway) {
countUnreadMessages(currentUnreadSnapshot, lastReadSnapshot)
} else {
0
}
}
suspend fun scrollConversationToBottom(animated: Boolean) {
programmaticBottomScroll = true
@@ -2045,6 +2086,7 @@ fun ChatScreen(
items(messages.size, key = { messages[it].id }) { index ->
val message = messages[index]
val processNotification = message.hermesProcessNotificationOrNull()
// Skip empty bubbles (content stripped by annotation parser, no tool calls,
// no attachments). Attachments keep the bubble alive for inbound media;
@@ -2053,6 +2095,7 @@ fun ChatScreen(
message.toolCalls.isEmpty() &&
message.attachments.isEmpty() &&
message.cards.isEmpty() &&
message.backgroundTask == null &&
!message.isStreaming
) return@items
@@ -2072,10 +2115,43 @@ fun ChatScreen(
DateSeparator(timestamp = message.timestamp)
}
MessageBubble(
val hasBackgroundTask = message.backgroundTask != null
val shouldRenderBubble =
!hasBackgroundTask ||
message.content.isNotBlank() ||
message.thinkingContent.isNotBlank() ||
message.attachments.isNotEmpty() ||
message.cards.isNotEmpty()
message.backgroundTask?.let { task ->
BackgroundTaskCard(
task = task,
toolCalls = message.toolCalls,
showTimeline = toolDisplay != "off",
modifier = Modifier
.padding(
top = if (isFirstInGroup) 6.dp else 2.dp,
bottom = if (shouldRenderBubble) 3.dp else 0.dp,
)
.animateItem(),
)
}
if (processNotification != null) {
SyntheticProcessNotificationNotice(
notification = processNotification,
modifier = Modifier
.padding(top = if (isFirstInGroup) 6.dp else 2.dp)
.animateItem(),
)
} else if (shouldRenderBubble) MessageBubble(
message = message,
modifier = Modifier
.padding(top = if (isFirstInGroup) 6.dp else 1.dp)
.padding(
top = if (hasBackgroundTask) 1.dp
else if (isFirstInGroup) 6.dp
else 1.dp,
)
.animateItem(),
maxBubbleWidth = maxBubbleWidth,
showThinking = showThinking,
@@ -2172,7 +2248,7 @@ fun ChatScreen(
}
}
if (toolDisplay != "off") {
if (toolDisplay != "off" && !hasBackgroundTask) {
// Subagent children (taskIndex != null) group
// into lanes after the top-level tool cards;
// the null group renders exactly as before.
@@ -2231,7 +2307,16 @@ fun ChatScreen(
.zIndex(8f)
) {
SmallFloatingActionButton(
modifier = Modifier.size(48.dp),
modifier = Modifier
.size(48.dp)
.semantics {
contentDescription = if (unreadMessageCount > 0) {
"Scroll to bottom, $unreadMessageCount unread " +
if (unreadMessageCount == 1) "message" else "messages"
} else {
"Scroll to bottom"
}
},
onClick = {
haptic.performHapticFeedback(HapticFeedbackType.TextHandleMove)
scope.launch {
@@ -2244,10 +2329,25 @@ fun ChatScreen(
},
containerColor = MaterialTheme.colorScheme.primaryContainer
) {
Icon(
Icons.Filled.KeyboardArrowDown,
contentDescription = "Scroll to bottom"
)
BadgedBox(
badge = {
if (unreadMessageCount > 0) {
Badge(
modifier = Modifier.clearAndSetSemantics { },
) {
Text(
if (unreadMessageCount > 99) "99+"
else unreadMessageCount.toString(),
)
}
}
},
) {
Icon(
Icons.Filled.KeyboardArrowDown,
contentDescription = null,
)
}
}
}
@@ -2261,6 +2361,14 @@ fun ChatScreen(
}
}
if (isGatewayTransport) {
GatewayBackgroundProcessStrip(
processes = backgroundProcesses,
loading = backgroundProcessesLoading,
onClick = { showBackgroundProcesses = true },
)
}
// Inline slash command autocomplete
AnimatedVisibility(visible = showAutocomplete) {
InlineAutocomplete(
@@ -2659,24 +2767,34 @@ fun ChatScreen(
}
},
onVoice = {
if (voiceReady) {
requestVoiceMode()
} else {
android.widget.Toast.makeText(
context,
when (standardVoiceAvailability) {
com.hermesandroid.relay.viewmodel.StandardVoiceAvailability.SignInRequired ->
standardVoiceSignInRouteHint?.let { route ->
"Voice needs a one-time sign-in on the $route route — open Manage"
} ?: "Voice needs dashboard sign-in — open Manage to sign in"
com.hermesandroid.relay.viewmodel.StandardVoiceAvailability.Unsupported ->
"This Hermes build has no voice routes — update hermes-agent or pair Relay"
else ->
"Voice needs a reachable Hermes dashboard or Relay voice route"
},
android.widget.Toast.LENGTH_SHORT,
).show()
}
dispatchChatVoiceAction(
isDemoMode = isDemoMode,
voiceReady = voiceReady,
onDemoNotice = {
Toast.makeText(
context,
"Voice is unavailable in the offline demo — connect to Hermes to use it",
Toast.LENGTH_LONG,
).show()
},
onStartVoice = requestVoiceMode,
onSetupNotice = {
Toast.makeText(
context,
when (standardVoiceAvailability) {
com.hermesandroid.relay.viewmodel.StandardVoiceAvailability.SignInRequired ->
standardVoiceSignInRouteHint?.let { route ->
"Voice needs a one-time sign-in on the $route route — open Manage"
} ?: "Voice needs dashboard sign-in — open Manage to sign in"
com.hermesandroid.relay.viewmodel.StandardVoiceAvailability.Unsupported ->
"This Hermes build has no voice routes — update hermes-agent or pair Relay"
else ->
"Voice needs a reachable Hermes dashboard or Relay voice route"
},
Toast.LENGTH_SHORT,
).show()
},
)
},
onStop = {
chatViewModel.cancelStream()
@@ -2918,6 +3036,18 @@ fun ChatScreen(
)
}
if (showBackgroundProcesses) {
GatewayBackgroundProcessSheet(
processes = backgroundProcesses,
loading = backgroundProcessesLoading,
stoppingProcessIds = stoppingProcessIds,
onRefresh = chatViewModel::refreshBackgroundProcesses,
onStop = chatViewModel::stopBackgroundProcess,
onDismissProcess = chatViewModel::dismissBackgroundProcess,
onDismiss = { showBackgroundProcesses = false },
)
}
// Agent info sheet — one consolidated surface for agent state (profile,
// personality, connection summary). Replaces the old AlertDialog and the
// two top-bar chips (ProfilePicker + PersonalityPicker). Tap target is
@@ -0,0 +1,101 @@
package com.hermesandroid.relay.ui.screens
import com.hermesandroid.relay.data.ChatMessage
/**
* Compact record of the conversation content that was visible at the bottom.
* Only the tail can grow without increasing [messageCount], so keeping one
* visible-content revision avoids hashing the entire transcript on each token.
*/
internal data class ChatUnreadSnapshot(
val messageCount: Int,
val lastMessageId: String?,
val lastVisibleRevision: Int,
)
internal fun List<ChatMessage>.toUnreadSnapshot(): ChatUnreadSnapshot {
val last = lastOrNull()
return ChatUnreadSnapshot(
messageCount = size,
lastMessageId = last?.id,
lastVisibleRevision = last?.visibleUnreadRevision() ?: 0,
)
}
internal fun countUnreadMessages(
current: ChatUnreadSnapshot,
lastRead: ChatUnreadSnapshot,
): Int {
if (current.messageCount < lastRead.messageCount) return 0
val appended = (current.messageCount - lastRead.messageCount).coerceAtLeast(0)
if (appended > 0) return appended
if (current.messageCount == 0) return 0
return if (
current.lastMessageId != lastRead.lastMessageId ||
current.lastVisibleRevision != lastRead.lastVisibleRevision
) {
1
} else {
0
}
}
internal enum class ChatVoiceAction {
ShowDemoNotice,
StartVoice,
ShowSetupNotice,
}
/** Demo is offline, so it must win before any permission or route check. */
internal fun resolveChatVoiceAction(
isDemoMode: Boolean,
voiceReady: Boolean,
): ChatVoiceAction = when {
isDemoMode -> ChatVoiceAction.ShowDemoNotice
voiceReady -> ChatVoiceAction.StartVoice
else -> ChatVoiceAction.ShowSetupNotice
}
internal inline fun dispatchChatVoiceAction(
isDemoMode: Boolean,
voiceReady: Boolean,
onDemoNotice: () -> Unit,
onStartVoice: () -> Unit,
onSetupNotice: () -> Unit,
) {
when (resolveChatVoiceAction(isDemoMode, voiceReady)) {
ChatVoiceAction.ShowDemoNotice -> onDemoNotice()
ChatVoiceAction.StartVoice -> onStartVoice()
ChatVoiceAction.ShowSetupNotice -> onSetupNotice()
}
}
private fun ChatMessage.visibleUnreadRevision(): Int {
var revision = 17
revision = 31 * revision + content.length
revision = 31 * revision + thinkingContent.length
revision = 31 * revision + toolCalls.fold(1) { acc, call ->
var callRevision = 17
callRevision = 31 * callRevision + (call.id?.hashCode() ?: 0)
callRevision = 31 * callRevision + call.name.hashCode()
callRevision = 31 * callRevision + (call.result?.hashCode() ?: 0)
callRevision = 31 * callRevision + (call.error?.hashCode() ?: 0)
callRevision = 31 * callRevision + (call.success?.hashCode() ?: 0)
callRevision = 31 * callRevision + call.isComplete.hashCode()
callRevision = 31 * callRevision + call.isGenerating.hashCode()
31 * acc + callRevision
}
revision = 31 * revision + attachments.fold(1) { acc, attachment ->
var attachmentRevision = 17
attachmentRevision = 31 * attachmentRevision + attachment.contentType.hashCode()
attachmentRevision = 31 * attachmentRevision + (attachment.fileName?.hashCode() ?: 0)
attachmentRevision = 31 * attachmentRevision + attachment.state.hashCode()
attachmentRevision = 31 * attachmentRevision + (attachment.errorMessage?.hashCode() ?: 0)
31 * acc + attachmentRevision
}
revision = 31 * revision + cards.hashCode()
revision = 31 * revision + cardDispatches.hashCode()
revision = 31 * revision + (backgroundTask?.hashCode() ?: 0)
return revision
}
@@ -70,10 +70,15 @@ import com.hermesandroid.relay.data.BargeInSensitivity
import com.hermesandroid.relay.data.Profile
import com.hermesandroid.relay.data.VoiceAudioRoute
import com.hermesandroid.relay.data.VoiceEngineMode
import com.hermesandroid.relay.data.VoiceModePreset
import com.hermesandroid.relay.data.VoiceModePresetState
import com.hermesandroid.relay.data.VoicePreferencesRepository
import com.hermesandroid.relay.data.VoicePresetPromotionSettings
import com.hermesandroid.relay.data.VoiceSettings
import com.hermesandroid.relay.data.detectVoiceModePreset
import com.hermesandroid.relay.network.relay.RealtimeProviderInfo
import com.hermesandroid.relay.network.relay.RealtimeVoiceConfig
import com.hermesandroid.relay.network.relay.RealtimeVoicePromotion
import com.hermesandroid.relay.network.relay.RelayVoiceClient
import com.hermesandroid.relay.network.relay.VoiceConfig
import com.hermesandroid.relay.network.relay.VoiceOutputConfig
@@ -189,6 +194,97 @@ fun VoiceSettingsScreen(
// are shown here as well as the inline "unavailable" labels.
val snackbarHost = LocalSnackbarHost.current
val scope = rememberCoroutineScope()
var presetApplying by remember { mutableStateOf(false) }
val presetState = VoiceModePresetState(
voiceSettings = voiceSettings,
bargeInPreferences = bargeInPrefs,
promotion = configState.realtimeConfig?.promotion?.toPresetSettings(),
)
val activePreset = detectVoiceModePreset(presetState)
val presetsReady = voiceClient != null && presetState.promotion != null
fun applyPreset(preset: VoiceModePreset) {
if (presetApplying) return
val client = voiceClient
val priorPromotion = presetState.promotion
if (client == null || priorPromotion == null) {
scope.launch {
snackbarHost.showSnackbar(
"Realtime Agent background settings must be available before applying a preset.",
)
}
return
}
val target = preset.applyTo(presetState)
val update = preset.promotionUpdate
scope.launch {
presetApplying = true
try {
// Server first: if the relay rejects a preset, local controls
// stay untouched and the UI cannot falsely report it active.
val result = client.updateRealtimeAgentPromotion(
promotionEnabled = update.enabled,
promoteAfterMs = update.promoteAfterMs,
spokenHandoff = update.spokenHandoff,
resultDelivery = update.resultDelivery,
backgroundDefaultMode = update.backgroundDefaultMode,
progressSpokenAfterMs = update.progressSpokenAfterMs,
progressRepeatMs = update.progressRepeatMs,
maxBackgroundRuns = update.maxBackgroundRuns,
)
if (result.isFailure) {
snackbarHost.showHumanError(
classifyError(result.exceptionOrNull(), context = "voice_config"),
)
return@launch
}
settingsViewModel.setRealtimeConfig(result.getOrNull())
try {
// One DataStore transaction covers Voice + barge-in.
prefsRepo.applyModePreset(preset)
} catch (error: Exception) {
// The network and DataStore cannot share one transaction.
// Restore the captured relay values so a local write
// failure does not leave a half-applied preset.
val rollback = client.updateRealtimeAgentPromotion(
promotionEnabled = priorPromotion.enabled,
promoteAfterMs = priorPromotion.promoteAfterMs,
spokenHandoff = priorPromotion.spokenHandoff,
resultDelivery = priorPromotion.resultDelivery,
backgroundDefaultMode = priorPromotion.backgroundDefaultMode,
progressSpokenAfterMs = priorPromotion.progressSpokenAfterMs,
progressRepeatMs = priorPromotion.progressRepeatMs,
maxBackgroundRuns = priorPromotion.maxBackgroundRuns,
)
if (rollback.isSuccess) {
settingsViewModel.setRealtimeConfig(rollback.getOrNull())
snackbarHost.showHumanError(
classifyError(error, context = "voice_config"),
)
} else {
snackbarHost.showSnackbar(
"Preset partly applied: Relay settings changed, but phone " +
"settings could not be saved. Reapply a preset to recover.",
)
}
return@launch
}
voiceViewModel.setInteractionMode(
when (target.voiceSettings.interactionMode) {
"hold" -> InteractionMode.HoldToTalk
"continuous" -> InteractionMode.Continuous
else -> InteractionMode.TapToTalk
},
)
snackbarHost.showSnackbar("${preset.displayName} preset applied")
} finally {
presetApplying = false
}
}
}
// WP-V2/V3: point the screen's prefs repo at the active (connection,
// profile) scope so the per-profile engine/route/enhanced toggles read and
@@ -265,6 +361,13 @@ fun VoiceSettingsScreen(
selectedProfile = selectedProfile,
)
VoiceModePresetCard(
activePreset = activePreset,
enabled = presetsReady,
applying = presetApplying,
onSelect = ::applyPreset,
)
// --- Voice for this profile: engine + route ---
VoiceForThisProfileCard(
currentEngine = currentEngine,
@@ -457,6 +560,98 @@ private fun scopeFallback(config: Any?): Boolean = when (config) {
else -> false
}
private fun RealtimeVoicePromotion.toPresetSettings(): VoicePresetPromotionSettings =
VoicePresetPromotionSettings(
enabled = enabled,
promoteAfterMs = promoteAfterMs,
backgroundDefaultMode = backgroundDefaultMode,
spokenHandoff = spokenHandoff,
progressSpokenAfterMs = progressSpokenAfterMs,
progressRepeatMs = progressRepeatMs,
resultDelivery = resultDelivery,
maxBackgroundRuns = maxBackgroundRuns,
)
// ---------------------------------------------------------------------------
// Mode presets — compact bundles over controls already present on this screen.
// ---------------------------------------------------------------------------
@Composable
private fun VoiceModePresetCard(
activePreset: VoiceModePreset?,
enabled: Boolean,
applying: Boolean,
onSelect: (VoiceModePreset) -> Unit,
) {
SectionCard(title = "Mode preset") {
Text(
text = "Tune interaction, interruption, trace, and long-task delivery together.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(8.dp))
// Two rows keep each target about 148dp wide on a 360dp screen after
// screen/card padding; all four labels remain readable without tiny
// type or ambiguous abbreviations.
Column(verticalArrangement = Arrangement.spacedBy(6.dp)) {
VoiceModePreset.entries.chunked(2).forEach { rowPresets ->
SingleChoiceSegmentedButtonRow(modifier = Modifier.fillMaxWidth()) {
rowPresets.forEachIndexed { index, preset ->
SegmentedButton(
shape = SegmentedButtonDefaults.itemShape(
index = index,
count = rowPresets.size,
),
onClick = { onSelect(preset) },
selected = activePreset == preset,
enabled = enabled && !applying,
) {
Text(
text = preset.shortLabel,
style = MaterialTheme.typography.labelSmall,
maxLines = 1,
overflow = TextOverflow.Ellipsis,
)
}
}
}
}
}
Spacer(Modifier.height(8.dp))
Text(
text = activePreset?.displayName ?: "Custom",
style = MaterialTheme.typography.labelLarge,
color = MaterialTheme.colorScheme.onSurface,
)
Text(
text = activePreset?.description
?: "Your manual values do not exactly match a preset.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
if (!enabled) {
Text(
text = "Connect Relay voice so background delivery can be " +
"applied with the local controls.",
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
if (applying) {
Spacer(Modifier.height(4.dp))
LinearProgressIndicator(modifier = Modifier.fillMaxWidth())
}
Spacer(Modifier.height(4.dp))
Text(
text = "Engine, route, provider, model, voice, and credentials stay unchanged.",
style = MaterialTheme.typography.labelSmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
// ---------------------------------------------------------------------------
// Voice for this profile — engine + STT/TTS route (per-profile prefs).
// ---------------------------------------------------------------------------
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,305 @@
package com.hermesandroid.relay.viewmodel
import com.hermesandroid.relay.network.upstream.GatewayProcess
import com.hermesandroid.relay.network.upstream.GatewayProcessCapability
import com.hermesandroid.relay.network.upstream.GatewayProcessEvent
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Job
import kotlinx.coroutines.delay
import kotlinx.coroutines.flow.MutableStateFlow
import kotlinx.coroutines.flow.StateFlow
import kotlinx.coroutines.flow.asStateFlow
import kotlinx.coroutines.launch
/**
* Small adapter seam around [com.hermesandroid.relay.network.upstream.GatewayChatClient].
* Keeping the controller on this interface makes its session/race behavior
* testable without opening a WebSocket.
*/
internal interface GatewayProcessSource {
val capability: StateFlow<GatewayProcessCapability>
suspend fun listProcesses(): Result<List<GatewayProcess>>
suspend fun killProcess(processId: String): Result<Unit>
fun setEventListener(listener: ((GatewayProcessEvent) -> Unit)?)
/** False when background polling would reopen a deliberately closed socket. */
fun isPollingAllowed(): Boolean
}
/**
* Owns the active chat's upstream background-process snapshot.
*
* The gateway RPCs are scoped by a private, live session id. The UI only knows
* the durable chat id, so a session becomes queryable only after [sessionReady]
* confirms that the source has resumed/created its matching live session.
* Every async result is fenced by both binding/session generation and request
* sequence so a slow response from a previous chat can never repaint this one.
*/
internal class GatewayProcessController(
private val scope: CoroutineScope,
private val pollIntervalMs: Long = 5_000L,
private val outputTailLimit: Int = 4_000,
) {
private val _processes = MutableStateFlow<List<GatewayProcess>>(emptyList())
val processes: StateFlow<List<GatewayProcess>> = _processes.asStateFlow()
private val _capability = MutableStateFlow(GatewayProcessCapability.Unknown)
val capability: StateFlow<GatewayProcessCapability> = _capability.asStateFlow()
private val _loading = MutableStateFlow(false)
val loading: StateFlow<Boolean> = _loading.asStateFlow()
private val _stoppingProcessIds = MutableStateFlow<Set<String>>(emptySet())
val stoppingProcessIds: StateFlow<Set<String>> = _stoppingProcessIds.asStateFlow()
private var source: GatewayProcessSource? = null
private var selectedSessionId: String? = null
private var selectedScopeKey: String? = null
private var readySessionId: String? = null
private var generation = 0L
private var refreshSequence = 0L
private var allProcesses: List<GatewayProcess> = emptyList()
/** A dismissal applies only to this concrete process identity. */
private val dismissedIdentities = mutableMapOf<String, ProcessIdentity>()
private var capabilityJob: Job? = null
private var refreshJob: Job? = null
private var pollJob: Job? = null
/** Replace the gateway client and clear all session-owned state. */
fun bind(newSource: GatewayProcessSource?, sessionId: String?, scopeKey: String? = null) {
if (
source === newSource &&
selectedSessionId == sessionId &&
selectedScopeKey == scopeKey
) return
source?.setEventListener(null)
capabilityJob?.cancel()
source = newSource
resetForSession(sessionId, scopeKey)
_capability.value = newSource?.capability?.value ?: GatewayProcessCapability.Unknown
if (newSource == null) return
newSource.setEventListener { event ->
// Gateway callbacks arrive on the socket/callback dispatcher. Keep
// snapshot transitions serialized with refresh/kill mutations.
scope.launch {
if (source === newSource) handleEvent(newSource, event)
}
}
capabilityJob = scope.launch {
newSource.capability.collect { value ->
if (source !== newSource) return@collect
_capability.value = value
if (value == GatewayProcessCapability.Unsupported) {
clearSnapshot(cancelDismissals = true)
_loading.value = false
}
}
}
}
/** Select a durable chat id. The live gateway session may follow later. */
fun selectSession(sessionId: String?, scopeKey: String? = null) {
if (selectedSessionId == sessionId && selectedScopeKey == scopeKey) return
resetForSession(sessionId, scopeKey)
}
/**
* Admit process RPCs for [sessionId] after the gateway has created/resumed
* that chat's live session. Stale ready callbacks are ignored.
*/
fun sessionReady(sessionId: String) {
if (sessionId != selectedSessionId || source == null) return
val firstReady = readySessionId != sessionId
readySessionId = sessionId
// A same-session ready callback can mean the socket was resumed after
// being offline. Re-list even when this chat was already admitted: a
// process may have started/finished while no event stream was present.
refresh(showLoading = firstReady && allProcesses.isEmpty())
}
fun refresh(showLoading: Boolean = allProcesses.isEmpty()) {
val currentSource = source ?: return
val sessionId = selectedSessionId ?: return
if (readySessionId != sessionId) return
if (_capability.value == GatewayProcessCapability.Unsupported) return
val expectedGeneration = generation
val request = ++refreshSequence
refreshJob?.cancel()
refreshJob = scope.launch {
if (showLoading && allProcesses.isEmpty()) _loading.value = true
val result = currentSource.listProcesses()
if (!owns(currentSource, sessionId, expectedGeneration) || request != refreshSequence) {
return@launch
}
result.onSuccess(::applySnapshot)
_loading.value = false
}
}
/** Stop one running process; the authoritative follow-up snapshot wins. */
fun stop(processId: String, onFailure: (String) -> Unit = {}) {
val currentSource = source ?: return
val sessionId = selectedSessionId ?: return
if (readySessionId != sessionId) return
if (processId in _stoppingProcessIds.value) return
if (allProcesses.none { it.id == processId && it.isRunning }) return
val expectedGeneration = generation
_stoppingProcessIds.value = _stoppingProcessIds.value + processId
scope.launch {
val result = currentSource.killProcess(processId)
if (!owns(currentSource, sessionId, expectedGeneration)) return@launch
_stoppingProcessIds.value = _stoppingProcessIds.value - processId
result.fold(
onSuccess = { refresh(showLoading = false) },
onFailure = { error ->
onFailure(error.message ?: "Couldn't stop background process")
},
)
}
}
/** Finished rows are dismissed locally; upstream history remains untouched. */
fun dismiss(processId: String) {
val process = allProcesses.firstOrNull { it.id == processId && !it.isRunning } ?: return
dismissedIdentities[processId] = process.identity()
publishVisibleSnapshot()
}
fun close() {
source?.setEventListener(null)
source = null
capabilityJob?.cancel()
capabilityJob = null
resetForSession(null, null)
_capability.value = GatewayProcessCapability.Unknown
}
private fun resetForSession(sessionId: String?, scopeKey: String?) {
generation += 1
refreshSequence += 1
selectedSessionId = sessionId
selectedScopeKey = scopeKey
readySessionId = null
refreshJob?.cancel()
refreshJob = null
pollJob?.cancel()
pollJob = null
_loading.value = false
_stoppingProcessIds.value = emptySet()
clearSnapshot(cancelDismissals = true)
}
private fun clearSnapshot(cancelDismissals: Boolean) {
allProcesses = emptyList()
_processes.value = emptyList()
pollJob?.cancel()
pollJob = null
if (cancelDismissals) {
dismissedIdentities.clear()
}
}
private fun applySnapshot(incoming: List<GatewayProcess>) {
val incomingById = incoming.associateBy(GatewayProcess::id)
// A server can eventually reuse a process id. Never let a local
// dismissal hide the new command instance.
dismissedIdentities.entries.removeAll { (id, dismissedIdentity) ->
val next = incomingById[id]
next == null || next.identity() != dismissedIdentity
}
allProcesses = incoming
publishVisibleSnapshot()
updatePoller()
}
private fun publishVisibleSnapshot() {
_processes.value = allProcesses.filterNot { process ->
dismissedIdentities[process.id] == process.identity()
}
}
private fun updatePoller() {
val shouldPoll = allProcesses.any(GatewayProcess::isRunning) &&
readySessionId == selectedSessionId &&
source != null &&
source?.isPollingAllowed() == true &&
_capability.value != GatewayProcessCapability.Unsupported
if (!shouldPoll) {
pollJob?.cancel()
pollJob = null
return
}
if (pollJob?.isActive == true) return
val expectedGeneration = generation
pollJob = scope.launch {
while (generation == expectedGeneration && allProcesses.any(GatewayProcess::isRunning)) {
delay(pollIntervalMs)
if (
generation != expectedGeneration ||
!allProcesses.any(GatewayProcess::isRunning) ||
source?.isPollingAllowed() != true
) {
break
}
// Avoid repeatedly cancelling a slow RPC. An invalidation can
// own this interval; the next tick remains the safety net.
if (refreshJob?.isActive != true) refresh(showLoading = false)
}
if (generation == expectedGeneration) pollJob = null
}
}
private fun handleEvent(currentSource: GatewayProcessSource, event: GatewayProcessEvent) {
val sessionId = selectedSessionId ?: return
if (!owns(currentSource, sessionId, generation) || readySessionId != sessionId) return
when (event) {
is GatewayProcessEvent.Invalidated,
is GatewayProcessEvent.TerminalClosed -> refresh(showLoading = false)
is GatewayProcessEvent.Output -> {
val index = allProcesses.indexOfFirst { it.id == event.processId }
if (index < 0) {
// Output can beat tool.complete/process.list by a frame.
refresh(showLoading = false)
return
}
val process = allProcesses[index]
val tail = (process.outputTail.orEmpty() + event.chunk).takeLast(outputTailLimit)
allProcesses = allProcesses.toMutableList().also {
it[index] = process.copy(outputTail = tail)
}
publishVisibleSnapshot()
}
}
}
private fun owns(
expectedSource: GatewayProcessSource,
expectedSessionId: String,
expectedGeneration: Long,
): Boolean =
source === expectedSource &&
selectedSessionId == expectedSessionId &&
readySessionId == expectedSessionId &&
generation == expectedGeneration
private data class ProcessIdentity(
val id: String,
val command: String,
val pid: Long?,
val startedAt: String?,
)
private fun GatewayProcess.identity() = ProcessIdentity(id, command, pid, startedAt)
}
@@ -37,6 +37,9 @@ import com.hermesandroid.relay.network.shared.LocalDispatchResult
import com.hermesandroid.relay.util.HumanError
import com.hermesandroid.relay.util.classifyError
import com.hermesandroid.relay.voice.VoiceIntentSyncBuilder
import com.hermesandroid.relay.voice.VoiceCommandAction
import com.hermesandroid.relay.voice.VoiceCommandContext
import com.hermesandroid.relay.voice.VoiceCommandInterpreter
// === PHASE3-voice-intents: voice→bridge intent routing ===
import com.hermesandroid.relay.voice.IntentResult
import com.hermesandroid.relay.voice.LocalBridgeDispatcher
@@ -663,6 +666,15 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
private var realtimeAmplitudeDecayJob: Job? = null
private var firstFrameWatchdogJob: Job? = null
private var continuousLoopArmed: Boolean = false
/**
* Remembers an explicit hands-free pause across the one manual mic turn
* needed to say "resume continuous listening". A normal utterance after
* that tap still clears the flag and keeps the existing tap-to-rearm
* behavior.
*/
private var continuousListeningPaused: Boolean = false
/** One final-transcript window opened specifically by barge-in playback. */
private var responseInterruptedForVoiceCommand: Boolean = false
private var lastRealtimeAudioDeltaAtMs: Long = 0L
/**
@@ -1123,6 +1135,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
if (mode != InteractionMode.Continuous) {
continuousLoopArmed = false
continuousListeningPaused = false
continuousResumeJob?.cancel()
continuousResumeJob = null
}
@@ -1139,6 +1152,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
if (mode == InteractionMode.Continuous && _uiState.value.voiceMode) {
continuousLoopArmed = true
continuousListeningPaused = false
when (_uiState.value.state) {
VoiceState.Idle, VoiceState.Error -> startListening()
else -> Unit
@@ -1327,6 +1341,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
resetBrokeredToolSpeechState()
resetRealtimeSpeechCoalescer()
continuousLoopArmed = false
continuousListeningPaused = false
continuousResumeJob?.cancel()
continuousResumeJob = null
realtimeAmplitudeDecayJob?.cancel()
@@ -1650,6 +1665,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
streamObserverJob?.cancel()
streamObserverJob = null
continuousLoopArmed = false
continuousListeningPaused = false
continuousResumeJob?.cancel()
continuousResumeJob = null
realtimeAmplitudeDecayJob?.cancel()
@@ -1694,6 +1710,9 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
// ---------------------------------------------------------------------
fun startListening() {
// A direct mic tap starts a normal capture. Only the recorder opened by
// onBargeInDetected may carry response-interruption command context.
responseInterruptedForVoiceCommand = false
val rec = recorder
if (rec == null) {
setError("Recorder not initialized")
@@ -1708,6 +1727,10 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
continuousResumeJob?.cancel()
continuousResumeJob = null
// A manual mic tap after an explicit pause arms this capture so it can
// become either the exact "resume" command or a normal prompt. Keep
// continuousListeningPaused until the final transcript boundary: the
// command handler clears it explicitly in either branch.
continuousLoopArmed = _uiState.value.interactionMode == InteractionMode.Continuous
lastRealtimeAudioDeltaAtMs = 0L
realtimeAmplitudeDecayJob?.cancel()
@@ -1792,6 +1815,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
listeningStartedAtMs = 0L
if (shouldDiscardVoiceCapture(captureDurationMs, inputPcm.size)) {
responseInterruptedForVoiceCommand = false
try { file.delete() } catch (_: Exception) { /* ignore */ }
DiagnosticsLog.record(
category = DiagnosticCategory.Voice,
@@ -1824,6 +1848,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
durationMs < MIN_VOICE_CAPTURE_DURATION_MS
private fun cancelListeningWithoutProcessing(title: String, detail: String? = null) {
responseInterruptedForVoiceCommand = false
silenceWatchdogJob?.cancel()
silenceWatchdogJob = null
listeningStartedAtMs = 0L
@@ -1865,6 +1890,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
*/
fun pauseContinuousMode() {
continuousLoopArmed = false
continuousListeningPaused = _uiState.value.interactionMode == InteractionMode.Continuous
continuousResumeJob?.cancel()
continuousResumeJob = null
realtimeAmplitudeDecayJob?.cancel()
@@ -1887,6 +1913,29 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
}
/**
* Re-arm an explicitly paused Continuous loop and immediately return to
* live listening. This is intentionally a no-op outside Continuous mode;
* a spoken resume phrase must never switch the user's saved interaction
* mode behind their back.
*/
fun resumeContinuousMode() {
if (_uiState.value.interactionMode != InteractionMode.Continuous) return
continuousListeningPaused = false
continuousLoopArmed = true
continuousResumeJob?.cancel()
continuousResumeJob = null
_uiState.update {
it.copy(
state = VoiceState.Idle,
amplitude = 0f,
outputAudioActive = false,
responseText = "Continuous listening resumed.",
)
}
startListening()
}
/**
* Silence-based auto-stop watchdog for the current Listening turn.
*
@@ -2337,6 +2386,129 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
// Voice turn processing
// ---------------------------------------------------------------------
private fun voiceCommandContext(
responseActiveOverride: Boolean? = null,
allowNewChat: Boolean = true,
responseWasInterrupted: Boolean = false,
): VoiceCommandContext {
val ui = _uiState.value
val backgroundPhase = ui.backgroundRun?.phase
val backgroundTaskActive = backgroundPhase == BackgroundRunPhase.RUNNING ||
backgroundPhase == BackgroundRunPhase.RECONNECTING
val backgroundDeliveryActive = backgroundPhase == BackgroundRunPhase.DELIVERING
val responseActive = responseWasInterrupted ||
(responseActiveOverride ?: (
ui.state == VoiceState.Speaking ||
chatViewModel?.isStreaming?.value == true ||
providerRealtimeAgentTurnActive.get()
))
return VoiceCommandContext(
responseActive = responseActive,
backgroundTaskActive = backgroundTaskActive,
backgroundAnswerAvailable = backgroundPhase == BackgroundRunPhase.DONE &&
realtimeAgentControl != null,
continuousModeSelected = ui.interactionMode == InteractionMode.Continuous,
continuousListeningActive = ui.interactionMode == InteractionMode.Continuous &&
continuousLoopArmed &&
!continuousListeningPaused,
continuousListeningPaused = continuousListeningPaused,
canStartNewChat = allowNewChat &&
!responseActive &&
!backgroundTaskActive &&
!backgroundDeliveryActive &&
chatViewModel?.isStreaming?.value != true,
)
}
/**
* Intercept one committed transcript at the last local boundary before it
* would become a normal Hermes prompt. Returns the typed action when the
* transcript was consumed, or null when it must continue through ordinary
* intent/chat routing.
*/
private fun tryHandleFinalVoiceCommand(
transcript: String,
responseActiveOverride: Boolean? = null,
fromRealtime: Boolean = false,
realtimeControl: RealtimeAgentSessionControl? = null,
): VoiceCommandAction? {
// A persistent Realtime Agent websocket is bound to the chat session
// it opened with. Starting a new chat requires an explicit session
// reset/rebind boundary that this callback does not own, so leave that
// typed action for the Standard path (and a future realtime coordinator).
val responseWasInterrupted = responseInterruptedForVoiceCommand
responseInterruptedForVoiceCommand = false
val context = voiceCommandContext(
responseActiveOverride = responseActiveOverride,
allowNewChat = !fromRealtime,
responseWasInterrupted = responseWasInterrupted,
)
val action = VoiceCommandInterpreter.interpretFinalTranscript(transcript, context)
if (action == null) {
// The user manually tapped the mic after an explicit pause and said
// an ordinary prompt. Preserve the existing tap-to-rearm behavior.
if (continuousListeningPaused && continuousLoopArmed) {
continuousListeningPaused = false
}
return null
}
Log.i(TAG, "Hands-free voice command action=$action source=${if (fromRealtime) "realtime" else "stt"}")
DiagnosticsLog.record(
category = DiagnosticCategory.Voice,
severity = DiagnosticSeverity.Info,
title = "Hands-free voice command",
detail = action.name,
)
if (fromRealtime) {
chatViewModel?.discardRealtimeAgentLocalCommandTurn(rtAssistantMessageId)
}
val backgroundOwnsResponse = preserveRealtimeTurnOnStop(
_uiState.value.backgroundRun?.phase,
)
when (action) {
VoiceCommandAction.StopResponse -> interruptSpeaking()
VoiceCommandAction.CancelBackgroundTask -> cancelBackgroundRun()
VoiceCommandAction.PauseContinuousListening -> pauseContinuousMode()
VoiceCommandAction.ResumeContinuousListening -> {
if (fromRealtime && !backgroundOwnsResponse) {
realtimeControl?.cancel()
}
resumeContinuousMode()
}
VoiceCommandAction.RepeatBackgroundAnswer -> {
// Cancel only the provider's would-be conversational answer,
// then use the existing relay respeak command for the durable
// result. The completed background task itself cannot be lost.
if (fromRealtime) realtimeControl?.cancel()
respeakBackgroundResult()
}
VoiceCommandAction.StartNewChat -> {
chatViewModel?.createNewChat()
val resumeContinuous = _uiState.value.interactionMode == InteractionMode.Continuous
continuousListeningPaused = false
continuousLoopArmed = resumeContinuous
_uiState.update {
it.copy(
state = VoiceState.Idle,
outputAudioActive = false,
responseText = "New chat started.",
)
}
if (resumeContinuous) startListening()
}
}
// The callback-local response gate owns late audio/text for this
// command. Keep the session-global gate open so a preserved background
// task can deliver even when Pause/Stop called interruptSpeaking().
if (fromRealtime) realtimeAudioSuppressed = false
return action
}
private suspend fun processVoiceInput(
audioFile: File,
inputPcm: ByteArray,
@@ -2433,6 +2605,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
val transcribeResult = audioClient.transcribe(audioFile)
val sttLatencyMs = System.currentTimeMillis() - sttStartedAtMs
if (transcribeResult.isFailure) {
responseInterruptedForVoiceCommand = false
val err = transcribeResult.exceptionOrNull()
Log.w(TAG, "transcribe failed: ${err?.message}")
DiagnosticsLog.record(
@@ -2446,6 +2619,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
val userText = transcribeResult.getOrNull().orEmpty()
if (userText.isBlank()) {
responseInterruptedForVoiceCommand = false
DiagnosticsLog.record(
category = DiagnosticCategory.Voice,
severity = DiagnosticSeverity.Warning,
@@ -2464,6 +2638,18 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
)
}
// This is the Standard route's committed STT boundary. Commands are
// intentionally checked here — never while recorder amplitude or a
// partial transcript is still changing — and before bridge/chat intent
// routing can consume the phrase as an ordinary prompt.
_uiState.update {
it.copy(
outputAudioActive = false,
transcribedText = userText,
)
}
if (tryHandleFinalVoiceCommand(userText) != null) return
_uiState.update {
it.copy(
state = VoiceState.Thinking,
@@ -2826,6 +3012,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
var hadActiveTurnWhenSessionEnded = false
var suppressLocalCommandResponse = false
val result = try {
client.runRealtimeAgent(
prompt = userText,
@@ -2854,11 +3041,39 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
return@runRealtimeAgent
}
realtimeAgentControl = control
chatVm.applyRealtimeAgentEvent(
assistantMessageId = rtAssistantMessageId,
event = event,
showDetailedTrace = realtimeTraceDetails,
)
val providerResponseEvent = event.type == "voice.response.started" ||
event.type == "voice.response.delta" ||
event.type == "voice.playback_drain.requested" ||
event.type == "voice.output_audio.delta" ||
event.type == "voice.output_audio.done" ||
event.type == "voice.response.done"
val deliveryResponseEvent = !event.delivery.isNullOrBlank()
if (suppressLocalCommandResponse && deliveryResponseEvent) {
// A completed background task owns this response, not the
// intercepted local command. Let authoritative delivery
// through and retire the command-only gate.
suppressLocalCommandResponse = false
realtimeAudioSuppressed = false
}
val suppressCommandResponse =
suppressLocalCommandResponse && providerResponseEvent
if (!suppressCommandResponse) {
chatVm.applyRealtimeAgentEvent(
assistantMessageId = rtAssistantMessageId,
event = event,
showDetailedTrace = realtimeTraceDetails,
)
}
if (suppressCommandResponse) {
if (event.type == "voice.response.done") {
providerRealtimeAgentTurnActive.set(
realtimeTurnActiveAfterResponseDone(_uiState.value.backgroundRun?.phase),
)
suppressLocalCommandResponse = false
realtimeAudioSuppressed = false
}
return@runRealtimeAgent
}
when (event.type) {
"voice.input_transcript.delta" -> {
event.delta?.let { inputTranscript.append(it) }
@@ -2871,15 +3086,34 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
}
"voice.input_transcript.final" -> {
// A committed transcript starts a new turn boundary. A
// cancelled local command may not receive response.done,
// so never carry its suppression into the next utterance.
suppressLocalCommandResponse = false
inputTranscript.clear()
inputTranscript.append(event.text ?: rtUserText)
val finalTranscript = inputTranscript.toString()
_uiState.update {
it.copy(
state = VoiceState.Thinking,
outputAudioActive = false,
transcribedText = inputTranscript.toString(),
transcribedText = finalTranscript,
)
}
val command = tryHandleFinalVoiceCommand(
transcript = finalTranscript,
responseActiveOverride = providerRealtimeAgentTurnActive.get(),
fromRealtime = true,
realtimeControl = control,
)
if (command != null) {
// Re-speak output is intentionally allowed through;
// it is the requested durable answer, not the
// provider's conversational response to this phrase.
suppressLocalCommandResponse =
command != VoiceCommandAction.RepeatBackgroundAnswer
return@runRealtimeAgent
}
}
"voice.response.started", "hermes.run.started" -> {
providerRealtimeAgentTurnActive.set(true)
@@ -4721,6 +4955,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
// interruptSpeaking landed us in Idle — flip to Listening and
// pre-warm the recorder so the first ~100 ms of user speech
// isn't clipped by recorder cold-start.
responseInterruptedForVoiceCommand = true
_uiState.update {
it.copy(
state = VoiceState.Listening,
@@ -4734,6 +4969,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
try {
rec.startRecording()
} catch (t: Throwable) {
responseInterruptedForVoiceCommand = false
Log.w(TAG, "barge-in pre-warm recorder failed: ${t.message}")
surfaceError(t, context = "record")
return
@@ -4805,6 +5041,8 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
return@launch
}
responseInterruptedForVoiceCommand = false
// Silence after interrupt. If resume is off, drop the tail
// and return to Idle (the user wanted a hard cancel semantic).
if (!prefs.resumeAfterInterruption) {
@@ -0,0 +1,107 @@
package com.hermesandroid.relay.voice
import java.util.Locale
/**
* Local actions that can be requested from a committed voice transcript.
*
* These are deliberately separate from normal Hermes prompts. Callers must
* invoke [VoiceCommandInterpreter.interpretFinalTranscript] only after STT (or
* a realtime provider) has emitted a final transcript; partial transcripts are
* never safe command boundaries.
*/
internal enum class VoiceCommandAction {
StopResponse,
CancelBackgroundTask,
PauseContinuousListening,
ResumeContinuousListening,
RepeatBackgroundAnswer,
StartNewChat,
}
/** State gates that keep an exact command phrase from becoming a global hotword. */
internal data class VoiceCommandContext(
val responseActive: Boolean = false,
val backgroundTaskActive: Boolean = false,
val backgroundAnswerAvailable: Boolean = false,
val continuousModeSelected: Boolean = false,
val continuousListeningActive: Boolean = false,
val continuousListeningPaused: Boolean = false,
val canStartNewChat: Boolean = false,
)
/**
* Conservative, exact-only interpreter for hands-free Voice controls.
*
* False negatives are preferred: a phrase must match one complete normalized
* utterance and its corresponding state gate. There is no prefix, substring,
* edit-distance, or fuzzy matching, so ordinary prompts such as "How do I stop
* talking too quickly?" continue to Hermes unchanged.
*/
internal object VoiceCommandInterpreter {
private val stopResponsePhrases = setOf(
"stop speaking",
"stop talking",
"stop the response",
"stop your response",
)
private val cancelBackgroundTaskPhrases = setOf(
"cancel the background task",
"cancel my background task",
"cancel that background task",
"cancel the running background task",
)
private val pauseContinuousPhrases = setOf(
"pause",
"pause continuous listening",
"pause hands free listening",
)
private val resumeContinuousPhrases = setOf(
"resume",
"resume continuous listening",
"resume hands free listening",
)
private val repeatBackgroundAnswerPhrases = setOf(
"repeat that",
"repeat the background answer",
"repeat the last background answer",
"repeat that background answer",
)
private val newChatPhrases = setOf(
"new chat",
"start a new chat",
"open a new chat",
"create a new chat",
)
fun interpretFinalTranscript(
rawTranscript: String,
context: VoiceCommandContext,
): VoiceCommandAction? {
val phrase = normalize(rawTranscript)
if (phrase.isEmpty()) return null
return when {
context.backgroundTaskActive && phrase in cancelBackgroundTaskPhrases ->
VoiceCommandAction.CancelBackgroundTask
context.responseActive && phrase in stopResponsePhrases ->
VoiceCommandAction.StopResponse
context.continuousModeSelected &&
context.continuousListeningActive &&
phrase in pauseContinuousPhrases -> VoiceCommandAction.PauseContinuousListening
context.continuousModeSelected &&
context.continuousListeningPaused &&
phrase in resumeContinuousPhrases -> VoiceCommandAction.ResumeContinuousListening
context.backgroundAnswerAvailable && phrase in repeatBackgroundAnswerPhrases ->
VoiceCommandAction.RepeatBackgroundAnswer
context.canStartNewChat && phrase in newChatPhrases -> VoiceCommandAction.StartNewChat
else -> null
}
}
private fun normalize(raw: String): String = raw
.lowercase(Locale.ROOT)
.replace(Regex("[\\p{Punct}\\p{P}]"), " ")
.replace(Regex("\\s+"), " ")
.trim()
}
@@ -0,0 +1,121 @@
package com.hermesandroid.relay.data
import androidx.datastore.core.DataStore
import androidx.datastore.preferences.core.PreferenceDataStoreFactory
import androidx.datastore.preferences.core.Preferences
import androidx.datastore.preferences.core.edit
import androidx.datastore.preferences.core.stringPreferencesKey
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.cancel
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertNull
import org.junit.Before
import org.junit.Rule
import org.junit.Test
import org.junit.rules.TemporaryFolder
class ChatTurnCheckpointStoreTest {
@get:Rule
val tempFolder = TemporaryFolder()
private lateinit var scope: CoroutineScope
private lateinit var dataStore: DataStore<Preferences>
private lateinit var store: DataStoreChatTurnCheckpointStore
private var now = 10_000L
@Before
fun setUp() {
scope = CoroutineScope(Dispatchers.IO + SupervisorJob())
val file = tempFolder.newFile("chat_checkpoint.preferences_pb")
file.delete()
dataStore = PreferenceDataStoreFactory.create(
scope = scope,
produceFile = { file },
)
store = DataStoreChatTurnCheckpointStore(dataStore) { now }
}
@After
fun tearDown() {
scope.cancel()
}
@Test
fun fullRichTurn_roundTrips() = runTest {
val checkpoint = sampleCheckpoint()
store.write(checkpoint)
assertEquals(checkpoint, store.read())
}
@Test
fun corruptJson_isDiscarded() = runTest {
dataStore.edit { preferences ->
preferences[stringPreferencesKey("chat_inflight_turn_checkpoint_v1")] = "{broken"
}
assertNull(store.read())
assertNull(store.read())
}
@Test
fun staleCheckpoint_isDiscarded() = runTest {
val checkpoint = sampleCheckpoint().copy(updatedAt = now)
store.write(checkpoint)
now += ChatTurnCheckpoint.MAX_AGE_MS + 1L
assertNull(store.read())
assertNull(store.read())
}
private fun sampleCheckpoint() = ChatTurnCheckpoint(
contextKey = "connection-a/profile-default",
sessionId = "stored-42",
liveSessionId = "live-42",
transport = "gateway",
user = ChatTurnUserCheckpoint("user-1", "research this", 1_000L),
assistant = ChatTurnAssistantCheckpoint(
id = "assistant-1",
content = "Working on it",
timestamp = 1_001L,
thinkingContent = "I should inspect the source",
isThinkingStreaming = true,
agentName = "Hermes",
toolCalls = listOf(
ChatTurnToolCheckpoint(
id = "tool-1",
name = "terminal",
isComplete = false,
startedAt = 1_002L,
),
),
backgroundTask = ChatTurnBackgroundTaskCheckpoint(
id = "run-1",
title = "Research",
tier = "durable",
phase = BackgroundTaskPhase.RUNNING.name,
statusLine = "Checking sources",
startedAt = 1_003L,
),
),
turnStatus = "Running terminal",
priorUserMessageCount = 3,
baselineAssistantCount = 3,
pendingAsk = ChatTurnAskCheckpoint(
kind = "APPROVAL",
text = "Allow command?",
timeoutSeconds = 0,
messageId = "ask-1",
cardKey = "approval-1",
receivedAt = 1_004L,
),
startedAt = 1_001L,
updatedAt = now,
)
}
@@ -0,0 +1,106 @@
package com.hermesandroid.relay.data
import org.junit.Assert.assertEquals
import org.junit.Assert.assertNull
import org.junit.Test
class HermesProcessNotificationTest {
@Test
fun completionEnvelopeParsesWithoutChangingCanonicalMessageRole() {
val content = """
[IMPORTANT: Background process 42 completed normally (exit code 0).
Command: ./gradlew test
Output:
BUILD SUCCESSFUL]
""".trimIndent()
val message = ChatMessage(
id = "server-user-row",
role = MessageRole.USER,
content = content,
timestamp = 1L,
)
val parsed = message.hermesProcessNotificationOrNull()
assertEquals(MessageRole.USER, message.role)
assertEquals("42", parsed?.processId)
assertEquals(
"Background process 42 completed normally (exit code 0).",
parsed?.headline,
)
assertEquals(
"Command: ./gradlew test\nOutput:\nBUILD SUCCESSFUL",
parsed?.detail,
)
}
@Test
fun watchMatchEnvelopePreservesMultilineDetail() {
val content = """
[IMPORTANT: Background process proc-7 matched watch pattern "ready".
Command: python server.py
Matched output:
Server ready
Listening on 127.0.0.1]
""".trimIndent()
val parsed = HermesProcessNotificationParser.parse(content)
assertEquals("proc-7", parsed?.processId)
assertEquals(
"Command: python server.py\nMatched output:\nServer ready\nListening on 127.0.0.1",
parsed?.detail,
)
}
@Test
fun compactEnvelopeWithoutOutputStillParses() {
val parsed = HermesProcessNotificationParser.parse(
"[IMPORTANT: Background process 123 finished]",
)
assertEquals("123", parsed?.processId)
assertEquals("Background process 123 finished", parsed?.headline)
assertNull(parsed?.detail)
}
@Test
fun nonProcessImportantMessageIsNotClaimed() {
assertNull(
HermesProcessNotificationParser.parse(
"[IMPORTANT: Process watch was disabled]",
),
)
}
@Test
fun markerEmbeddedInHumanTextIsNotClaimed() {
assertNull(
HermesProcessNotificationParser.parse(
"Hermes said [IMPORTANT: Background process 12 finished] yesterday",
),
)
}
@Test
fun malformedEnvelopeWithoutStatusIsNotClaimed() {
assertNull(
HermesProcessNotificationParser.parse(
"[IMPORTANT: Background process 12]",
),
)
}
@Test
fun assistantCopyOfProcessEnvelopeUsesNormalPresentation() {
val message = ChatMessage(
id = "assistant-copy",
role = MessageRole.ASSISTANT,
content = "[IMPORTANT: Background process 123 finished]",
timestamp = 1L,
)
assertNull(message.hermesProcessNotificationOrNull())
}
}
@@ -0,0 +1,172 @@
package com.hermesandroid.relay.data
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Test
class VoiceModePresetTest {
private val current = VoiceModePresetState(
voiceSettings = VoiceSettings(
engineMode = VoiceEngineMode.HermesVoiceOutput.storageValue,
audioRoute = VoiceAudioRoute.Relay.storageValue,
interactionMode = "tap",
silenceThresholdMs = 3000L,
realtimeTraceDetails = false,
realtimePersistentSession = false,
realtimeModel = "custom-realtime-model",
realtimeVoice = "custom-realtime-voice",
enhancedVoice = "custom-output-voice",
enhancedModel = "custom-output-model",
enhancedAudioTags = true,
enhancedPersona = "Warm and precise",
enhancedLanguage = "en-US",
),
bargeInPreferences = BargeInPreferences(
enabled = true,
sensitivity = BargeInSensitivity.High,
resumeAfterInterruption = false,
),
promotion = VoicePresetPromotionSettings(
enabled = true,
promoteAfterMs = 42000,
backgroundDefaultMode = "foreground",
spokenHandoff = true,
progressSpokenAfterMs = 32000,
progressRepeatMs = 123000,
resultDelivery = "notify_then_speak",
maxBackgroundRuns = 4,
),
)
@Test
fun handsFreeMapsContinuousListeningAndPreservesExperimentalBargeInChoice() {
val source = current.copy(
bargeInPreferences = BargeInPreferences(
enabled = false,
sensitivity = BargeInSensitivity.High,
resumeAfterInterruption = false,
),
)
val target = VoiceModePreset.HandsFree.applyTo(source)
assertEquals("continuous", target.voiceSettings.interactionMode)
assertEquals(1250L, target.voiceSettings.silenceThresholdMs)
assertTrue(target.voiceSettings.realtimeTraceDetails)
assertTrue(target.voiceSettings.realtimePersistentSession)
assertFalse(target.bargeInPreferences.enabled)
assertEquals(BargeInSensitivity.High, target.bargeInPreferences.sensitivity)
assertFalse(target.bargeInPreferences.resumeAfterInterruption)
assertEquals(6000, target.promotion?.promoteAfterMs)
assertTrue(target.promotion?.spokenHandoff == true)
assertEquals(15000, target.promotion?.progressSpokenAfterMs)
assertEquals(90000, target.promotion?.progressRepeatMs)
assertEquals("speak_verbatim", target.promotion?.resultDelivery)
}
@Test
fun lowLatencyMapsShortestSilenceAndFastPromotion() {
val target = VoiceModePreset.LowLatency.applyTo(current)
assertEquals("tap", target.voiceSettings.interactionMode)
assertEquals(750L, target.voiceSettings.silenceThresholdMs)
assertFalse(target.voiceSettings.realtimeTraceDetails)
assertTrue(target.voiceSettings.realtimePersistentSession)
assertFalse(target.bargeInPreferences.enabled)
assertEquals(2500, target.promotion?.promoteAfterMs)
assertFalse(target.promotion?.spokenHandoff == true)
assertEquals(0, target.promotion?.progressSpokenAfterMs)
assertEquals("speak_when_idle", target.promotion?.resultDelivery)
}
@Test
fun carefulToolsKeepsRunsForegroundAndResultsExact() {
val target = VoiceModePreset.CarefulTools.applyTo(current)
assertEquals("hold", target.voiceSettings.interactionMode)
assertEquals(1750L, target.voiceSettings.silenceThresholdMs)
assertTrue(target.voiceSettings.realtimeTraceDetails)
assertFalse(target.bargeInPreferences.enabled)
assertFalse(target.promotion?.enabled == true)
assertEquals("foreground", target.promotion?.backgroundDefaultMode)
assertEquals("speak_verbatim", target.promotion?.resultDelivery)
}
@Test
fun quietVisualOnlyLeavesShortReplyBehaviorExplicitlyOutOfScope() {
val target = VoiceModePreset.QuietVisualOnly.applyTo(current)
assertEquals("tap", target.voiceSettings.interactionMode)
assertEquals(1250L, target.voiceSettings.silenceThresholdMs)
assertTrue(target.voiceSettings.realtimeTraceDetails)
assertFalse(target.bargeInPreferences.enabled)
assertTrue(target.promotion?.enabled == true)
assertFalse(target.promotion?.spokenHandoff == true)
assertEquals(0, target.promotion?.progressSpokenAfterMs)
assertEquals("visual_only", target.promotion?.resultDelivery)
}
@Test
fun detectorRecognizesEveryFullyAppliedPreset() {
VoiceModePreset.entries.forEach { preset ->
assertEquals(preset, detectVoiceModePreset(preset.applyTo(current)))
}
}
@Test
fun manualDivergenceReportsCustom() {
val handsFree = VoiceModePreset.HandsFree.applyTo(current)
val diverged = handsFree.copy(
voiceSettings = handsFree.voiceSettings.copy(interactionMode = "tap"),
)
assertNull(detectVoiceModePreset(diverged))
}
@Test
fun missingPromotionSnapshotNeverClaimsAnActivePreset() {
val noPromotion = current.copy(promotion = null)
assertNull(detectVoiceModePreset(VoiceModePreset.HandsFree.applyTo(noPromotion)))
}
@Test
fun presetsPreserveVoiceIdentityRoutingAndConcurrency() {
VoiceModePreset.entries.forEach { preset ->
val target = preset.applyTo(current)
assertEquals(current.voiceSettings.engineMode, target.voiceSettings.engineMode)
assertEquals(current.voiceSettings.audioRoute, target.voiceSettings.audioRoute)
assertEquals(current.voiceSettings.realtimeModel, target.voiceSettings.realtimeModel)
assertEquals(current.voiceSettings.realtimeVoice, target.voiceSettings.realtimeVoice)
assertEquals(current.voiceSettings.enhancedVoice, target.voiceSettings.enhancedVoice)
assertEquals(current.voiceSettings.enhancedModel, target.voiceSettings.enhancedModel)
assertEquals(
current.voiceSettings.enhancedAudioTags,
target.voiceSettings.enhancedAudioTags,
)
assertEquals(current.voiceSettings.enhancedPersona, target.voiceSettings.enhancedPersona)
assertEquals(current.voiceSettings.enhancedLanguage, target.voiceSettings.enhancedLanguage)
assertEquals(
current.promotion?.maxBackgroundRuns,
target.promotion?.maxBackgroundRuns,
)
}
}
@Test
fun disabledBargeInPresetsPreserveHiddenSensitivityPreferences() {
listOf(
VoiceModePreset.LowLatency,
VoiceModePreset.CarefulTools,
VoiceModePreset.QuietVisualOnly,
).forEach { preset ->
val target = preset.applyTo(current)
assertEquals(BargeInSensitivity.High, target.bargeInPreferences.sensitivity)
assertFalse(target.bargeInPreferences.resumeAfterInterruption)
}
}
}
@@ -0,0 +1,40 @@
package com.hermesandroid.relay.network.relay
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Test
class RealtimeVoiceEventParsingTest {
@Test
fun brokerOkFieldMapsToTheSharedSuccessState() {
val success = Json.parseToJsonElement("""{"ok":true}""").jsonObject
val failure = Json.parseToJsonElement("""{"ok":false}""").jsonObject
assertTrue(realtimeEventSuccess(success) == true)
assertFalse(realtimeEventSuccess(failure) == true)
}
@Test
fun explicitSuccessWinsWhenBothFieldsArePresent() {
val event = Json.parseToJsonElement(
"""{"success":false,"ok":true}""",
).jsonObject
assertEquals(false, realtimeEventSuccess(event))
}
@Test
fun absentOrMalformedFieldsRemainUnknown() {
assertNull(realtimeEventSuccess(Json.parseToJsonElement("{}").jsonObject))
assertNull(
realtimeEventSuccess(
Json.parseToJsonElement("""{"ok":"not-a-boolean"}""").jsonObject,
),
)
}
}
@@ -3,6 +3,10 @@ package com.hermesandroid.relay.network.upstream
import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.ChatSession
import com.hermesandroid.relay.data.ChatTurnAssistantCheckpoint
import com.hermesandroid.relay.data.ChatTurnCheckpoint
import com.hermesandroid.relay.data.ChatTurnToolCheckpoint
import com.hermesandroid.relay.data.ChatTurnUserCheckpoint
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.data.RealtimeTurnTrace
import com.hermesandroid.relay.data.ToolCall
@@ -1450,6 +1454,73 @@ class ChatHandlerTest {
assertFalse(orphan.realtimeTurn!!.syncedToServer)
}
@Test
fun restoreInFlightTurn_restoresThinkingAndToolState_withoutDuplicatingPersistedUser() {
handler.addUserMessage(createUserMessage("old-user", "Earlier"))
handler.addPlaceholderMessage(
ChatMessage(
id = "old-assistant",
role = MessageRole.ASSISTANT,
content = "Earlier answer",
timestamp = 2L,
isStreaming = true,
),
)
handler.onStreamComplete("old-assistant")
// Models history having persisted the pending user before Android
// reopens; the restore must not append a second identical row.
handler.addUserMessage(createUserMessage("server-user", "Run the checks"))
val checkpoint = ChatTurnCheckpoint(
contextKey = "connection/profile",
sessionId = "stored-1",
liveSessionId = "live-1",
transport = "gateway",
user = ChatTurnUserCheckpoint("local-user", "Run the checks", 3L),
assistant = ChatTurnAssistantCheckpoint(
id = "assistant-live",
content = "I am checking",
timestamp = 4L,
thinkingContent = "Inspect the project first",
isThinkingStreaming = true,
toolCalls = listOf(
ChatTurnToolCheckpoint(
id = "tool-1",
name = "terminal",
isComplete = false,
startedAt = 5L,
),
ChatTurnToolCheckpoint(
id = "tool-2",
name = "search",
result = "3 matches",
success = true,
isComplete = true,
startedAt = 6L,
completedAt = 7L,
),
),
),
turnStatus = "Running terminal",
priorUserMessageCount = 1,
baselineAssistantCount = 1,
startedAt = 4L,
updatedAt = 8L,
)
handler.restoreInFlightTurn(checkpoint, upstreamAssistantText = "I am checking the tests")
assertEquals(2, handler.messages.value.count { it.role == MessageRole.USER })
val restored = handler.messages.value.single { it.id == "assistant-live" }
assertEquals("I am checking the tests", restored.content)
assertEquals("Inspect the project first", restored.thinkingContent)
assertTrue(restored.isThinkingStreaming)
assertFalse(restored.toolCalls[0].isComplete)
assertTrue(restored.toolCalls[1].isComplete)
assertEquals(true, restored.toolCalls[1].success)
assertTrue(handler.isStreaming.value)
assertEquals("Running terminal", handler.turnStatus.value)
}
// --- Helper ---
private fun createUserMessage(id: String, content: String) = ChatMessage(
@@ -4,7 +4,9 @@ import com.hermesandroid.relay.network.upstream.models.UsageInfo
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.async
import kotlinx.coroutines.cancel
import kotlinx.coroutines.delay
import kotlinx.coroutines.runBlocking
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
@@ -24,6 +26,8 @@ import okhttp3.mockwebserver.RecordedRequest
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNotNull
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
@@ -51,6 +55,12 @@ class GatewayClientHarness(
var failTicketMint = false
var resumeFails = false
@Volatile
var recoveryRunning = false
@Volatile
var recoveryAssistant = ""
@Volatile
var steerStatus = "queued"
@@ -109,9 +119,47 @@ class GatewayClientHarness(
}
"session.resume" ->
if (resumeFails) null
else buildJsonObject { put("session_id", "live-resumed") }
else recoveryPayload("live-resumed")
"session.activate" -> recoveryPayload(
(params["session_id"] as? JsonPrimitive)?.contentOrNull ?: "live-activated",
)
"prompt.submit" -> buildJsonObject { put("ok", true) }
"session.interrupt" -> buildJsonObject { put("ok", true) }
"process.list" -> buildJsonObject {
put(
"processes",
json.parseToJsonElement(
"""
[
{
"session_id": "proc-17",
"command": "./gradlew test",
"cwd": "/workspace/app",
"pid": 4812,
"started_at": "2026-07-10T09:30:00",
"uptime_seconds": 42,
"status": "running",
"output_preview": "running tests",
"output_tail": "running tests\n42 tests completed",
"notify_on_complete": true,
"session_scoped": true,
"watch_patterns": ["BUILD SUCCESSFUL"],
"watch_hit": false
},
{
"session_id": "proc-18",
"command": "npm run lint",
"uptime_seconds": 7,
"status": "exited",
"exit_code": 1,
"detached": true
}
]
""".trimIndent(),
),
)
}
"process.kill" -> buildJsonObject { put("status", "killed") }
"session.steer" -> buildJsonObject {
put("status", steerStatus)
put("text", (params["text"] as? JsonPrimitive)?.contentOrNull ?: "")
@@ -195,6 +243,19 @@ class GatewayClientHarness(
private val autoRespondEnabled = autoRespond
private fun recoveryPayload(sessionId: String): JsonObject = buildJsonObject {
put("session_id", sessionId)
put("running", recoveryRunning)
put("status", if (recoveryRunning) "streaming" else "idle")
if (recoveryRunning) {
put("inflight", buildJsonObject {
put("user", "research this")
put("assistant", recoveryAssistant)
put("streaming", true)
})
}
}
init {
server.dispatcher = object : Dispatcher() {
override fun dispatch(request: RecordedRequest): MockResponse {
@@ -245,13 +306,16 @@ class GatewayClientHarness(
fun awaitPendingAck(): PendingAck =
pendingAcks.poll(5, TimeUnit.SECONDS) ?: error("suppressed ack never captured")
/** Release a withheld ack with a generic success result. */
fun releaseAck(ack: PendingAck) {
/** Release a withheld ack with a caller-supplied or generic success result. */
fun releaseAck(
ack: PendingAck,
result: JsonObject = buildJsonObject { put("ok", true) },
) {
ack.ws.send(
buildJsonObject {
put("jsonrpc", "2.0")
put("id", ack.id)
put("result", buildJsonObject { put("ok", true) })
put("result", result)
}.toString(),
)
}
@@ -293,11 +357,14 @@ class GatewayChatClientTest {
private var unsupportedMarked = false
private class Recorder {
val starts = AtomicInteger(0)
val textDeltas = ConcurrentLinkedQueue<String>()
val thinkingDeltas = ConcurrentLinkedQueue<String>()
val sessionIds = ConcurrentLinkedQueue<String>()
val errors = ConcurrentLinkedQueue<String>()
val interactions = ConcurrentLinkedQueue<GatewayAsk>()
val toolStarts = ConcurrentLinkedQueue<Pair<String, String>>()
val toolDone = ConcurrentLinkedQueue<Pair<String, String?>>()
// ConcurrentLinkedQueue rejects nulls — unnamed generating events store "".
val toolGenerating = ConcurrentLinkedQueue<String>()
@@ -308,10 +375,11 @@ class GatewayChatClientTest {
val callbacks = GatewayTurnCallbacks(
onSessionId = { sessionIds += it },
onStart = { starts.incrementAndGet() },
onTextDelta = { textDeltas += it },
onThinkingDelta = { thinkingDeltas += it },
onToolCallStart = { _, _ -> },
onToolCallDone = { _, _ -> },
onToolCallStart = { id, name -> toolStarts += id to name },
onToolCallDone = { id, result -> toolDone += id to result },
onToolCallFailed = { _, _ -> },
onTurnComplete = { },
onComplete = { completeLatch.countDown() },
@@ -424,6 +492,187 @@ class GatewayChatClientTest {
assertTrue(r.preflightFailures.isEmpty())
}
@Test
fun `unsolicited assistant turn for resumed session streams without sendTurn`() = runBlocking {
val r = Recorder()
val registrations = ConcurrentLinkedQueue<String>()
val processEvents = ConcurrentLinkedQueue<GatewayProcessEvent>()
val processEventLatch = CountDownLatch(1)
client.setProcessEventListener {
processEvents += it
processEventLatch.countDown()
}
client.setUnsolicitedTurnProvider { storedSessionId ->
registrations += storedSessionId
GatewayInboundTurnRegistration(
callbacks = r.callbacks,
onHandle = { true },
)
}
assertTrue(client.prewarmAwait("stored-session"))
val serverWs = harness.awaitServerSocket()
serverWs.send(harness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
harness.eventFrame(
"message.delta",
buildJsonObject { put("text", "Background task finished.") },
"live-resumed",
),
)
serverWs.send(
harness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Background task finished.") },
"live-resumed",
),
)
assertTrue("unsolicited turn never completed", r.completeLatch.await(5, TimeUnit.SECONDS))
assertTrue("turn completion did not invalidate process inventory", processEventLatch.await(5, TimeUnit.SECONDS))
assertEquals(listOf("stored-session"), registrations.toList())
assertEquals(1, r.starts.get())
assertEquals(listOf("Background task finished."), r.textDeltas.toList())
assertEquals(
listOf(GatewayProcessEvent.Invalidated(GatewayProcessEvent.Trigger.MESSAGE_COMPLETE)),
processEvents.toList(),
)
assertFalse(harness.rpcLog.any { it.first == "prompt.submit" })
}
@Test
fun `unsolicited starts without exact live session are ignored`() = runBlocking {
val r = Recorder()
val registrations = AtomicInteger(0)
client.setUnsolicitedTurnProvider {
registrations.incrementAndGet()
GatewayInboundTurnRegistration(r.callbacks) { true }
}
assertTrue(client.prewarmAwait("stored-session"))
val serverWs = harness.awaitServerSocket()
serverWs.send(harness.eventFrame("message.start", null, null))
serverWs.send(harness.eventFrame("message.start", null, "someone-else"))
assertFalse("foreign turn was accepted", r.completeLatch.await(300, TimeUnit.MILLISECONDS))
assertEquals(0, registrations.get())
assertEquals(0, r.starts.get())
}
@Test
fun `cold prewarm reports resumed stored session once`() = runBlocking {
val resumedSessions = ConcurrentLinkedQueue<String>()
val resumedLatch = CountDownLatch(1)
client.setColdPrewarmSessionReadyListener { storedSessionId ->
resumedSessions += storedSessionId
resumedLatch.countDown()
}
assertTrue(client.prewarmAwait("stored-session"))
assertTrue("cold resume was not reported", resumedLatch.await(5, TimeUnit.SECONDS))
assertTrue(client.prewarmAwait("stored-session"))
Thread.sleep(100)
assertEquals(listOf("stored-session"), resumedSessions.toList())
}
@Test
fun `newer prewarm selection wins when an older resume completes late`() = runBlocking {
harness.suppressAckMethods += "session.resume"
val old = async(Dispatchers.IO) { client.prewarmAwait("old-session") }
val oldAck = harness.awaitPendingAck()
// Starting the newer request advances the desired-session generation
// even though it must wait for the older request's connect mutex.
val newer = async(Dispatchers.IO) { client.prewarmAwait("new-session") }
delay(100)
harness.releaseAck(
oldAck,
buildJsonObject { put("session_id", "live-old") },
)
assertFalse(old.await())
val newerAck = harness.awaitPendingAck()
harness.releaseAck(
newerAck,
buildJsonObject { put("session_id", "live-new") },
)
assertTrue(newer.await())
client.listProcesses().getOrThrow()
val params = harness.awaitRpc("process.list")
assertEquals("live-new", (params["session_id"] as? JsonPrimitive)?.contentOrNull)
}
@Test
fun `same stored session id is resumed again when profile namespace changes`() = runBlocking {
var profile = "profile-a"
client.sessionProfileProvider = { profile }
assertTrue(client.prewarmAwait("same-stored-id"))
profile = "profile-b"
assertTrue(client.prewarmAwait("same-stored-id"))
val resumes = harness.awaitRpcCount("session.resume", 2)
assertEquals(
"profile-a",
(resumes[0]["profile"] as? JsonPrimitive)?.contentOrNull,
)
assertEquals(
"profile-b",
(resumes[1]["profile"] as? JsonPrimitive)?.contentOrNull,
)
}
@Test
fun `unsolicited error clears turn so the next unsolicited response can arrive`() = runBlocking {
val recorders = ConcurrentLinkedQueue<Recorder>()
client.setUnsolicitedTurnProvider {
val recorder = Recorder()
recorders += recorder
GatewayInboundTurnRegistration(recorder.callbacks) { true }
}
assertTrue(client.prewarmAwait("stored-session"))
val serverWs = harness.awaitServerSocket()
serverWs.send(harness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
harness.eventFrame(
"error",
buildJsonObject { put("message", "first failed") },
"live-resumed",
),
)
val first = awaitRecorder(recorders, 1)
assertTrue(first.completeLatch.await(5, TimeUnit.SECONDS))
assertEquals(listOf("first failed"), first.errors.toList())
serverWs.send(harness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
harness.eventFrame(
"message.complete",
buildJsonObject { put("text", "second worked") },
"live-resumed",
),
)
val second = awaitRecorder(recorders, 2)
assertTrue(second.completeLatch.await(5, TimeUnit.SECONDS))
assertEquals(listOf("second worked"), second.textDeltas.toList())
}
private fun awaitRecorder(recorders: ConcurrentLinkedQueue<Recorder>, count: Int): Recorder {
val deadline = System.currentTimeMillis() + 5_000
while (System.currentTimeMillis() < deadline) {
if (recorders.size >= count) return recorders.elementAt(count - 1)
Thread.sleep(20)
}
error("recorder $count was never registered")
}
@Test
fun `model options refresh flag rides gateway rpc only on explicit refresh`() = runBlocking {
val normal = client.modelOptions().getOrThrow()
@@ -437,6 +686,154 @@ class GatewayChatClientTest {
assertTrue((refreshParams["refresh"] as? JsonPrimitive)?.booleanOrNull == true)
}
@Test
fun `process list uses live session id and parses typed snapshot`() = runBlocking {
assertTrue(client.prewarmAwait("stored-session"))
val processes = client.listProcesses().getOrThrow()
val params = harness.awaitRpc("process.list")
assertEquals("live-resumed", (params["session_id"] as? JsonPrimitive)?.contentOrNull)
assertEquals(GatewayProcessCapability.Supported, client.processCapability.value)
assertEquals(2, processes.size)
assertEquals(
GatewayProcess(
id = "proc-17",
command = "./gradlew test",
cwd = "/workspace/app",
pid = 4812L,
startedAt = "2026-07-10T09:30:00",
uptimeSeconds = 42L,
status = "running",
outputPreview = "running tests",
outputTail = "running tests\n42 tests completed",
notifyOnComplete = true,
sessionScoped = true,
watchPatterns = listOf("BUILD SUCCESSFUL"),
),
processes[0],
)
assertTrue(processes[0].isRunning)
assertEquals(1, processes[1].exitCode)
assertTrue(processes[1].detached)
assertFalse(processes[1].isRunning)
}
@Test
fun `process kill uses exact live session and process id`() = runBlocking {
assertTrue(client.prewarmAwait("stored-session"))
assertTrue(client.killProcess("proc-17").isSuccess)
val params = harness.awaitRpc("process.kill")
assertEquals("live-resumed", (params["session_id"] as? JsonPrimitive)?.contentOrNull)
assertEquals("proc-17", (params["process_id"] as? JsonPrimitive)?.contentOrNull)
assertEquals(GatewayProcessCapability.Supported, client.processCapability.value)
}
@Test
fun `process method not found disables repeat probes for current socket`() = runBlocking {
harness.methodNotFound.add("process.list")
assertTrue(client.prewarmAwait("stored-session"))
assertTrue(client.listProcesses().isFailure)
assertEquals(GatewayProcessCapability.Unsupported, client.processCapability.value)
assertEquals(1, harness.rpcLog.count { it.first == "process.list" })
assertTrue(client.listProcesses().isFailure)
assertEquals(1, harness.rpcLog.count { it.first == "process.list" })
}
@Test
fun `process events bypass active turn gate but require exact live session`() = runBlocking {
val events = ConcurrentLinkedQueue<GatewayProcessEvent>()
val eventLatch = CountDownLatch(5)
client.setProcessEventListener {
events += it
eventLatch.countDown()
}
assertTrue(client.prewarmAwait("stored-session"))
val serverWs = harness.awaitServerSocket()
serverWs.send(
harness.eventFrame(
"agent.terminal.output",
buildJsonObject { put("process_id", "foreign"); put("chunk", "do not leak") },
"someone-else",
),
)
serverWs.send(
harness.eventFrame(
"tool.complete",
buildJsonObject { put("name", "browser"); put("tool_id", "tool-ignored") },
"live-resumed",
),
)
serverWs.send(
harness.eventFrame(
"tool.complete",
buildJsonObject { put("name", "terminal"); put("tool_id", "tool-1") },
"live-resumed",
),
)
serverWs.send(
harness.eventFrame(
"status.update",
buildJsonObject { put("kind", "process"); put("text", "process proc-17 completed") },
"live-resumed",
),
)
serverWs.send(
harness.eventFrame(
"agent.terminal.output",
buildJsonObject { put("process_id", "proc-17"); put("chunk", "BUILD SUCCESSFUL\n") },
"live-resumed",
),
)
serverWs.send(
harness.eventFrame(
"terminal.close",
buildJsonObject { put("process_id", "proc-17") },
"live-resumed",
),
)
serverWs.send(
harness.eventFrame(
"message.complete",
buildJsonObject { put("text", "foreign turn") },
"someone-else",
),
)
serverWs.send(
harness.eventFrame(
"message.complete",
buildJsonObject { put("text", "missing session id") },
null,
),
)
// Upstream can omit tool lifecycle events for a background launch;
// every exact-session turn completion is therefore a list fallback.
serverWs.send(
harness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Started as proc-17") },
"live-resumed",
),
)
assertTrue("process events were dropped without an active turn", eventLatch.await(5, TimeUnit.SECONDS))
assertEquals(
listOf(
GatewayProcessEvent.Invalidated(GatewayProcessEvent.Trigger.TOOL_COMPLETE),
GatewayProcessEvent.Invalidated(GatewayProcessEvent.Trigger.STATUS_UPDATE),
GatewayProcessEvent.Output("proc-17", "BUILD SUCCESSFUL\n"),
GatewayProcessEvent.TerminalClosed("proc-17"),
GatewayProcessEvent.Invalidated(GatewayProcessEvent.Trigger.MESSAGE_COMPLETE),
),
events.toList(),
)
}
@Test
fun `foreign session events are dropped`() {
val r = Recorder()
@@ -1130,6 +1527,147 @@ class GatewayChatClientTest {
assertEquals(1, harness.rpcLog.count { it.first == "prompt.submit" })
}
@Test
fun `recoverTurn activates exact live session and continues deltas and tool events`() {
harness.recoveryRunning = true
harness.recoveryAssistant = "partial answer"
val recorder = Recorder()
val recovery = runBlocking {
client.recoverTurn(
storedId = "stored-42",
preferredLiveId = "live-original",
callbacks = recorder.callbacks,
).getOrThrow()
}
assertTrue(recovery.running)
assertEquals("live-original", recovery.liveSessionId)
assertEquals("partial answer", recovery.inflight?.assistant)
assertNotNull(recovery.handle)
assertEquals(1, harness.rpcLog.count { it.first == "session.activate" })
assertEquals(0, harness.rpcLog.count { it.first == "session.resume" })
val serverWs = harness.awaitServerSocket()
serverWs.send(
harness.eventFrame(
"reasoning.delta",
buildJsonObject { put("text", "still thinking") },
"live-original",
),
)
serverWs.send(
harness.eventFrame(
"tool.start",
buildJsonObject {
put("tool_id", "tool-1")
put("name", "terminal")
},
"live-original",
),
)
serverWs.send(
harness.eventFrame(
"tool.complete",
buildJsonObject {
put("tool_id", "tool-1")
put("name", "terminal")
put("summary", "tests passed")
},
"live-original",
),
)
serverWs.send(
harness.eventFrame(
"message.delta",
buildJsonObject { put("text", " final") },
"live-original",
),
)
serverWs.send(
harness.eventFrame(
"message.complete",
buildJsonObject { put("text", "partial answer final") },
"live-original",
),
)
assertTrue(recorder.completeLatch.await(5, TimeUnit.SECONDS))
assertEquals(listOf("still thinking"), recorder.thinkingDeltas.toList())
assertEquals(listOf("tool-1" to "terminal"), recorder.toolStarts.toList())
assertEquals(listOf("tool-1" to "tests passed"), recorder.toolDone.toList())
assertEquals(listOf(" final"), recorder.textDeltas.toList())
}
@Test
fun `recoverTurn falls back to durable resume when activate is unsupported`() {
harness.methodNotFound += "session.activate"
harness.recoveryRunning = true
val recorder = Recorder()
val recovery = runBlocking {
client.recoverTurn(
storedId = "stored-42",
preferredLiveId = "expired-live-id",
callbacks = recorder.callbacks,
).getOrThrow()
}
assertTrue(recovery.running)
assertEquals("live-resumed", recovery.liveSessionId)
assertNotNull(recovery.handle)
assertEquals(1, harness.rpcLog.count { it.first == "session.activate" })
assertEquals(1, harness.rpcLog.count { it.first == "session.resume" })
val serverWs = harness.awaitServerSocket()
serverWs.send(
harness.eventFrame(
"message.complete",
buildJsonObject { put("text", "recovered") },
"live-resumed",
),
)
assertTrue(recorder.completeLatch.await(5, TimeUnit.SECONDS))
assertEquals(listOf("recovered"), recorder.textDeltas.toList())
}
@Test
fun `recoverTurn returns no handle for an already-settled session`() {
harness.recoveryRunning = false
val recovery = runBlocking {
client.recoverTurn(
storedId = "stored-42",
preferredLiveId = "live-original",
callbacks = Recorder().callbacks,
).getOrThrow()
}
assertFalse(recovery.running)
assertEquals("idle", recovery.status)
assertNull(recovery.handle)
assertFalse(client.hasActiveTurn())
assertTrue(harness.rpcLog.none { it.first == "session.interrupt" })
}
@Test
fun `detaching recovered handle does not interrupt server turn`() {
harness.recoveryRunning = true
val recovery = runBlocking {
client.recoverTurn(
storedId = "stored-42",
preferredLiveId = "live-original",
callbacks = Recorder().callbacks,
).getOrThrow()
}
recovery.handle!!.detach()
Thread.sleep(100)
assertFalse(client.hasActiveTurn())
assertTrue(harness.rpcLog.none { it.first == "session.interrupt" })
}
@Test
fun `idle watchdog does not fire while events keep arriving slowly`() {
rebuildClient(turnIdleTimeoutMs = 1_000L)
@@ -26,6 +26,7 @@ class GatewayEventMapperTest {
val subagentEvents = mutableListOf<GatewaySubagentEvent>()
val interactions = mutableListOf<GatewayAsk>()
val sessionIds = mutableListOf<String>()
var starts = 0
var turnCompletes = 0
var completes = 0
var usage: UsageInfo? = null
@@ -34,6 +35,7 @@ class GatewayEventMapperTest {
val callbacks = GatewayTurnCallbacks(
onSessionId = { sessionIds += it },
onStart = { starts++ },
onTextDelta = { textDeltas += it },
onThinkingDelta = { thinkingDeltas += it },
onToolCallStart = { id, name -> toolStarts += id to name },
@@ -52,7 +54,10 @@ class GatewayEventMapperTest {
private fun obj(jsonText: String): JsonObject =
Json.parseToJsonElement(jsonText) as JsonObject
private fun mapperWith(recorder: Recorder) = GatewayEventMapper(recorder.callbacks)
private fun mapperWith(
recorder: Recorder,
dedupeAdjacentMessageStarts: Boolean = false,
) = GatewayEventMapper(recorder.callbacks, dedupeAdjacentMessageStarts)
// --- The feature: live thinking ---
@@ -118,6 +123,19 @@ class GatewayEventMapperTest {
assertEquals(0, r.turnCompletes)
mapper.onEvent("message.start", null)
assertEquals(1, r.turnCompletes)
assertEquals(2, r.starts)
}
@Test
fun `adjacent duplicate message starts are one boundary`() {
val r = Recorder()
val mapper = mapperWith(r, dedupeAdjacentMessageStarts = true)
mapper.onEvent("message.start", null)
mapper.onEvent("message.start", null)
assertEquals(1, r.starts)
assertEquals(0, r.turnCompletes)
}
@Test
@@ -0,0 +1,119 @@
package com.hermesandroid.relay.ui.components
import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.data.AttachmentState
import org.junit.Assert.assertEquals
import org.junit.Test
class AttachmentGalleryLayoutTest {
@Test
fun `one image keeps the existing single attachment renderer`() {
val items = attachmentLayoutItems(listOf(image("one.png")))
assertEquals(listOf(AttachmentLayoutItem.Single(0)), items)
}
@Test
fun `two loaded images collapse into one gallery`() {
val items = attachmentLayoutItems(
listOf(image("one.png"), image("two.jpg")),
)
assertEquals(listOf(AttachmentLayoutItem.Gallery(listOf(0, 1))), items)
}
@Test
fun `files split image runs so mixed attachment order is preserved`() {
val items = attachmentLayoutItems(
listOf(
file("notes.pdf", "application/pdf"),
image("one.png"),
file("readme.txt", "text/plain"),
image("two.jpg"),
),
)
assertEquals(
listOf(
AttachmentLayoutItem.Single(0),
AttachmentLayoutItem.Single(1),
AttachmentLayoutItem.Single(2),
AttachmentLayoutItem.Single(3),
),
items,
)
}
@Test
fun `loading and failed images split runs and remain retryable standalone cards`() {
val items = attachmentLayoutItems(
listOf(
image("ready.png"),
image("loading.png", AttachmentState.LOADING),
image("failed.png", AttachmentState.FAILED),
image("ready-too.png"),
),
)
assertEquals(
listOf(
AttachmentLayoutItem.Single(0),
AttachmentLayoutItem.Single(1),
AttachmentLayoutItem.Single(2),
AttachmentLayoutItem.Single(3),
),
items,
)
}
@Test
fun `gallery rows use two columns and leave an odd final image spanning the row`() {
assertEquals(listOf(listOf(0, 1)), galleryRows(2))
assertEquals(listOf(listOf(0, 1), listOf(2)), galleryRows(3))
assertEquals(listOf(listOf(0, 1), listOf(2, 3)), galleryRows(4))
}
@Test
fun `only four bounded thumbnails are composed for a large gallery`() {
assertEquals(listOf(0, 1, 2, 3), galleryPreviewIndices(12))
}
@Test
fun `separate contiguous image runs become separate galleries`() {
val items = attachmentLayoutItems(
listOf(
image("one.png"),
image("two.png"),
file("notes.pdf", "application/pdf"),
image("three.png"),
image("four.png"),
),
)
assertEquals(
listOf(
AttachmentLayoutItem.Gallery(listOf(0, 1)),
AttachmentLayoutItem.Single(2),
AttachmentLayoutItem.Gallery(listOf(3, 4)),
),
items,
)
}
private fun image(
name: String,
state: AttachmentState = AttachmentState.LOADED,
) = Attachment(
contentType = if (name.endsWith(".jpg")) "image/jpeg" else "image/png",
content = "bytes",
fileName = name,
state = state,
)
private fun file(name: String, mime: String) = Attachment(
contentType = mime,
content = "bytes",
fileName = name,
)
}
@@ -0,0 +1,113 @@
package com.hermesandroid.relay.ui.components
import androidx.compose.material3.MaterialTheme
import androidx.compose.ui.test.junit4.v2.createComposeRule
import androidx.compose.ui.test.onAllNodesWithText
import androidx.compose.ui.test.assertIsNotEnabled
import androidx.compose.ui.test.onNodeWithContentDescription
import androidx.compose.ui.test.onNodeWithTag
import androidx.compose.ui.test.onNodeWithText
import androidx.compose.ui.test.performClick
import androidx.compose.ui.test.performTouchInput
import androidx.compose.ui.test.swipeLeft
import com.hermesandroid.relay.data.Attachment
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith
import androidx.test.ext.junit.runners.AndroidJUnit4
import org.robolectric.annotation.Config
import org.robolectric.annotation.GraphicsMode
@RunWith(AndroidJUnit4::class)
@GraphicsMode(GraphicsMode.Mode.NATIVE)
@Config(qualifiers = "w360dp-h720dp-xhdpi")
class AttachmentGalleryUiTest {
@get:Rule
val compose = createComposeRule()
@Test
fun `tile opens its page and the viewer swipes across the image group`() {
val attachments = listOf("one.png", "two.png", "three.png").map { name ->
Attachment(
contentType = "image/png",
content = ONE_PIXEL_PNG,
fileName = name,
)
}
compose.setContent {
MaterialTheme {
AttachmentGallery(attachments = attachments)
}
}
compose.onNodeWithContentDescription("3 image gallery").assertExists()
compose.onNodeWithTag("attachment-gallery-tile-0").performClick()
compose.onNodeWithText("1 / 3").assertExists()
compose.onNodeWithTag("attachment-gallery-pager").performTouchInput { swipeLeft() }
compose.waitUntil(timeoutMillis = 5_000) {
compose.onAllNodesWithText("2 / 3").fetchSemanticsNodes().isNotEmpty()
}
compose.onNodeWithText("2 / 3").assertExists()
}
@Test
fun `tapped tile opens the matching page`() {
val attachments = listOf("one.png", "two.png", "three.png").map { name ->
Attachment(
contentType = "image/png",
content = ONE_PIXEL_PNG,
fileName = name,
)
}
compose.setContent {
MaterialTheme {
AttachmentGallery(attachments = attachments)
}
}
compose.onNodeWithTag("attachment-gallery-tile-1").performClick()
compose.waitUntil(timeoutMillis = 5_000) {
compose.onAllNodesWithText("2 / 3").fetchSemanticsNodes().isNotEmpty()
}
compose.onNodeWithText("2 / 3").assertExists()
}
@Test
fun `sensitive page actions stay disabled until that page is revealed`() {
val attachments = listOf(
Attachment(
contentType = "image/png",
content = ONE_PIXEL_PNG,
fileName = "safe.png",
),
Attachment(
contentType = "image/png",
content = ONE_PIXEL_PNG,
fileName = "sensitive.png",
sensitive = true,
),
)
compose.setContent {
MaterialTheme {
AttachmentGallery(attachments = attachments)
}
}
compose.onNodeWithTag("attachment-gallery-tile-0").performClick()
compose.onNodeWithTag("attachment-gallery-pager").performTouchInput { swipeLeft() }
compose.waitUntil(timeoutMillis = 5_000) {
compose.onAllNodesWithText("2 / 2").fetchSemanticsNodes().isNotEmpty()
}
compose.onNodeWithContentDescription("Share").assertIsNotEnabled()
}
private companion object {
const val ONE_PIXEL_PNG =
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk+A8AAQUBAScY42YAAAAASUVORK5CYII="
}
}
@@ -0,0 +1,36 @@
package com.hermesandroid.relay.ui.components
import com.hermesandroid.relay.data.BackgroundTaskPhase
import com.hermesandroid.relay.data.BackgroundTaskState
import com.hermesandroid.relay.data.ToolCall
import org.junit.Assert.assertEquals
import org.junit.Test
class BackgroundTaskCardTest {
@Test
fun phaseLabelsUseSharedTaskVocabulary() {
assertEquals("Working", backgroundTaskPhaseLabel(BackgroundTaskPhase.RUNNING))
assertEquals("Needs input", backgroundTaskPhaseLabel(BackgroundTaskPhase.WAITING))
assertEquals("Delivering", backgroundTaskPhaseLabel(BackgroundTaskPhase.DELIVERING))
assertEquals("Complete", backgroundTaskPhaseLabel(BackgroundTaskPhase.COMPLETE))
assertEquals("Failed", backgroundTaskPhaseLabel(BackgroundTaskPhase.FAILED))
assertEquals("Cancelled", backgroundTaskPhaseLabel(BackgroundTaskPhase.CANCELLED))
}
@Test
fun metaUsesLargestCompletedCountAndQueuedDepth() {
val task = BackgroundTaskState(
id = "run-1",
title = "Check release",
completedToolCount = 1,
queuedCount = 2,
)
val calls = listOf(
ToolCall(name = "one", args = null, result = null, success = true, isComplete = true),
ToolCall(name = "two", args = null, result = null, success = true, isComplete = true),
)
assertEquals("2 steps · +2 queued", backgroundTaskMeta(task, calls))
}
}
@@ -0,0 +1,52 @@
package com.hermesandroid.relay.ui.components
import org.junit.Assert.assertFalse
import org.junit.Assert.assertTrue
import org.junit.Test
class DotMatrixIndicatorTest {
@Test
fun animationRunsOnlyWhenEveryMotionGateAllowsIt() {
assertTrue(
shouldAnimateDotMatrix(
appAnimationsEnabled = true,
osAnimationsEnabled = true,
touchExplorationEnabled = false,
),
)
}
@Test
fun appAnimationPreferenceCanParkTheIndicator() {
assertFalse(
shouldAnimateDotMatrix(
appAnimationsEnabled = false,
osAnimationsEnabled = true,
touchExplorationEnabled = false,
),
)
}
@Test
fun systemReduceMotionCanParkTheIndicator() {
assertFalse(
shouldAnimateDotMatrix(
appAnimationsEnabled = true,
osAnimationsEnabled = false,
touchExplorationEnabled = false,
),
)
}
@Test
fun talkBackTouchExplorationCanParkTheIndicator() {
assertFalse(
shouldAnimateDotMatrix(
appAnimationsEnabled = true,
osAnimationsEnabled = true,
touchExplorationEnabled = true,
),
)
}
}
@@ -0,0 +1,23 @@
package com.hermesandroid.relay.ui.components
import org.junit.Assert.assertEquals
import org.junit.Test
class GatewayBackgroundProcessesTest {
@Test
fun `elapsed time stays compact from seconds through hours`() {
assertEquals("0s", formatElapsed(-5))
assertEquals("42s", formatElapsed(42))
assertEquals("2m 5s", formatElapsed(125))
assertEquals("1h 2m", formatElapsed(3_725))
}
@Test
fun `terminal output strips ansi control and normalizes progress carriage returns`() {
val raw =
"\u001B[31mFAIL\u001B[0m\r50%\r100%\u001B]0;secret title\u0007\n" +
"ok\t!\u0000"
assertEquals("FAIL\n50%\n100%\nok\t!", sanitizeTerminalText(raw))
}
}
@@ -0,0 +1,147 @@
package com.hermesandroid.relay.ui.components
import org.junit.Assert.assertEquals
import org.junit.Assert.assertTrue
import org.junit.Test
class MarkdownStreamingParserTest {
@Test
fun activeParagraph_remainsRawUntilAStableBlockBoundary() {
assertEquals(
listOf(StreamingMarkdownBlock.Text("A paragraph still arriving")),
parseStreamingMarkdownBlocks("A paragraph still arriving"),
)
// A single newline is a Markdown soft break, not a stable block split.
assertEquals(
listOf(StreamingMarkdownBlock.Text("line one\nline two")),
parseStreamingMarkdownBlocks("line one\nline two"),
)
}
@Test
fun blankLine_promotesSettledPrefixToRealMarkdown() {
assertEquals(
listOf(
StreamingMarkdownBlock.Markdown("## Stable heading"),
StreamingMarkdownBlock.Text("- one\n- two\n\nTail still arriving"),
),
parseStreamingMarkdownBlocks(
"## Stable heading\n\n- one\n- two\n\nTail still arriving",
),
)
}
@Test
fun openFence_keepsBlankLinesInsideTheActiveCodeBlock() {
assertEquals(
listOf(
StreamingMarkdownBlock.Markdown("Intro"),
StreamingMarkdownBlock.Code(
language = "kotlin",
code = "val first = 1\n\nval second = 2",
),
),
parseStreamingMarkdownBlocks(
"Intro\n\n```kotlin\nval first = 1\n\nval second = 2",
),
)
}
@Test
fun closedFence_staysOnTheStreamingCodeSurfaceUntilFinal() {
val blocks = parseStreamingMarkdownBlocks(
"```kotlin\nval answer = 42\n```\n\nNext paragraph",
)
assertEquals(2, blocks.size)
assertEquals(
StreamingMarkdownBlock.Code(
language = "kotlin",
code = "val answer = 42",
),
blocks[0],
)
assertEquals(StreamingMarkdownBlock.Text("Next paragraph"), blocks[1])
}
@Test
fun longerFence_isNotClosedByShorterFenceInsideCode() {
assertEquals(
listOf(
StreamingMarkdownBlock.Code(
language = "markdown",
code = "```\ninside\n```",
),
),
parseStreamingMarkdownBlocks(
"````markdown\n```\ninside\n```\n````",
),
)
}
@Test
fun table_staysRawUntilTheFinalCommonMarkParse() {
val content = "| Name | Value |\n| --- | --- |\n| Alpha | 1 |\n| Beta |"
assertEquals(
listOf(StreamingMarkdownBlock.Text(content)),
parseStreamingMarkdownBlocks(content),
)
}
@Test
fun incompleteTableDelimiter_doesNotPrematurelyPromoteTheTable() {
val content = "| Name | Value |\n| --- | --"
assertEquals(
listOf(StreamingMarkdownBlock.Text(content)),
parseStreamingMarkdownBlocks(content),
)
}
@Test
fun escapedHeaderPipe_doesNotMakeTheTableLookSettled() {
val content = "| Name \\| alias | Value |\n| --- | --- |\n| Alpha | 1 |\n| Beta |"
assertEquals(
listOf(StreamingMarkdownBlock.Text(content)),
parseStreamingMarkdownBlocks(content),
)
}
@Test
fun listAndContinuation_stayTogetherUntilFinal() {
val content = "- first paragraph\n\n continuation\n- second"
assertEquals(
listOf(StreamingMarkdownBlock.Text(content)),
parseStreamingMarkdownBlocks(content),
)
}
@Test
fun lazyBlockQuoteContinuation_isNeverSplitIntoASettledPrefix() {
val content = "> quoted line\n\nlazy continuation"
assertEquals(
listOf(StreamingMarkdownBlock.Text(content)),
parseStreamingMarkdownBlocks(content),
)
}
@Test
fun crlfInput_isNormalizedWithoutLeakingCarriageReturns() {
val blocks = parseStreamingMarkdownBlocks("First\r\n\r\nSecond")
assertEquals(
listOf(
StreamingMarkdownBlock.Markdown("First"),
StreamingMarkdownBlock.Text("Second"),
),
blocks,
)
assertTrue(blocks.none { it.toString().contains('\r') })
}
}
@@ -0,0 +1,149 @@
package com.hermesandroid.relay.ui.screens
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.data.ToolCall
import org.junit.Assert.assertEquals
import org.junit.Test
class ChatUnreadStateTest {
@Test
fun appendedMessagesAreCountedFromTheLastReadSnapshot() {
val readMessages = listOf(message("user", "Hello"))
val currentMessages = readMessages + listOf(
message("assistant", "Hi there", MessageRole.ASSISTANT),
message("system", "Connection restored", MessageRole.SYSTEM),
)
assertEquals(
2,
countUnreadMessages(
current = currentMessages.toUnreadSnapshot(),
lastRead = readMessages.toUnreadSnapshot(),
),
)
}
@Test
fun streamingGrowthCountsTheBubbleOnce() {
val readMessages = listOf(message("assistant", "Partial", MessageRole.ASSISTANT))
val firstUpdate = listOf(message("assistant", "Partial answer", MessageRole.ASSISTANT))
val secondUpdate = listOf(message("assistant", "Partial answer completed", MessageRole.ASSISTANT))
val lastRead = readMessages.toUnreadSnapshot()
assertEquals(1, countUnreadMessages(firstUpdate.toUnreadSnapshot(), lastRead))
assertEquals(1, countUnreadMessages(secondUpdate.toUnreadSnapshot(), lastRead))
}
@Test
fun visibleToolProgressCountsAsUnreadContent() {
val pending = message("assistant", "Working", MessageRole.ASSISTANT).copy(
toolCalls = listOf(
ToolCall(
id = "tool-1",
name = "search",
args = null,
result = null,
success = null,
),
),
)
val completed = pending.copy(
toolCalls = pending.toolCalls.map {
it.copy(result = "done", success = true, isComplete = true)
},
)
assertEquals(
1,
countUnreadMessages(
listOf(completed).toUnreadSnapshot(),
listOf(pending).toUnreadSnapshot(),
),
)
}
@Test
fun streamingFlagOnlyChangeDoesNotCreateUnreadContent() {
val streaming = message("assistant", "Complete text", MessageRole.ASSISTANT)
.copy(isStreaming = true)
val settled = streaming.copy(isStreaming = false)
assertEquals(
0,
countUnreadMessages(
listOf(settled).toUnreadSnapshot(),
listOf(streaming).toUnreadSnapshot(),
),
)
}
@Test
fun aTranscriptThatShrinksDoesNotCreateUnreadContent() {
val lastRead = listOf(
message("user", "Question"),
message("assistant", "Answer", MessageRole.ASSISTANT),
)
val current = listOf(message("user", "Question"))
assertEquals(
0,
countUnreadMessages(current.toUnreadSnapshot(), lastRead.toUnreadSnapshot()),
)
}
@Test
fun demoModeTakesPriorityOverLiveVoiceReadiness() {
assertEquals(
ChatVoiceAction.ShowDemoNotice,
resolveChatVoiceAction(isDemoMode = true, voiceReady = true),
)
assertEquals(
ChatVoiceAction.ShowDemoNotice,
resolveChatVoiceAction(isDemoMode = true, voiceReady = false),
)
}
@Test
fun demoDispatchNeverInvokesTheLiveVoiceCallback() {
var demoNotices = 0
var voiceStarts = 0
var setupNotices = 0
dispatchChatVoiceAction(
isDemoMode = true,
voiceReady = true,
onDemoNotice = { demoNotices += 1 },
onStartVoice = { voiceStarts += 1 },
onSetupNotice = { setupNotices += 1 },
)
assertEquals(1, demoNotices)
assertEquals(0, voiceStarts)
assertEquals(0, setupNotices)
}
@Test
fun liveChatUsesTheExistingVoiceReadinessGate() {
assertEquals(
ChatVoiceAction.StartVoice,
resolveChatVoiceAction(isDemoMode = false, voiceReady = true),
)
assertEquals(
ChatVoiceAction.ShowSetupNotice,
resolveChatVoiceAction(isDemoMode = false, voiceReady = false),
)
}
private fun message(
id: String,
content: String,
role: MessageRole = MessageRole.USER,
) = ChatMessage(
id = id,
role = role,
content = content,
timestamp = 0L,
)
}
@@ -0,0 +1,683 @@
package com.hermesandroid.relay.viewmodel
import android.os.Handler
import android.os.Looper
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.ChatTurnAskCheckpoint
import com.hermesandroid.relay.data.ChatTurnAssistantCheckpoint
import com.hermesandroid.relay.data.ChatTurnCheckpoint
import com.hermesandroid.relay.data.ChatTurnCheckpointStore
import com.hermesandroid.relay.data.ChatTurnToolCheckpoint
import com.hermesandroid.relay.data.ChatTurnUserCheckpoint
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.network.upstream.ChatHandler
import com.hermesandroid.relay.network.upstream.DashboardApiClient
import com.hermesandroid.relay.network.upstream.GatewayChatClient
import com.hermesandroid.relay.network.upstream.GatewayClientHarness
import com.hermesandroid.relay.network.upstream.GatewayConnectionState
import com.hermesandroid.relay.network.upstream.HermesApiClient
import com.hermesandroid.relay.network.upstream.models.MessageItem
import kotlinx.coroutines.CompletableDeferred
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.cancel
import kotlinx.coroutines.runBlocking
import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import okhttp3.OkHttpClient
import okhttp3.WebSocket
import okhttp3.mockwebserver.Dispatcher
import okhttp3.mockwebserver.MockResponse
import okhttp3.mockwebserver.MockWebServer
import okhttp3.mockwebserver.RecordedRequest
import okhttp3.mockwebserver.SocketPolicy
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.robolectric.RobolectricTestRunner
import org.robolectric.Shadows.shadowOf
import org.robolectric.annotation.Config
import java.util.concurrent.TimeUnit
import java.util.concurrent.atomic.AtomicInteger
@RunWith(RobolectricTestRunner::class)
@Config(sdk = [34])
class ChatViewModelGatewayInboundTurnTest {
private class MemoryCheckpointStore(
var checkpoint: ChatTurnCheckpoint? = null,
) : ChatTurnCheckpointStore {
override suspend fun read(): ChatTurnCheckpoint? = checkpoint
override suspend fun write(checkpoint: ChatTurnCheckpoint) {
this.checkpoint = checkpoint
}
override suspend fun clear() {
checkpoint = null
}
}
private lateinit var gatewayHarness: GatewayClientHarness
private lateinit var apiServer: MockWebServer
private lateinit var gatewayScope: CoroutineScope
private lateinit var gatewayClient: GatewayChatClient
private lateinit var serverWs: WebSocket
private lateinit var handler: ChatHandler
private lateinit var viewModel: ChatViewModel
@Volatile
private var persistedHistory: List<MessageItem> = emptyList()
@Volatile
private var holdCompletionsStream = false
@Before
fun setUp() {
gatewayHarness = GatewayClientHarness()
apiServer = MockWebServer().apply {
dispatcher = object : Dispatcher() {
override fun dispatch(request: RecordedRequest): MockResponse =
if (holdCompletionsStream && request.path == "/v1/chat/completions") {
MockResponse().setSocketPolicy(SocketPolicy.NO_RESPONSE)
} else {
MockResponse().setResponseCode(404)
}
}
start()
}
gatewayScope = CoroutineScope(SupervisorJob() + Dispatchers.IO)
gatewayClient = GatewayChatClient(
initialDashboardClient = DashboardApiClient(
baseUrl = gatewayHarness.server.url("/").toString().trimEnd('/'),
okHttpClient = OkHttpClient(),
),
okHttpClient = OkHttpClient(),
// Match production ordering: Gateway callbacks are posted from the
// OkHttp WebSocket thread onto Android's main looper.
callbackDispatcher = { block ->
Handler(Looper.getMainLooper()).post(block)
},
scope = gatewayScope,
)
handler = ChatHandler().also { it.setSessionId(STORED_SESSION_ID) }
persistedHistory = emptyList()
holdCompletionsStream = false
viewModel = ChatViewModel().also {
it.initialize(
HermesApiClient(apiServer.url("/").toString(), "test-key"),
handler,
)
it.streamingEndpoint = "gateway"
it.setProfileMessageLoader {
Result.success(persistedHistory)
}
it.updateGatewayClient(gatewayClient)
}
assertTrue(runBlocking { gatewayClient.prewarmAwait(STORED_SESSION_ID) })
serverWs = gatewayHarness.awaitServerSocket()
shadowOf(Looper.getMainLooper()).idle()
}
@After
fun tearDown() {
viewModel.updateGatewayClient(null)
gatewayClient.shutdown()
gatewayScope.cancel()
gatewayHarness.shutdown()
apiServer.shutdown()
}
@Test
fun unsolicitedGatewayCompletionAppearsAsOneAssistantTurnAndSettles() {
// Upstream's process-completion poller currently emits this adjacent
// duplicate pair; it must still create exactly one placeholder.
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
awaitCondition {
handler.messages.value.singleOrNull()?.content == BACKGROUND_ANSWER
}
assertTrue(handler.isStreaming.value)
assertEquals(1, handler.messages.value.size)
assertEquals(MessageRole.ASSISTANT, handler.messages.value.single().role)
persistedHistory = persistedAnswerHistory()
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
awaitCondition { !handler.isStreaming.value }
shadowOf(Looper.getMainLooper()).idle()
awaitCondition {
handler.messages.value.singleOrNull()?.content == BACKGROUND_ANSWER
}
assertFalse(handler.messages.value.single().isStreaming)
assertFalse(gatewayHarness.rpcLog.any { it.first == "prompt.submit" })
}
@Test
fun queuedMainDispatchAdmitsBackgroundStartAfterLocalCompletion() {
viewModel.sendMessage("Local gateway turn")
gatewayHarness.awaitRpc("prompt.submit")
awaitCondition { gatewayClient.hasActiveTurn() }
// Queue the local completion and the server-initiated start back to
// back. The inbound admission must run behind the local completion on
// main, rather than reading stale activeStream state on the socket.
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Local answer") },
"live-resumed",
),
)
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
awaitCondition {
handler.messages.value.any {
it.id.startsWith("gateway-inbound-") && it.content == BACKGROUND_ANSWER
}
}
awaitCondition { !handler.isStreaming.value }
}
@Test
fun stopOnUnsolicitedTurnInterruptsTheGatewaySession() {
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", "Still composing") },
"live-resumed",
),
)
awaitCondition { handler.isStreaming.value }
viewModel.cancelStream()
gatewayHarness.awaitRpc("session.interrupt")
assertFalse(handler.isStreaming.value)
// Upstream can emit the interrupted turn's terminal event after the
// interrupt RPC. It is a drain marker, not a new background answer.
persistedHistory = persistedAnswerHistory("Canceled answer", "canceled-server-answer")
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Canceled answer") },
"live-resumed",
),
)
Thread.sleep(150)
shadowOf(Looper.getMainLooper()).idleFor(250, TimeUnit.MILLISECONDS)
assertFalse(handler.messages.value.any { it.content == "Canceled answer" })
assertTrue(handler.messages.value.any { "Stopped" in it.badges })
}
@Test
fun reopenedChatRestoresRichStateAndReattachesLiveGatewayTurn() {
val now = System.currentTimeMillis()
val checkpointStore = MemoryCheckpointStore(
ChatTurnCheckpoint(
contextKey = PROFILE_CONTEXT,
sessionId = STORED_SESSION_ID,
liveSessionId = "live-resumed",
transport = "gateway",
user = ChatTurnUserCheckpoint("pending-user", "Research this", now - 2_000L),
assistant = ChatTurnAssistantCheckpoint(
id = "pending-assistant",
content = "Partial",
timestamp = now - 1_900L,
thinkingContent = "Inspecting sources",
isThinkingStreaming = true,
toolCalls = listOf(
ChatTurnToolCheckpoint(
id = "tool-1",
name = "terminal",
isComplete = false,
startedAt = now - 1_500L,
),
),
),
turnStatus = "Running terminal",
priorUserMessageCount = 0,
baselineAssistantCount = 0,
pendingAsk = ChatTurnAskCheckpoint(
kind = "APPROVAL",
text = "Allow the command?",
timeoutSeconds = 0,
messageId = "ask-approval-1",
cardKey = "approval-1",
receivedAt = now - 1_000L,
),
startedAt = now - 1_900L,
updatedAt = now,
),
)
gatewayHarness.recoveryRunning = true
gatewayHarness.recoveryAssistant = "Partial answer from upstream"
viewModel.setChatTurnCheckpointStore(checkpointStore)
// Cold process start: ConnectionViewModel's fresh handler has not yet
// adopted the persisted lastSessionId when the profile context binds.
handler.setSessionId(null)
viewModel.switchProfileContext(PROFILE_CONTEXT, STORED_SESSION_ID)
viewModel.prewarmGateway()
gatewayHarness.awaitRpc("session.activate")
awaitCondition {
handler.messages.value.any { it.id == "pending-assistant" } &&
handler.isStreaming.value
}
val restored = handler.messages.value.single { it.id == "pending-assistant" }
assertEquals("Partial answer from upstream", restored.content)
assertEquals("Inspecting sources", restored.thinkingContent)
assertEquals("terminal", restored.toolCalls.single().name)
assertFalse(restored.toolCalls.single().isComplete)
assertEquals("Running terminal", handler.turnStatus.value)
assertEquals("approval-1", viewModel.pendingAsk.value?.cardKey)
assertTrue(handler.messages.value.any { it.id == "ask-approval-1" && it.cards.isNotEmpty() })
assertTrue(
handler.messages.value.indexOfFirst { it.id == "pending-assistant" } <
handler.messages.value.indexOfFirst { it.id == "ask-approval-1" },
)
serverWs.send(
gatewayHarness.eventFrame(
"tool.complete",
buildJsonObject {
put("tool_id", "tool-1")
put("name", "terminal")
put("summary", "done")
},
"live-resumed",
),
)
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", " and finished") },
"live-resumed",
),
)
persistedHistory = listOf(
MessageItem(id = "server-user", role = "user", content = JsonPrimitive("Research this")),
MessageItem(
id = "server-assistant",
role = "assistant",
content = JsonPrimitive("Partial answer from upstream and finished"),
),
)
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Partial answer from upstream and finished") },
"live-resumed",
),
)
awaitCondition { !handler.isStreaming.value }
shadowOf(Looper.getMainLooper()).idle()
awaitCondition { checkpointStore.checkpoint == null }
assertTrue(handler.messages.value.any {
it.role == MessageRole.ASSISTANT &&
it.content == "Partial answer from upstream and finished"
})
}
@Test
fun lateCanceledCompletionDrainsBeforeImmediateNextTurn() {
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", "Old partial") },
"live-resumed",
),
)
awaitCondition { handler.isStreaming.value }
viewModel.cancelStream()
gatewayHarness.awaitRpc("session.interrupt")
viewModel.sendMessage("Start the next turn")
// Let the bounded next-submit wait elapse first. The started turn's
// tombstone must still drain its eventual terminal event rather than
// routing it into the new active mapper.
Thread.sleep(2_100)
awaitCondition {
gatewayHarness.rpcLog.any { (method, params) ->
method == "prompt.submit" && params["text"] == JsonPrimitive("Start the next turn")
}
}
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Canceled answer") },
"live-resumed",
),
)
persistedHistory = persistedAnswerHistory("New answer", "new-server-answer")
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", "New answer") },
"live-resumed",
),
)
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", "New answer") },
"live-resumed",
),
)
awaitCondition { handler.messages.value.any { it.content == "New answer" } }
awaitCondition { !handler.isStreaming.value }
assertFalse(handler.messages.value.any { it.content == "Canceled answer" })
}
@Test
fun queuedMessageDrainsAfterUnsolicitedTurnCompletes() {
gatewayHarness.steerStatus = "rejected"
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", "Finishing background work") },
"live-resumed",
),
)
awaitCondition { handler.isStreaming.value }
viewModel.sendMessage("Run this next")
gatewayHarness.awaitRpc("session.steer")
awaitCondition { viewModel.queuedMessages.value == listOf("Run this next") }
persistedHistory = persistedAnswerHistory()
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Finishing background work") },
"live-resumed",
),
)
awaitCondition {
gatewayHarness.rpcLog.any { (method, params) ->
method == "prompt.submit" && params["text"] == JsonPrimitive("Run this next")
}
}
assertTrue(viewModel.queuedMessages.value.isEmpty())
}
@Test
fun coldForegroundPrewarmReloadsACompletionMissedWhileDisconnected() {
persistedHistory = persistedAnswerHistory()
serverWs.close(1012, "test disconnect")
awaitCondition { gatewayClient.connectionState.value == GatewayConnectionState.Idle }
viewModel.prewarmGateway()
gatewayHarness.awaitServerSocket()
gatewayHarness.awaitRpcCount("session.resume", 2)
awaitCondition {
handler.messages.value.singleOrNull()?.content == BACKGROUND_ANSWER
}
assertFalse(handler.isStreaming.value)
}
@Test
fun unsolicitedTurnDoesNotReplaceAnActiveForcedSseTurn() {
holdCompletionsStream = true
viewModel.sendVoiceMessage("local voice turn", "Respond for spoken playback")
awaitCondition { handler.isStreaming.value }
val localPlaceholderId = handler.messages.value.last().id
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
persistedHistory = persistedAnswerHistory()
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
Thread.sleep(150)
shadowOf(Looper.getMainLooper()).idle()
assertTrue("the forced SSE turn must still own streaming", handler.isStreaming.value)
assertTrue(handler.messages.value.any { it.id == localPlaceholderId && it.isStreaming })
assertFalse(handler.messages.value.any { it.content == BACKGROUND_ANSWER })
viewModel.cancelStream()
awaitCondition { !handler.isStreaming.value }
awaitCondition { handler.messages.value.any { it.content == BACKGROUND_ANSWER } }
}
@Test
fun acceptedInboundTurnSettlesAfterGatewayDowngrade() {
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
awaitCondition { handler.isStreaming.value }
viewModel.streamingEndpoint = "sessions"
viewModel.updateGatewayClient(null)
persistedHistory = persistedAnswerHistory()
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
awaitCondition { !handler.isStreaming.value }
awaitCondition {
handler.messages.value.singleOrNull()?.id == "persisted-background-answer"
}
assertEquals(BACKGROUND_ANSWER, handler.messages.value.single().content)
}
@Test
fun lateOldCompletionDoesNotClearNewGatewayTurnSteering() {
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", "Old inbound") },
"live-resumed",
),
)
awaitCondition { handler.isStreaming.value }
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", "Old inbound") },
"live-resumed",
),
)
val deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(5)
while (gatewayClient.hasActiveTurn() && System.nanoTime() < deadline) {
// Deliberately do not idle main: keep the old completion callback
// queued while the socket-side mapper reaches its terminal state.
Thread.sleep(20)
}
assertFalse("old gateway turn never ended", gatewayClient.hasActiveTurn())
viewModel.cancelStream()
viewModel.sendMessage("New gateway turn")
awaitCondition { gatewayClient.hasActiveTurn() }
// Pump the queued old callback. It no longer owns activeStream and must
// not clear the new turn's steering affordance.
shadowOf(Looper.getMainLooper()).idle()
assertTrue(viewModel.steerableTurn.value)
}
@Test
fun reconnectAfterMissedStartRecoversOnExactSessionCompletion() {
serverWs.close(1012, "missed start")
awaitCondition { gatewayClient.connectionState.value == GatewayConnectionState.Idle }
viewModel.prewarmGateway()
serverWs = gatewayHarness.awaitServerSocket()
gatewayHarness.awaitRpcCount("session.resume", 2)
// Reconnected midway through the synthetic turn: no message.start is
// replayed, so the delta is intentionally ignored and completion drives
// authoritative history recovery.
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
persistedHistory = persistedAnswerHistory()
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
awaitCondition { handler.messages.value.any { it.content == BACKGROUND_ANSWER } }
assertFalse(handler.isStreaming.value)
}
@Test
fun staleHistoryReadCannotEraseATurnCompletedDuringTheFetch() {
val loadCount = AtomicInteger(0)
val firstLoadStarted = CompletableDeferred<Unit>()
val releaseFirstLoad = CompletableDeferred<Unit>()
viewModel.setProfileMessageLoader {
when (loadCount.incrementAndGet()) {
1 -> {
firstLoadStarted.complete(Unit)
releaseFirstLoad.await()
Result.success(persistedAnswerHistory())
}
else -> Result.success(
persistedAnswerHistory() +
MessageItem(
id = "newer-local-answer",
sessionId = STORED_SESSION_ID,
role = "assistant",
content = JsonPrimitive("Newer answer"),
),
)
}
}
serverWs.send(gatewayHarness.eventFrame("message.start", null, "live-resumed"))
serverWs.send(
gatewayHarness.eventFrame(
"message.delta",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
serverWs.send(
gatewayHarness.eventFrame(
"message.complete",
buildJsonObject { put("text", BACKGROUND_ANSWER) },
"live-resumed",
),
)
awaitCondition { firstLoadStarted.isCompleted }
handler.addPlaceholderMessage(
ChatMessage(
id = "newer-local-answer",
role = MessageRole.ASSISTANT,
content = "Newer answer",
timestamp = System.currentTimeMillis(),
isStreaming = true,
),
)
handler.onStreamComplete("newer-local-answer")
releaseFirstLoad.complete(Unit)
awaitCondition {
handler.messages.value.any { it.content == "Newer answer" } && loadCount.get() >= 2
}
assertTrue(handler.messages.value.any { it.content == BACKGROUND_ANSWER })
}
private fun persistedAnswerHistory(
answer: String = BACKGROUND_ANSWER,
id: String = "persisted-background-answer",
): List<MessageItem> = listOf(
MessageItem(
id = id,
sessionId = STORED_SESSION_ID,
role = "assistant",
content = JsonPrimitive(answer),
),
)
private fun awaitCondition(condition: () -> Boolean) {
val deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(5)
while (System.nanoTime() < deadline) {
// Advance Robolectric's paused main clock so coroutine delay-based
// reconciliation retries can resume as they do on-device.
shadowOf(Looper.getMainLooper()).idleFor(20, TimeUnit.MILLISECONDS)
if (condition()) return
Thread.sleep(20)
}
assertTrue("condition not met; messages=${handler.messages.value}", condition())
}
companion object {
private const val STORED_SESSION_ID = "stored-session"
private const val PROFILE_CONTEXT = "connection-a/profile-default"
private const val BACKGROUND_ANSWER = "Background task finished."
}
}
@@ -1,5 +1,6 @@
package com.hermesandroid.relay.viewmodel
import com.hermesandroid.relay.data.BackgroundTaskPhase
import com.hermesandroid.relay.network.upstream.ChatHandler
import com.hermesandroid.relay.network.upstream.HermesApiClient
import com.hermesandroid.relay.network.relay.RealtimeVoiceEvent
@@ -10,6 +11,7 @@ import okhttp3.mockwebserver.RecordedRequest
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
@@ -79,6 +81,34 @@ class ChatViewModelRealtimeTurnTest {
assertFalse(handler.messages.value.any { it.content == "Listening..." })
}
@Test
fun localVoiceCommandRemovesItsSyntheticChatTurn() {
val assistantId = viewModel.startRealtimeAgentTurn(userText = "", chatSessionId = "session-1")
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "voice.input_transcript.final",
text = "pause",
raw = "{}",
),
)
viewModel.discardRealtimeAgentLocalCommandTurn(assistantId)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "voice.response.delta",
delta = "late provider response",
raw = "{}",
),
)
assertTrue(handler.messages.value.none { it.id == assistantId })
assertTrue(handler.messages.value.none { it.content == "pause" })
assertTrue(handler.messages.value.none { it.content == "Cancelled." })
assertNull(handler.lastSentMessage.value)
}
@Test
fun normalResponseCompletionAllowsLaterBackgroundSummaryOnSameTurn() {
val assistantId = viewModel.startRealtimeAgentTurn(userText = "Check Hermes", chatSessionId = "session-1")
@@ -108,4 +138,198 @@ class ChatViewModelRealtimeTurnTest {
assertTrue(handler.isStreaming.value)
assertEquals("I'll check. Final answer.", assistant.content)
}
@Test
fun backgroundRunKeepsOneChatIdentityFromPromotionThroughDelivery() {
val assistantId = viewModel.startRealtimeAgentTurn(
userText = "Check release readiness across all targets",
chatSessionId = "session-1",
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "hermes.run.promoted",
runId = "run-42",
tier = "durable",
queuedCount = 1,
raw = "{}",
),
)
val promoted = handler.messages.value.single { it.id == assistantId }.backgroundTask
assertEquals("run-42", promoted?.id)
assertEquals("Check release readiness across all targets", promoted?.title)
assertEquals(BackgroundTaskPhase.RUNNING, promoted?.phase)
assertEquals(1, promoted?.queuedCount)
assertEquals(2, handler.messages.value.size)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "hermes.run.progress",
runId = "run-42",
activeToolName = "shell_command",
completedToolCount = 2,
message = "Checking Android targets",
raw = "{}",
),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "hermes.run.queued",
runId = "run-42",
queuedCount = 3,
raw = "{}",
),
)
val running = handler.messages.value.single { it.id == assistantId }.backgroundTask
assertEquals(BackgroundTaskPhase.RUNNING, running?.phase)
assertEquals("Checking Android targets", running?.statusLine)
assertEquals(2, running?.completedToolCount)
assertEquals(3, running?.queuedCount)
// The provider's initial spoken handoff completes before Hermes does;
// that must not settle the still-running task card.
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(type = "voice.response.done", raw = "{}"),
)
assertEquals(
BackgroundTaskPhase.RUNNING,
handler.messages.value.single { it.id == assistantId }.backgroundTask?.phase,
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "hermes.run.background_completed",
runId = "run-42",
success = true,
queuedCount = 0,
raw = "{}",
),
)
assertEquals(
BackgroundTaskPhase.DELIVERING,
handler.messages.value.single { it.id == assistantId }.backgroundTask?.phase,
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(type = "voice.response.started", raw = "{}"),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(
type = "voice.response.delta",
delta = "All targets are ready.",
raw = "{}",
),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = assistantId,
event = RealtimeVoiceEvent(type = "voice.response.done", raw = "{}"),
)
val settled = handler.messages.value.single { it.id == assistantId }
assertEquals(BackgroundTaskPhase.COMPLETE, settled.backgroundTask?.phase)
assertEquals("All targets are ready.", settled.content)
assertEquals(2, handler.messages.value.size)
}
@Test
fun backgroundRunKeepsItsOwnerAfterANewerLocalVoiceCommand() {
val backgroundAssistantId = viewModel.startRealtimeAgentTurn(
userText = "Check every release target",
chatSessionId = "session-1",
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = backgroundAssistantId,
event = RealtimeVoiceEvent(
type = "hermes.run.promoted",
runId = "run-42",
tier = "durable",
raw = "{}",
),
)
// The provider handoff ends, but run ownership must outlive the turn's
// normal tracking maps while Hermes continues in the background.
viewModel.applyRealtimeAgentEvent(
assistantMessageId = backgroundAssistantId,
event = RealtimeVoiceEvent(type = "voice.response.done", raw = "{}"),
)
val commandAssistantId = viewModel.startRealtimeAgentTurn(
userText = "",
chatSessionId = "session-1",
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = commandAssistantId,
event = RealtimeVoiceEvent(
type = "voice.input_transcript.final",
text = "pause",
raw = "{}",
),
)
viewModel.discardRealtimeAgentLocalCommandTurn(commandAssistantId)
// The persistent Voice callback supplies its newest assistant id, but
// run-scoped and delivery events still belong to the initiating row.
viewModel.applyRealtimeAgentEvent(
assistantMessageId = commandAssistantId,
event = RealtimeVoiceEvent(
type = "hermes.run.progress",
runId = "run-42",
message = "Checking the final target",
completedToolCount = 3,
raw = "{}",
),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = commandAssistantId,
event = RealtimeVoiceEvent(
type = "hermes.run.background_completed",
runId = "run-42",
success = true,
raw = "{}",
),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = commandAssistantId,
event = RealtimeVoiceEvent(
type = "voice.response.started",
delivery = "forced_summary",
raw = "{}",
),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = commandAssistantId,
event = RealtimeVoiceEvent(
type = "voice.response.delta",
source = "provider",
delta = "Every release target is ready.",
delivery = "forced_summary",
raw = "{}",
),
)
viewModel.applyRealtimeAgentEvent(
assistantMessageId = commandAssistantId,
event = RealtimeVoiceEvent(
type = "voice.response.done",
delivery = "forced_summary",
raw = "{}",
),
)
val background = handler.messages.value.single { it.id == backgroundAssistantId }
assertEquals(BackgroundTaskPhase.COMPLETE, background.backgroundTask?.phase)
assertEquals("Every release target is ready.", background.content)
assertEquals(3, background.backgroundTask?.completedToolCount)
assertTrue(handler.messages.value.none { it.id == commandAssistantId })
assertTrue(handler.messages.value.none { it.content == "pause" })
assertEquals("Check every release target", handler.lastSentMessage.value)
}
}
@@ -0,0 +1,355 @@
package com.hermesandroid.relay.viewmodel
import com.hermesandroid.relay.network.upstream.GatewayProcess
import com.hermesandroid.relay.network.upstream.GatewayProcessCapability
import com.hermesandroid.relay.network.upstream.GatewayProcessEvent
import kotlinx.coroutines.CompletableDeferred
import kotlinx.coroutines.ExperimentalCoroutinesApi
import kotlinx.coroutines.NonCancellable
import kotlinx.coroutines.flow.MutableStateFlow
import kotlinx.coroutines.test.advanceTimeBy
import kotlinx.coroutines.test.runCurrent
import kotlinx.coroutines.test.runTest
import kotlinx.coroutines.withContext
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertTrue
import org.junit.Test
@OptIn(ExperimentalCoroutinesApi::class)
class GatewayProcessControllerTest {
@Test
fun pollsEveryFiveSecondsOnlyWhileAProcessIsRunning() = runTest {
val source = FakeProcessSource()
source.snapshot = listOf(process(id = "p1", status = "running"))
val controller = GatewayProcessController(this)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
assertEquals(1, source.listCalls)
assertEquals(listOf("p1"), controller.processes.value.map(GatewayProcess::id))
advanceTimeBy(4_999L)
runCurrent()
assertEquals(1, source.listCalls)
source.snapshot = listOf(process(id = "p1", status = "exited", exitCode = 0))
advanceTimeBy(1L)
runCurrent()
assertEquals(2, source.listCalls)
assertFalse(controller.processes.value.single().isRunning)
advanceTimeBy(20_000L)
runCurrent()
assertEquals(2, source.listCalls)
controller.close()
}
@Test
fun staleListResultCannotOverwriteNewSession() = runTest {
val oldResult = CompletableDeferred<Result<List<GatewayProcess>>>()
val source = FakeProcessSource().apply {
listHandler = { call ->
if (call == 1) {
// Model a transport that finishes an RPC even after the
// controller cancels the old refresh job.
withContext(NonCancellable) { oldResult.await() }
} else {
Result.success(listOf(process(id = "new", command = "new session")))
}
}
}
val controller = GatewayProcessController(this)
controller.bind(source, "chat-old")
controller.sessionReady("chat-old")
runCurrent()
assertEquals(1, source.listCalls)
controller.selectSession("chat-new")
controller.sessionReady("chat-new")
runCurrent()
assertEquals(listOf("new"), controller.processes.value.map(GatewayProcess::id))
oldResult.complete(Result.success(listOf(process(id = "old", command = "old session"))))
runCurrent()
assertEquals(listOf("new"), controller.processes.value.map(GatewayProcess::id))
controller.close()
}
@Test
fun messageCompleteFallbackRefreshesAndLiveOutputExtendsTheRecoverableTail() = runTest {
val source = FakeProcessSource().apply {
snapshot = listOf(
process(id = "p1", status = "running", outputTail = "before\n"),
)
}
val controller = GatewayProcessController(this, outputTailLimit = 12)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
source.emit(GatewayProcessEvent.Output("p1", "after-output"))
runCurrent()
assertEquals("after-output", controller.processes.value.single().outputTail)
source.snapshot = listOf(
process(id = "p1", status = "exited", exitCode = 7, outputTail = "final"),
)
source.emit(
GatewayProcessEvent.Invalidated(GatewayProcessEvent.Trigger.MESSAGE_COMPLETE),
)
runCurrent()
assertEquals(2, source.listCalls)
assertEquals(7, controller.processes.value.single().exitCode)
assertEquals("final", controller.processes.value.single().outputTail)
controller.close()
}
@Test
fun messageCompleteDiscoversRunningProcessAndStartsPolling() = runTest {
val source = FakeProcessSource()
val controller = GatewayProcessController(this)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
assertEquals(1, source.listCalls)
assertTrue(controller.processes.value.isEmpty())
source.snapshot = listOf(process(id = "p1", status = "running"))
source.emit(
GatewayProcessEvent.Invalidated(GatewayProcessEvent.Trigger.MESSAGE_COMPLETE),
)
runCurrent()
assertEquals(2, source.listCalls)
assertEquals(listOf("p1"), controller.processes.value.map(GatewayProcess::id))
source.snapshot = listOf(process(id = "p1", status = "exited", exitCode = 0))
advanceTimeBy(5_000L)
runCurrent()
assertEquals(3, source.listCalls)
assertFalse(controller.processes.value.single().isRunning)
advanceTimeBy(20_000L)
runCurrent()
assertEquals(3, source.listCalls)
controller.close()
}
@Test
fun finishedRowsStayUntilDismissedAndReappearForANewIdentity() = runTest {
val source = FakeProcessSource().apply {
snapshot = listOf(
process(
id = "p1",
status = "exited",
exitCode = 0,
startedAt = "first",
),
)
}
val controller = GatewayProcessController(this)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
advanceTimeBy(60_000L)
runCurrent()
assertEquals(1, controller.processes.value.size)
controller.dismiss("p1")
assertTrue(controller.processes.value.isEmpty())
source.snapshot = listOf(
process(
id = "p1",
command = "replacement",
status = "running",
startedAt = "second",
),
)
controller.refresh(showLoading = false)
runCurrent()
assertEquals("replacement", controller.processes.value.single().command)
controller.close()
}
@Test
fun stopExposesInFlightIdentityThenRefreshesAuthoritativeState() = runTest {
val killResult = CompletableDeferred<Result<Unit>>()
val source = FakeProcessSource().apply {
snapshot = listOf(process(id = "p1", status = "running"))
killHandler = { killResult.await() }
}
val controller = GatewayProcessController(this)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
controller.stop("p1")
runCurrent()
assertEquals(setOf("p1"), controller.stoppingProcessIds.value)
assertEquals(listOf("p1"), source.killCalls)
source.snapshot = listOf(process(id = "p1", status = "exited", exitCode = 143))
killResult.complete(Result.success(Unit))
runCurrent()
assertTrue(controller.stoppingProcessIds.value.isEmpty())
assertEquals(2, source.listCalls)
assertEquals(143, controller.processes.value.single().exitCode)
controller.close()
}
@Test
fun unsupportedCapabilityClearsSnapshotAndStopsPolling() = runTest {
val source = FakeProcessSource().apply {
snapshot = listOf(process(id = "p1", status = "running"))
}
val controller = GatewayProcessController(this)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
assertTrue(controller.processes.value.isNotEmpty())
source.capabilityState.value = GatewayProcessCapability.Unsupported
runCurrent()
assertTrue(controller.processes.value.isEmpty())
assertEquals(GatewayProcessCapability.Unsupported, controller.capability.value)
advanceTimeBy(10_000L)
runCurrent()
assertEquals(1, source.listCalls)
controller.close()
}
@Test
fun sameSessionReadyAfterReconnectRelistsEvenWithoutAnActivePoller() = runTest {
val source = FakeProcessSource()
val controller = GatewayProcessController(this)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
assertEquals(1, source.listCalls)
assertTrue(controller.processes.value.isEmpty())
source.snapshot = listOf(
process(id = "completed-offline", status = "exited", exitCode = 0),
)
controller.sessionReady("chat-a")
runCurrent()
assertEquals(2, source.listCalls)
assertEquals("completed-offline", controller.processes.value.single().id)
controller.close()
}
@Test
fun backgroundDisablesPollingUntilForegroundSessionReadyRefresh() = runTest {
val source = FakeProcessSource().apply {
snapshot = listOf(process(id = "p1", status = "running"))
}
val controller = GatewayProcessController(this)
controller.bind(source, "chat-a")
controller.sessionReady("chat-a")
runCurrent()
assertEquals(1, source.listCalls)
source.pollingAllowed = false
advanceTimeBy(5_000L)
runCurrent()
assertEquals(1, source.listCalls)
source.pollingAllowed = true
controller.sessionReady("chat-a")
runCurrent()
assertEquals(2, source.listCalls)
controller.close()
}
@Test
fun profileScopeChangeRejectsSameSessionIdStaleResult() = runTest {
val oldResult = CompletableDeferred<Result<List<GatewayProcess>>>()
val source = FakeProcessSource().apply {
listHandler = { call ->
if (call == 1) {
withContext(NonCancellable) { oldResult.await() }
} else {
Result.success(listOf(process(id = "new-profile")))
}
}
}
val controller = GatewayProcessController(this)
controller.bind(source, "same-id", scopeKey = "profile-a")
controller.sessionReady("same-id")
runCurrent()
controller.selectSession("same-id", scopeKey = "profile-b")
controller.sessionReady("same-id")
runCurrent()
oldResult.complete(Result.success(listOf(process(id = "old-profile"))))
runCurrent()
assertEquals(listOf("new-profile"), controller.processes.value.map(GatewayProcess::id))
controller.close()
}
private class FakeProcessSource : GatewayProcessSource {
val capabilityState = MutableStateFlow(GatewayProcessCapability.Unknown)
override val capability = capabilityState
var snapshot: List<GatewayProcess> = emptyList()
var listCalls = 0
var listHandler: suspend (Int) -> Result<List<GatewayProcess>> = { Result.success(snapshot) }
var killHandler: suspend (String) -> Result<Unit> = { Result.success(Unit) }
val killCalls = mutableListOf<String>()
var pollingAllowed = true
private var listener: ((GatewayProcessEvent) -> Unit)? = null
override suspend fun listProcesses(): Result<List<GatewayProcess>> {
listCalls += 1
return listHandler(listCalls)
}
override suspend fun killProcess(processId: String): Result<Unit> {
killCalls += processId
return killHandler(processId)
}
override fun setEventListener(listener: ((GatewayProcessEvent) -> Unit)?) {
this.listener = listener
}
override fun isPollingAllowed(): Boolean = pollingAllowed
fun emit(event: GatewayProcessEvent) {
listener?.invoke(event)
}
}
private fun process(
id: String,
command: String = "sleep 60",
status: String = "exited",
exitCode: Int? = null,
outputTail: String? = null,
startedAt: String? = "start",
) = GatewayProcess(
id = id,
command = command,
status = status,
exitCode = exitCode,
outputTail = outputTail,
startedAt = startedAt,
)
}
@@ -0,0 +1,169 @@
package com.hermesandroid.relay.voice
import org.junit.Assert.assertEquals
import org.junit.Assert.assertNull
import org.junit.Test
class VoiceCommandInterpreterTest {
@Test
fun `normalizes casing whitespace and terminal punctuation`() {
val action = VoiceCommandInterpreter.interpretFinalTranscript(
rawTranscript = " STOP TALKING!!! ",
context = VoiceCommandContext(responseActive = true),
)
assertEquals(VoiceCommandAction.StopResponse, action)
}
@Test
fun `normalizes hands-free punctuation without fuzzy matching`() {
val action = VoiceCommandInterpreter.interpretFinalTranscript(
rawTranscript = "Pause hands-free listening.",
context = VoiceCommandContext(
continuousModeSelected = true,
continuousListeningActive = true,
),
)
assertEquals(VoiceCommandAction.PauseContinuousListening, action)
}
@Test
fun `rejects partial transcripts and phrases embedded in ordinary prompts`() {
val active = VoiceCommandContext(responseActive = true)
assertNull(VoiceCommandInterpreter.interpretFinalTranscript("stop talk", active))
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"Can you stop talking about the old design?",
active,
),
)
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"Explain how to stop the response from timing out",
active,
),
)
}
@Test
fun `stop is response-only and never aliases background cancellation`() {
val bothActive = VoiceCommandContext(
responseActive = true,
backgroundTaskActive = true,
)
assertEquals(
VoiceCommandAction.StopResponse,
VoiceCommandInterpreter.interpretFinalTranscript("stop talking", bothActive),
)
assertEquals(
VoiceCommandAction.CancelBackgroundTask,
VoiceCommandInterpreter.interpretFinalTranscript(
"cancel the background task",
bothActive,
),
)
assertNull(VoiceCommandInterpreter.interpretFinalTranscript("cancel", bothActive))
assertNull(VoiceCommandInterpreter.interpretFinalTranscript("stop", bothActive))
}
@Test
fun `state gates stop and background cancellation`() {
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"stop talking",
VoiceCommandContext(responseActive = false),
),
)
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"cancel my background task",
VoiceCommandContext(backgroundTaskActive = false),
),
)
}
@Test
fun `pause and resume require continuous mode and the matching loop state`() {
assertEquals(
VoiceCommandAction.PauseContinuousListening,
VoiceCommandInterpreter.interpretFinalTranscript(
"pause",
VoiceCommandContext(
continuousModeSelected = true,
continuousListeningActive = true,
),
),
)
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"pause continuous listening",
VoiceCommandContext(
continuousModeSelected = true,
continuousListeningActive = false,
),
),
)
assertEquals(
VoiceCommandAction.ResumeContinuousListening,
VoiceCommandInterpreter.interpretFinalTranscript(
"resume",
VoiceCommandContext(
continuousModeSelected = true,
continuousListeningPaused = true,
),
),
)
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"resume continuous listening",
VoiceCommandContext(
continuousModeSelected = false,
continuousListeningPaused = true,
),
),
)
}
@Test
fun `repeat requires a delivered background answer`() {
assertEquals(
VoiceCommandAction.RepeatBackgroundAnswer,
VoiceCommandInterpreter.interpretFinalTranscript(
"repeat that",
VoiceCommandContext(backgroundAnswerAvailable = true),
),
)
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"repeat the last background answer",
VoiceCommandContext(backgroundAnswerAvailable = false),
),
)
}
@Test
fun `new chat is typed only at an allowed boundary`() {
assertEquals(
VoiceCommandAction.StartNewChat,
VoiceCommandInterpreter.interpretFinalTranscript(
"New chat!",
VoiceCommandContext(canStartNewChat = true),
),
)
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"start a new chat",
VoiceCommandContext(canStartNewChat = false),
),
)
assertNull(
VoiceCommandInterpreter.interpretFinalTranscript(
"Should I start a new chat for this?",
VoiceCommandContext(canStartNewChat = true),
),
)
}
}
+40
View File
@@ -1890,3 +1890,43 @@ convention and `CLAUDE.md` prose, never by structure or test:
- `.github/workflows/ci-android.yml` (boundary test in the explicit `--tests` list)
- `.github/workflows/ci-contract.yml` (vanilla-upstream route-contract job)
- `docs/plans/upstream-relay-isolation.md`
## ADR 35 — In-flight Chat recovery is session-centric and client-checkpointed
**Status:** Accepted (2026-07-10).
**Context.** A session-backed agent turn can outlive Android's Activity, process,
or WebSocket. Current upstream Hermes can reattach an exact live Gateway session
and reports whether it is still running plus its user/partial-assistant text, but
that server snapshot does not contain Android presentation state such as live
reasoning, tool/subagent card phases, lifecycle captions, pending ask cards, or a
client-owned background-task chip. Persisted history becomes authoritative only
after the turn settles and therefore cannot restore the in-between UI.
**Decision.** Treat Chat recovery as session-centric instead of request-centric:
1. Persist at most one active, session-backed turn in the shared Android
DataStore, scoped by connection/profile context and durable session id, with a
24-hour expiry. Store the visible assistant state and server-issued ask, but
never an entered password/secret or approval response.
2. On reopen, restore that UI immediately, then recover in this order: exact
`session.activate` using the saved live id; `session.resume` using the durable
id; bounded, positionally anchored history reconciliation.
3. Separate **detach** from **cancel**. Lifecycle teardown releases local callbacks
without `session.interrupt`; explicit Stop and session/profile/connection
switches retain their interrupt-and-clear behavior.
4. Apply the same history fallback to sessions-SSE transport drops and route
handoffs. A final persisted transcript replaces the checkpoint and clears it.
**Consequences.** Reopening Chat can continue the same assistant bubble with its
last-known reasoning and tool state instead of inventing a second prompt or empty
spinner. Tool state that changed while no client was attached remains explicitly
last-known until a new event or authoritative history reconcile arrives. Older
Hermes builds without live activation degrade to durable history recovery.
**Key files:**
- `app/src/main/kotlin/com/hermesandroid/relay/data/ChatTurnCheckpointStore.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/network/upstream/GatewayChatClient.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/network/upstream/ChatHandler.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/viewmodel/ChatViewModel.kt`
@@ -0,0 +1,197 @@
# Android 1.4.1 Chat + Voice enhancement batch
## Goal
Ship a cohesive Android 1.4.1 quality release that makes Chat easier to read and
use, keeps the authoritative background answer in Chat history, carries task-card
state through the live/in-process reconciliation path, and advances Voice from a
testbench-style overlay toward a dependable hands-free work surface. Cold-restart
reconstruction of client-only task metadata is a separate durability follow-up.
This plan starts from `dev` after the background-route-loss recovery fix. It does
not authorize a version bump, merge, push, tag, deployment, Play upload, or release.
## Local implementation status
- **Chat wave:** implemented — conservative streaming Markdown promotion, wide
tables, bounded multi-image galleries, motion/TalkBack handling, Demo mic gate,
unread tracking, and one first-class background-task assistant identity.
- **Voice wave:** implemented — exact/state-gated local commands, four safe mode
presets, relay event-schema aliases, and foreground exact-delivery validation,
generation-scoped confirmation, and one-fallback failure handling.
- **Deliberately not claimed:** shared Voice/Chat task-card rendering, cold-restart
trace/task metadata, provider-voiced DONE replay, spoken-truncation UI, OpenAI
rollover/OOB/async outputs, N-way runs, semantic VAD, and WebRTC.
## Product direction
Hermes Chat is a dense technical conversation surface, not a social-messaging
clone. Preserve the existing graphite/Material visual language and spend the
visual emphasis on structured content: stable Markdown, readable tables,
explicit task state, and compact tool timelines. Voice should reuse that same
task vocabulary so a background run has one identity whether it is heard,
glanced at in the overlay, or opened later in Chat.
### Chat design calibration
- **Subject / audience / job:** a mobile control room for technical Hermes users;
the Chat surface's job is to make long structured answers and active agent work
legible without turning the transcript into a dashboard.
- **Palette:** preserve the existing theme-reactive brand tokens, anchored in the
Hermes dark baseline: Background `#08090D`, Navy `#121426`, Navy2 `#191B31`,
Ink `#F7F6F0`, Relay `#AEBFFF`, with semantic Green `#58D36F`, Amber `#F2B14B`,
and Danger `#FF6B78`. Do not introduce feature-local colors.
- **Type:** use the selected app font for prose and headings; keep the current
chat-tuned Markdown scale (20/18/16sp headings, 14sp prose) and 13sp monospace
utility type. Stability and hierarchy matter more than making text larger.
- **Layout:** preserve the current message rhythm. Structured content may use the
full assistant-bubble measure, but new status/detail surfaces stay subordinate
to the answer.
- **Signature:** streaming structured content should feel physically stable —
headings, tables, code, and background-task state retain their identity from
first useful render through final settlement.
```text
assistant bubble
├─ settled Markdown blocks (final renderer)
├─ active tail (raw only while structurally incomplete)
├─ background task card (one identity, running → settled)
└─ expandable tool lane / gallery when present
```
The deliberate visual risk is the wider structured-content measure and table
overflow affordance; everything else remains quiet. This keeps the change
specific to technical Chat content instead of applying a generic messenger
restyle.
## Wave 1 — Chat — implemented locally
1. **Stable structured Markdown**
- Render blank-terminated, unambiguous top-level prose/headings with the real
Markdown renderer while the incomplete tail remains raw. Lists, quotes,
tables, HTML, and fences wait for the final parse to avoid CommonMark
re-parenting.
- Add horizontally scrollable wide tables with readable minimum columns and a
quiet overflow affordance.
- Keep code-block copy and horizontal scrolling unchanged.
2. **Attachment gallery**
- Render multiple images in one message as a compact grid.
- Open the selected image into a swipeable full-screen gallery while retaining
original-byte Share/Save behavior.
3. **Accessibility and small Chat affordances**
- Make the thinking indicator honor OS animation scale and TalkBack touch
exploration in addition to the app preference.
- Gate the Voice mic in Demo mode with a clear local explanation.
- Add an unread count to the jump-to-bottom affordance without disturbing
users who intentionally scrolled away.
4. **Background tasks become first-class Chat turns**
- Give each background run a stable short title and Chat-visible identity.
- Create one running entry at kickoff and settle the same entry with the final
result; do not create duplicate system/reply rows.
- Preserve the result through session history/reload where the current
upstream or relay surface can represent it honestly.
- Expand the entry into the existing `SubagentLane`/tool-timeline vocabulary.
5. **Ordinary Gateway background completion handoff**
- Accept upstream's unsolicited `message.start` → delta → completion lifecycle
for the exact active session even when Android did not call `sendTurn()`.
- Render the follow-up as a normal assistant response; Gateway events do not
carry the structured run metadata required by the Realtime Voice task card.
- Refresh persisted history after a cold foreground resume to recover a reply
that completed while the Gateway socket was closed.
6. **Ordinary Gateway background-process activity**
- Use upstream `process.list` as the authoritative, session-scoped snapshot and
`process.kill` for one exact process; never call the global `process.stop`.
- Refresh on terminal/process tool completion and process status events, merge
live `agent.terminal.output` chunks, and poll every five seconds only while a
listed process remains active and foreground polling is allowed.
- Show a compact status strip above the composer. Open a mobile bottom sheet for
running/recent rows, elapsed time, expandable output, Stop, and Dismiss.
- Preserve upstream's synthetic completion row as user-role history while
presenting it as a non-editable process notice. Feature-detect the RPCs and
hide the activity surface against older Hermes builds.
- Fence snapshots and actions by connection, profile context, stored session,
and exact live session. Redact raw output from logcat and strip terminal
control sequences in the plain-text mobile viewer.
## Wave 2 — Voice — implemented subset + explicit residuals
1. **Hands-free command layer and presets**
- Implemented exact final-transcript commands for stop, explicit background
cancel, repeat, safe pause/resume, and Standard new-chat. Realtime new-chat
remains gated on a clean session-rebind boundary.
- Implemented Hands-free, Low latency, Careful tools, and Quiet/visual-only
presets as compositions of existing settings, with manual settings still
available and experimental barge-in never silently enabled.
2. **Shared task status and timeline**
- Chat now owns a stable task identity and expandable tool detail. Voice keeps
its existing background chip for 1.4.1; shared card/timeline rendering and a
compact waiting-on-user/current-objective summary remain deferred.
3. **Exact-delivery hardening and polish**
- Implemented provider send/request failure fallback, direct exact text where
supported, structured instruction routing, answer-aware validation, and
generation-scoped delivery confirmation across foreground/background paths.
- Barge-in preemption-as-visible-text, provider-voiced DONE-chip replay, and a
"full answer in Chat" truncation affordance remain open.
4. **Durability and provider follow-ups that fit the batch**
- App-restart persistence for unsynced realtime trace/result state remains
open and still needs an explicit connection/profile/session key.
- The current five-voice xAI fallback, dynamic discovery, and expressive
speech-tag controls were already present; only live settings verification
remains.
- Keep OpenAI 60-minute rollover, out-of-band exact delivery, async function
output, semantic VAD, N-way concurrent runs, and WebRTC behind separate
empirical/design gates unless the implementation can be proven locally
without weakening ADR 29/32/33.
## Verification
### Local results
- Google Play and sideload focused suites passed for the ten new Chat/Voice test
classes. The final cross-turn task-ownership/command-cleanup class was then
forced through a clean `--rerun-tasks` execution.
- `:app:lintGooglePlayDebug` and `:app:lintSideloadDebug` passed. The first lint
attempt exposed an Android lint UAST crash for `internal` shared attachment
composables; ordinary top-level visibility removed that analyzer bug.
- The realtime routes, promotion, validation, xAI, and OpenAI provider slice is
94/94 green via `python -m unittest`.
- No remote-host deployment, ADB install, version bump, or release action ran.
### Owner on-device checklist
- Stream headings/prose and wide tables; confirm stable reflow, readable columns,
horizontal scrolling, and the overflow fade.
- Open 2–5 image messages; verify selected-page entry, swiping, edge pan/zoom,
sensitive reveal/action gating, and Share/Save against original bytes.
- Scroll away during a reply and verify unread counts; exercise TalkBack/reduced
motion and confirm Demo mic never attempts transcription.
- Run a background task, then a quick second Voice turn/local command; confirm the
initiating Chat card alone receives progress, tools, delivery, and final answer.
- In ordinary Gateway Chat, start a terminal background process and confirm the
process strip, sheet, elapsed state, live/output-tail expansion, exact Stop,
completion/Dismiss, synthetic process notice, and unsolicited agent follow-up.
Switch chats and reconnect once to verify strict session isolation and snapshot
recovery.
- Exercise Standard/Realtime stop-vs-explicit-cancel, pause/resume, repeat, Standard
new-chat Continuous rearm, and ordinary command-like prompts that must reach Hermes.
- Apply all four presets, force a manual divergence to Custom, and confirm route,
provider, model, voice, concurrency, credentials, and barge-in choices are preserved.
- Re-test foreground exact delivery and provider send/request failure: one fallback,
one completion boundary, then the terminal error. Recheck route-loss recovery,
waveform start, final-syllable tail, PCM tap/static, tool order, and screen wake lock.
## Release bookkeeping
- Update `[Unreleased]` CHANGELOG entries with public, concise 1.4.1 bullets.
- Record factual implementation and verification in DEVLOG.
- Remove completed items from TODO or mark only residual verification.
- Do not bump Android/plugin versions or create release artifacts until the
owner explicitly starts release preparation.
+11 -5
View File
@@ -85,12 +85,18 @@ This app is a community project and is not affiliated with or endorsed by NousRe
Paste into Play Console → **What's new** (≤500 characters):
```
v1.4.0 — Realtime voice that finishes the job.
v1.4.1 - Chat that keeps up
• Long voice tasks can queue, keep running while you ask quick follow-ups, and deliver answers in the selected realtime voice.
• Voice sessions recover more reliably after background or route changes and clear stale task states.
• Refresh model catalogs on demand; add opt-in notification rules and multi-device Bridge targeting.
• Safer startup, server-address handling, long chat turns, and credential media access.
Chat
* Follow background work from a live process strip; its result appears automatically in the same conversation.
* Reopen while an answer runs: partial text, thinking, tool progress, and approvals return.
Voice
* Speak commands to pause, resume, cancel, repeat a result, or start Standard voice chat.
* Pick Hands-free, Low latency, Careful tools, or Quiet presets.
Polish
* Multi-image galleries plus smoother streaming Markdown and long tables.
```
## Category
+2 -2
View File
@@ -1,6 +1,6 @@
[versions]
appVersionName = "1.4.0"
appVersionCode = "22"
appVersionName = "1.4.1"
appVersionCode = "23"
agp = "9.2.1"
kotlin = "2.4.0"
compose-bom = "2026.06.01"
+1 -1
View File
@@ -3,7 +3,7 @@
"label": "Relay",
"description": "Paired devices, bridge activity, media inspection, and remote access for hermes-relay",
"icon": "Activity",
"version": "1.4.0",
"version": "1.4.1",
"tab": {
"path": "/relay",
"position": "after:skills"
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "hermes-relay-dashboard",
"version": "1.4.0",
"version": "1.4.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "hermes-relay-dashboard",
"version": "1.4.0",
"version": "1.4.1",
"devDependencies": {
"esbuild": "^0.25.12",
"qrcode": "^1.5.4"
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "hermes-relay-dashboard",
"version": "1.4.0",
"version": "1.4.1",
"private": true,
"description": "Hermes-Relay dashboard plugin frontend (IIFE bundle). Loaded verbatim by the hermes-agent dashboard via the Plugin SDK global.",
"scripts": {
+1 -1
View File
@@ -1,6 +1,6 @@
name: hermes-relay
manifest_version: 1
version: 1.4.0
version: 1.4.1
description: "Hermes-Relay plugin for QR pairing, relay sessions, dashboard management, remote desktop/phone tooling, and optional legacy compatibility diagnostics. Standard chat, Manage, and dashboard voice remain vanilla upstream Hermes surfaces."
author: Axiom Labs
# All three are OPTIONAL — only needed if you use the relay's extra /
+1 -1
View File
@@ -19,7 +19,7 @@ See ``plugin/relay/server.py`` for the aiohttp server,
# CLI releases use desktop/package.json and cli-v* tags. The /health endpoint
# reports this plugin version, and stale values make live diagnosis harder than
# it should be.
__version__ = "1.4.0"
__version__ = "1.4.1"
from .server import create_app, main # noqa: E402 — must come after __version__
+207 -45
View File
@@ -213,6 +213,10 @@ class RealtimeAgentSession:
native_forced_hermes_turn_active: bool = False
native_forced_summary_active: bool = False
native_forced_summary_done: bool = False
# Monotonic ownership token for delivery-confirm alarms. A later result or
# preemption invalidates an older alarm even if the session-global summary
# flags have since been reset for another turn.
native_forced_summary_generation: int = 0
native_forced_summary_response_id: str | None = None
native_forced_summary_result: dict[str, Any] | None = None
native_forced_summary_buffer: list[dict[str, Any]] = field(default_factory=list)
@@ -1672,6 +1676,7 @@ class RealtimeAgentHandler:
"session_id": session.session_id,
"chat_session_id": session.chat_session_id,
"response_id": event.response_id,
"delivery": "forced_summary",
}
)
continue
@@ -1717,6 +1722,7 @@ class RealtimeAgentHandler:
"source": "provider",
"delta": delta,
"response_id": event.response_id,
"delivery": "forced_summary",
}
if session.native_forced_summary_committed:
# Early-committed: stream live.
@@ -1754,10 +1760,14 @@ class RealtimeAgentHandler:
if self._should_forward_provider_response_event(session, event):
if session.native_forced_summary_committed:
# Early-committed: stream audio live.
await self._send_provider_audio_delta(ws, session, event)
payload = await self._provider_audio_delta_event(ws, session, event)
if payload is not None:
payload["delivery"] = "forced_summary"
await self._send(ws, session, payload)
else:
payload = await self._provider_audio_delta_event(ws, session, event)
if payload is not None:
payload["delivery"] = "forced_summary"
session.native_forced_summary_buffer.append(payload)
continue
if not self._should_forward_provider_response_event(session, event):
@@ -1785,6 +1795,7 @@ class RealtimeAgentHandler:
audio_done_event = {
"type": SERVER_EVT_OUTPUT_AUDIO_DONE,
"response_id": event.response_id,
"delivery": "forced_summary",
}
if session.native_forced_summary_committed:
await self._send(ws, session, audio_done_event)
@@ -2087,6 +2098,7 @@ class RealtimeAgentHandler:
"event_log_path": str(session.event_log_path),
"chat_session_id": session.chat_session_id,
"response_id": None if response_id == "__blank_response_id__" else response_id,
"delivery": "forced_summary",
},
)
self._mark_provider_response_done(session, event)
@@ -2149,13 +2161,33 @@ class RealtimeAgentHandler:
"delivery": "fallback",
},
)
await self._render_provider_audio(
fallback_audio_completed = await self._render_provider_audio(
ws,
session,
fallback_text,
{},
response_id=fallback_response_id,
response_done_fields={
"source": "hermes",
"delivery": "fallback",
},
)
if not fallback_audio_completed:
await self._send(
ws,
session,
{
"type": SERVER_EVT_RESPONSE_DONE,
"provider": session.provider,
"model": session.model,
"voice": session.voice,
"event_log_path": str(session.event_log_path),
"chat_session_id": session.chat_session_id,
"response_id": fallback_response_id,
"source": "hermes",
"delivery": "fallback",
},
)
return
buffered = list(session.native_forced_summary_buffer)
@@ -2188,6 +2220,7 @@ class RealtimeAgentHandler:
"event_log_path": str(session.event_log_path),
"chat_session_id": session.chat_session_id,
"response_id": None if response_id == "__blank_response_id__" else response_id,
"delivery": "forced_summary",
},
)
self._mark_provider_response_done(session, event)
@@ -2240,8 +2273,26 @@ class RealtimeAgentHandler:
)
if not delivered:
return
if not await self._request_provider_response(ws, session, connection, call.call_id):
if not await self._request_provider_response(
ws,
session,
connection,
call.call_id,
result=result,
transcript=session.promoted_transcript or "",
):
return
confirm = asyncio.create_task(
self._confirm_background_delivery(
ws,
session,
dict(result),
delivery_generation=session.native_forced_summary_generation,
)
)
confirm.add_done_callback(
_log_task_failure(session, "forced_summary_tool_delivery_confirm_task")
)
async def _handle_provider_tool_call(
self,
@@ -2319,8 +2370,28 @@ class RealtimeAgentHandler:
tool_name=call.name,
reason="tool_result_summary",
)
if not await self._request_provider_response(ws, session, connection, call.call_id):
delivery_result = result if call.name == "hermes_run_task" else None
if not await self._request_provider_response(
ws,
session,
connection,
call.call_id,
result=delivery_result,
transcript=str(call.arguments.get("text") or ""),
):
return
if delivery_result is not None:
confirm = asyncio.create_task(
self._confirm_background_delivery(
ws,
session,
dict(delivery_result),
delivery_generation=session.native_forced_summary_generation,
)
)
confirm.add_done_callback(
_log_task_failure(session, "foreground_delivery_confirm_task")
)
def _provider_response_had_audio(
self,
@@ -2393,13 +2464,6 @@ class RealtimeAgentHandler:
await connection.send_tool_result(call_id, result)
return True
except Exception as exc:
await self._send_error(
ws,
session,
"Realtime provider closed before Hermes result could be delivered.",
error_code="provider_tool_result_send_failed",
call_id=call_id,
)
self._log(
session,
"voice.provider_tool_result_send_failed",
@@ -2409,6 +2473,20 @@ class RealtimeAgentHandler:
"error": str(exc),
},
)
if _result_answer_text(result):
await self._speak_fallback_answer(
ws,
session,
result,
reason="provider_tool_result_send_failed",
)
await self._send_error(
ws,
session,
"Realtime provider closed before Hermes result could be delivered.",
error_code="provider_tool_result_send_failed",
call_id=call_id,
)
await self._close_native_session(session, "provider_tool_result_send_failed")
return False
@@ -2418,18 +2496,38 @@ class RealtimeAgentHandler:
session: RealtimeAgentSession,
connection: RealtimeAgentConnection,
call_id: str,
*,
result: dict[str, Any] | None = None,
transcript: str = "",
) -> bool:
instructions = None
exact_text = None
if result is not None:
session.native_forced_summary_generation += 1
session.native_forced_summary_active = True
session.native_forced_summary_done = False
session.native_forced_summary_response_id = None
session.native_forced_summary_result = dict(result)
session.native_forced_summary_buffer.clear()
session.native_forced_summary_committed = False
session.native_forced_summary_text_parts.clear()
instructions = _result_delivery_prompt(
session.result_delivery,
transcript,
result,
)
if (
session.result_delivery == "speak_verbatim"
and not _answer_is_structured(result)
):
exact_text = _forced_summary_fallback_text(result)
try:
await connection.request_response()
await connection.request_response(
instructions=instructions,
exact_text=exact_text,
)
return True
except Exception as exc:
await self._send_error(
ws,
session,
"Realtime provider closed before it could summarize the Hermes result.",
error_code="provider_response_request_failed",
call_id=call_id,
)
self._log(
session,
"voice.provider_response_request_failed",
@@ -2439,6 +2537,20 @@ class RealtimeAgentHandler:
"error": str(exc),
},
)
if result is not None and _result_answer_text(result):
await self._speak_fallback_answer(
ws,
session,
result,
reason="provider_response_request_failed",
)
await self._send_error(
ws,
session,
"Realtime provider closed before it could summarize the Hermes result.",
error_code="provider_response_request_failed",
call_id=call_id,
)
await self._close_native_session(session, "provider_response_request_failed")
return False
@@ -2455,6 +2567,7 @@ class RealtimeAgentHandler:
if session.native_forced_hermes_turn_active:
return
session.native_forced_hermes_turn_active = True
session.native_forced_summary_generation += 1
session.native_forced_summary_active = False
session.native_forced_summary_done = False
session.native_forced_summary_response_id = None
@@ -2507,6 +2620,7 @@ class RealtimeAgentHandler:
# keeps the realtime voice; speak_verbatim differs only in the
# instructions (read the answer as written vs. natural summary).
# The forced-summary validator + relay-TTS fallback backstop both.
session.native_forced_summary_generation += 1
session.native_forced_summary_active = True
session.native_forced_summary_done = False
session.native_forced_summary_response_id = None
@@ -2517,10 +2631,17 @@ class RealtimeAgentHandler:
with contextlib.suppress(Exception):
await connection.cancel_response()
try:
exact_text = None
if (
session.result_delivery == "speak_verbatim"
and not _answer_is_structured(result)
):
exact_text = _forced_summary_fallback_text(result)
await connection.request_response(
instructions=_result_delivery_prompt(
session.result_delivery, transcript, result
)
),
exact_text=exact_text,
)
except Exception as exc: # noqa: BLE001 - the answer must still land
self._log(
@@ -2541,7 +2662,12 @@ class RealtimeAgentHandler:
# delivery lands within the confirm window, force a text emit so
# a dead/stalled provider response can't silently eat the answer.
confirm = asyncio.create_task(
self._confirm_background_delivery(ws, session, dict(result))
self._confirm_background_delivery(
ws,
session,
dict(result),
delivery_generation=session.native_forced_summary_generation,
)
)
confirm.add_done_callback(
_log_task_failure(session, "delivery_confirm_task")
@@ -3200,7 +3326,12 @@ class RealtimeAgentHandler:
# answer is force-emitted as text — it can never be silently lost to
# filler the validator missed, a dead response, or a stray cancel.
confirm = asyncio.create_task(
self._confirm_background_delivery(ws, session, dict(result))
self._confirm_background_delivery(
ws,
session,
dict(result),
delivery_generation=session.native_forced_summary_generation,
)
)
confirm.add_done_callback(
_log_task_failure(session, "delivery_confirm_task")
@@ -3211,11 +3342,15 @@ class RealtimeAgentHandler:
ws: web.WebSocketResponse,
session: RealtimeAgentSession,
result: dict[str, Any],
*,
delivery_generation: int,
) -> None:
"""Force a text emit when a spoken background delivery never lands."""
await asyncio.sleep(_DELIVERY_CONFIRM_SECONDS)
if session.closed:
return
if session.native_forced_summary_generation != delivery_generation:
return
if session.native_forced_summary_done or session.native_forced_summary_committed:
return
self._log(
@@ -3332,7 +3467,12 @@ class RealtimeAgentHandler:
# Same delivered-or-alarm guarantee as the attached background path:
# a resume-injected summary that never lands is force-emitted as text.
confirm = asyncio.create_task(
self._confirm_background_delivery(ws, session, dict(result))
self._confirm_background_delivery(
ws,
session,
dict(result),
delivery_generation=session.native_forced_summary_generation,
)
)
confirm.add_done_callback(
_log_task_failure(session, "delivery_confirm_task")
@@ -3472,13 +3612,33 @@ class RealtimeAgentHandler:
"delivery": "fallback",
},
)
await self._render_provider_audio(
fallback_audio_completed = await self._render_provider_audio(
ws,
session,
fallback_text,
{},
response_id=response_id,
response_done_fields={
"source": "hermes",
"delivery": "fallback",
},
)
if not fallback_audio_completed:
await self._send(
ws,
session,
{
"type": SERVER_EVT_RESPONSE_DONE,
"provider": session.provider,
"model": session.model,
"voice": session.voice,
"event_log_path": str(session.event_log_path),
"chat_session_id": session.chat_session_id,
"response_id": response_id,
"source": "hermes",
"delivery": "fallback",
},
)
async def _preempt_pending_forced_summary(
self,
@@ -3500,6 +3660,7 @@ class RealtimeAgentHandler:
was_active = session.native_forced_summary_active
pending = session.native_forced_summary_result
committed = session.native_forced_summary_committed
session.native_forced_summary_generation += 1
session.native_forced_summary_active = False
session.native_forced_summary_done = False
session.native_forced_summary_response_id = None
@@ -3537,6 +3698,7 @@ class RealtimeAgentHandler:
cancel_current: bool = True,
) -> None:
transcript = session.promoted_transcript or ""
session.native_forced_summary_generation += 1
session.native_forced_summary_active = True
session.native_forced_summary_done = False
session.native_forced_summary_response_id = None
@@ -4276,7 +4438,8 @@ class RealtimeAgentHandler:
payload: dict[str, Any],
*,
response_id: str | None = None,
) -> None:
response_done_fields: dict[str, Any] | None = None,
) -> bool:
queue: asyncio.Queue[dict[str, Any] | None] = asyncio.Queue()
loop = asyncio.get_running_loop()
output_path = session.event_log_path.with_suffix(".wav")
@@ -4332,7 +4495,7 @@ class RealtimeAgentHandler:
except (ProviderUnavailable, ProviderRunError) as exc:
session.floor.release(FloorMouth.RELAY_TTS)
await self._send_error(ws, session, str(exc), provider=session.provider)
return
return False
except Exception as exc:
session.floor.release(FloorMouth.RELAY_TTS)
await self._send_error(
@@ -4341,7 +4504,7 @@ class RealtimeAgentHandler:
f"realtime agent provider failed: {exc.__class__.__name__}: {exc}",
provider=session.provider,
)
return
return False
await self._send(
ws,
@@ -4361,24 +4524,23 @@ class RealtimeAgentHandler:
if not keep_tap:
with contextlib.suppress(OSError):
Path(response.audio_path).unlink()
await self._send(
ws,
session,
{
"type": "voice.response.done",
"provider": response.provider,
"model": response.model,
"voice": response.voice,
"audio_path": str(response.audio_path) if keep_tap else "",
"event_log_path": str(session.event_log_path),
"chat_session_id": session.chat_session_id,
"final_text": text,
"response_id": response_id,
"metrics": response.metrics.to_dict(),
"metadata": _safe_metadata(response.metadata),
},
)
response_done = {
"type": "voice.response.done",
"provider": response.provider,
"model": response.model,
"voice": response.voice,
"audio_path": str(response.audio_path) if keep_tap else "",
"event_log_path": str(session.event_log_path),
"chat_session_id": session.chat_session_id,
"final_text": text,
"response_id": response_id,
"metrics": response.metrics.to_dict(),
"metadata": _safe_metadata(response.metadata),
}
response_done.update(response_done_fields or {})
await self._send(ws, session, response_done)
session.floor.release(FloorMouth.RELAY_TTS)
return True
def _provider_options(
self,
@@ -5532,7 +5694,7 @@ def _bad_forced_summary_reason(text: str, answer: str | None = None) -> str | No
if len(normalized) <= 80 and any(
phrase in normalized
for phrase in ("got it", "sure", "okay", "ok")
) and "check" in normalized:
) and "check" in normalized and normalized != answer_normalized:
return "short_acknowledgement"
# Positive check (blocklists chase phrasings; this one doesn't): the
# summary must share content with the answer it claims to deliver.
+198 -4
View File
@@ -324,6 +324,8 @@ class FakeNativeConnection:
self.text_inputs: list[str] = []
self.tool_results: list[tuple[str, dict[str, Any]]] = []
self.context_items: list[tuple[str, str]] = []
self.exact_response_texts: list[str] = []
self.response_instructions: list[str] = []
self.request_response_count = 0
self.clear_count = 0
self.cancelled = False
@@ -366,6 +368,9 @@ class FakeNativeConnection:
self.request_response_count += 1
if instructions:
self.text_inputs.append(instructions)
self.response_instructions.append(instructions)
if exact_text:
self.exact_response_texts.append(exact_text)
async def close(self) -> None:
self.closed = True
@@ -1847,18 +1852,31 @@ class RealtimeAgentRoutesTests(AioHTTPTestCase):
"realtime_agent",
)
self._server().realtime_agent.sessions[
body["session_id"]
].result_delivery = "speak_verbatim"
await ws.send_json({"type": "playback.drained", "call_id": "call-1"})
for _ in range(20):
if fake_provider.connection.request_response_count:
break
await asyncio.sleep(0.01)
self.assertEqual(fake_provider.connection.request_response_count, 1)
self.assertEqual(
fake_provider.connection.exact_response_texts,
["Sure, I checked that."],
)
self.assertIn(
"word for word as written",
fake_provider.connection.response_instructions[-1],
)
finally:
await ws.close()
async def test_provider_native_suppresses_tool_response_done_until_followup(
self,
) -> None:
previous_confirm = broker_module._DELIVERY_CONFIRM_SECONDS
broker_module._DELIVERY_CONFIRM_SECONDS = 0.5
token = await self._make_session()
fake_broker = FakeHermesToolBroker()
fake_provider = FakeNativeProvider()
@@ -1921,6 +1939,13 @@ class RealtimeAgentRoutesTests(AioHTTPTestCase):
await fake_provider.connection.emit(
ProviderEvent(ProviderEventKind.RESPONSE_STARTED, response_id="resp-followup")
)
await fake_provider.connection.emit(
ProviderEvent(
ProviderEventKind.OUTPUT_TEXT_DELTA,
response_id="resp-followup",
payload={"delta": "Sure, I checked that."},
)
)
await fake_provider.connection.emit(
ProviderEvent(
ProviderEventKind.AUDIO_DELTA,
@@ -1936,7 +1961,7 @@ class RealtimeAgentRoutesTests(AioHTTPTestCase):
)
events: list[dict[str, Any]] = []
for _ in range(10):
for _ in range(30):
event = await self._next_ws_event(ws)
events.append(event)
if event["type"] == "voice.response.done":
@@ -1947,9 +1972,69 @@ class RealtimeAgentRoutesTests(AioHTTPTestCase):
]
self.assertEqual(len(done_events), 1)
self.assertEqual(done_events[0]["response_id"], "resp-followup")
delivery_events = [
event
for event in events
if event["type"]
in {
"voice.response.started",
"voice.response.delta",
"voice.output_audio.delta",
"voice.output_audio.done",
"voice.response.done",
}
]
self.assertTrue(delivery_events, events)
self.assertTrue(
all(event.get("delivery") == "forced_summary" for event in delivery_events),
delivery_events,
)
await asyncio.sleep(0.55)
log_text = self._server().realtime_agent.sessions[
body["session_id"]
].event_log_path.read_text(encoding="utf-8")
self.assertNotIn("voice.realtime_agent.delivery_unconfirmed", log_text)
finally:
broker_module._DELIVERY_CONFIRM_SECONDS = previous_confirm
await ws.close()
async def test_delivery_confirm_ignores_a_stale_generation(self) -> None:
previous_confirm = broker_module._DELIVERY_CONFIRM_SECONDS
broker_module._DELIVERY_CONFIRM_SECONDS = 0
token = await self._make_session()
fake_provider = FakeNativeProvider()
handler = self._server().realtime_agent
handler.native_providers["xai_realtime"] = fake_provider
try:
resp = await self.client.post(
"/voice/realtime-agent/session",
json={
"provider": "xai_realtime",
"model": "grok-voice-latest",
"voice": "leo",
"chat_session_id": "chat-123",
},
headers=self._bearer(token),
)
self.assertEqual(resp.status, 200)
body = await resp.json()
session = handler.sessions[body["session_id"]]
session.native_forced_summary_generation = 2
session.native_forced_summary_done = False
session.native_forced_summary_committed = False
event_seq_before = session.event_seq
await handler._confirm_background_delivery(
None, # type: ignore[arg-type] - stale generation returns before WS use
session,
{"answer": "stale answer"},
delivery_generation=1,
)
self.assertEqual(event_seq_before, session.event_seq)
finally:
broker_module._DELIVERY_CONFIRM_SECONDS = previous_confirm
async def test_provider_native_suppresses_status_response_done_until_followup(
self,
) -> None:
@@ -2451,15 +2536,124 @@ class RealtimeAgentRoutesTests(AioHTTPTestCase):
)
events: list[dict[str, Any]] = []
for _ in range(20):
for _ in range(60):
event = await self._next_ws_event(ws)
events.append(event)
if event["type"] == "voice.error":
if event.get("error_code") == "provider_tool_result_send_failed":
break
error = next(event for event in events if event["type"] == "voice.error")
error = next(
event
for event in events
if event.get("error_code") == "provider_tool_result_send_failed"
)
self.assertEqual(error["error_code"], "provider_tool_result_send_failed")
self.assertEqual(error["call_id"], "call-closing")
fallback = [
event
for event in events
if event["type"] == "voice.response.delta"
and event.get("delivery") == "fallback"
]
self.assertEqual(1, len(fallback), events)
self.assertEqual("Sure, I checked that.", fallback[0]["delta"])
self.assertLess(events.index(fallback[0]), events.index(error))
self.assertEqual(
1,
sum(event["type"] == "voice.response.done" for event in events),
events,
)
for _ in range(500):
if handler.sessions[body["session_id"]].closed:
break
await asyncio.sleep(0.01)
self.assertTrue(handler.sessions[body["session_id"]].closed)
finally:
await ws.close()
async def test_provider_native_response_request_failure_falls_back_once(self) -> None:
token = await self._make_session()
fake_broker = FakeHermesToolBroker()
fake_provider = FakeNativeProvider()
fake_provider.connection.fail_request_response = True
handler = self._server().realtime_agent
handler.hermes = fake_broker
handler.native_providers["xai_realtime"] = fake_provider
resp = await self.client.post(
"/voice/realtime-agent/session",
json={
"provider": "xai_realtime",
"model": "grok-voice-latest",
"voice": "leo",
"chat_session_id": "chat-123",
},
headers=self._bearer(token),
)
self.assertEqual(resp.status, 200)
body = await resp.json()
ws = await self.client.ws_connect(
body["websocket_path"],
headers=self._bearer(token),
)
try:
self.assertEqual((await self._next_ws_event(ws))["type"], "voice.session.ready")
await fake_provider.connection.emit(
ProviderEvent(
ProviderEventKind.FUNCTION_CALL_COMPLETED,
payload={
"call": ToolCallEvent(
call_id="call-response-fails",
name="hermes_run_task",
arguments={
"text": "Check before the response request fails.",
"session_id": "chat-123",
},
)
},
)
)
for _ in range(20):
event = await self._next_ws_event(ws)
if event["type"] == "voice.playback_drain.requested":
break
else:
self.fail("expected playback drain request")
await ws.send_json(
{"type": "playback.drained", "call_id": "call-response-fails"}
)
events: list[dict[str, Any]] = []
for _ in range(60):
event = await self._next_ws_event(ws)
events.append(event)
if event.get("error_code") == "provider_response_request_failed":
break
error = next(
event
for event in events
if event.get("error_code") == "provider_response_request_failed"
)
self.assertEqual(error["error_code"], "provider_response_request_failed")
fallback = [
event
for event in events
if event["type"] == "voice.response.delta"
and event.get("delivery") == "fallback"
]
self.assertEqual(1, len(fallback), events)
self.assertEqual("Sure, I checked that.", fallback[0]["delta"])
self.assertLess(events.index(fallback[0]), events.index(error))
self.assertEqual(
1,
sum(event["type"] == "voice.response.done" for event in events),
events,
)
for _ in range(500):
if handler.sessions[body["session_id"]].closed:
break
await asyncio.sleep(0.01)
self.assertTrue(handler.sessions[body["session_id"]].closed)
finally:
await ws.close()
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "hermes-relay"
version = "1.4.0"
version = "1.4.1"
description = "Hermes-Relay plugin — Android device control toolset, QR pairing CLI, and WSS relay server for hermes-agent"
requires-python = ">=3.11"
dependencies = [