Compare commits

..
Author SHA1 Message Date
Bailey Dixon 0a5649d016 docs: refresh upstream API baseline references 2026-06-04 21:59:52 -04:00
Bailey Dixon 7d667b9096 feat(android): prefer upstream skills endpoint (#63) 2026-06-04 21:03:59 -04:00
Bailey Dixon f7edffcd81 Merge remote-tracking branch 'origin/main' into dev
# Conflicts:
#	CHANGELOG.md
#	DEVLOG.md
#	RELEASE_NOTES.md
2026-05-26 21:36:26 -04:00
Bailey DixonandClaude Opus 4.7 b53df95bbf docs: correct stale VoicePlayer "MediaPlayer" references to Media3 ExoPlayer
VoicePlayer was migrated to a single Media3 ExoPlayer (gapless TTS queue)
in the V5 voice-quality pass, but two current-state descriptions still
called it a MediaPlayer — the CLAUDE.md Key Files row and the decisions.md
voice references. The CLAUDE.md drift actively misled a crash diagnosis.
Also note audioSessionId is now a thread-safe @Volatile cache.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 21:00:06 -04:00
Bailey Dixon f879fbbd51 Merge pull request #60 from Codename-11/fix/voice-barge-in-wrong-thread
fix(android): read ExoPlayer audio session id on main thread
2026-05-26 20:41:28 -04:00
Bailey DixonandClaude Opus 4.7 a586f3dd60 fix(android): read ExoPlayer audio session id on main thread
Voice chat crashed the instant Hermes began replying when barge-in was
enabled and audio used the legacy /voice/synthesize (Media3) path:

  IllegalStateException: Player is accessed on the wrong thread.
  Current thread: 'DefaultDispatcher-worker-4', Expected thread: 'main'

BargeInListener runs its mic reader on Dispatchers.IO and, to attach
AcousticEchoCanceler, polls an audioSessionIdProvider lambda. On the
legacy path that provider read exoPlayer.audioSessionId directly.
ExoPlayer is thread-confined — its getAudioSessionId() getter calls
verifyApplicationThread() and throws off-main. (The realtime PCM path
was immune: it provides an AudioTrack session id, which is thread-safe.)

VoicePlayer.audioSessionId now serves a @Volatile cache populated from
main-thread Media3 callbacks (AnalyticsListener.onAudioSessionIdChanged
plus a belt-and-braces read in onIsPlayingChanged), so it is safe to
read from any thread.

Adds VoicePlayerTest coverage: the getter reflects the cached id, never
re-invokes the thread-confined getter, and defaults to 0 before the
audio track is allocated.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 20:31:23 -04:00
Bailey Dixon d41e1b659b docs: add relay architecture spec page 2026-05-26 19:22:06 -04:00
Bailey Dixon de9988322b Merge pull request #59 from Codename-11/fix/realtime-voice-error-display
fix(voice): classify realtime voice.error for a clear, actionable message
2026-05-24 17:51:59 -04:00
Bailey DixonandClaude Opus 4.7 ff09c6116b fix(voice): classify realtime voice.error instead of showing raw provider string
The in-stream voice.error handler set uiState.error to the raw relay message
(e.g. 'xAI Realtime auth is not configured ...'). Route it through surfaceError
-> classifyError('voice_config') so provider-auth and other relay failures show
a clear, actionable banner ('Realtime provider auth unavailable ...') plus a
one-shot errorEvents snackbar with a Voice settings action, matching how the
result-failure path already surfaces errors. Raw detail is still recorded to the
Diagnostics log.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 17:42:26 -04:00
Bailey Dixon 7c5b7729c1 Merge pull request #58 from Codename-11/feature/realtime-persistent-session
feat(voice): persistent Realtime Agent conversation (one socket across turns)
2026-05-24 16:13:35 -04:00
Bailey DixonandClaude Opus 4.7 304d8e26c8 feat(voice): persistent realtime-agent session (VoiceViewModel wiring)
Wire the persistent session end to end. Realtime Agent voice now opens one
provider session/socket on the first turn (runRealtimeAgent persistent mode in
realtimeSessionJob) and feeds subsequent utterances on realtimeTurnChannel, so
the provider keeps the live conversation across turns.

- Per-turn event state hoisted to fields so the session-lived callback serves
  every turn; submitRealtimeTurn / the open path reset it per turn.
- onRealtimeTurnComplete finalizes each spoken turn (re-arms continuous listen);
  closeRealtimeSession tears down on exit / engine switch / onCleared / error.
- VoicePreferences.realtimePersistentSession (default true) + a Voice Settings ->
  Realtime Agent -> Persistent session toggle fall back to the one-shot path.

Compiles clean (compileSideloadDebugKotlin). Needs on-device validation
(multi-turn follow-ups, barge-in, background promotion mid-conversation,
exit/re-enter) — flag lets you fall back without a rebuild.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 16:03:02 -04:00
Bailey DixonandClaude Opus 4.7 742d899042 feat(voice): persistent realtime-agent session foundation (client)
Add opt-in persistent mode to RelayVoiceClient.runRealtimeAgent: when a
turnInputs ReceiveChannel is supplied, the WebSocket stays open across turns
(voice.response.done fires onTurnComplete instead of closing), subsequent
RealtimeTurnInputs are sent on the same socket with monotonic chunk ids, the
idle/turn guards scope to an active turn only, and the call ends when the channel
closes. One-shot path (turnInputs=null) is byte-for-byte unchanged.

Relay needs no change — _handle_provider_native_ws already loops over
input_audio/commit/response on one socket. Plan: docs/plans/2026-05-24-realtime-persistent-session.md.

VoiceViewModel wiring follows in the next commit.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 15:27:48 -04:00
Bailey Dixon 783fc34e01 Merge pull request #57 from Codename-11/feature/realtime-voice-loop-and-logging
feat(voice): log realtime Hermes run lifecycle (ADR 33 observability)
2026-05-24 15:18:31 -04:00
Bailey DixonandClaude Opus 4.7 ac51a95842 feat(voice): log realtime Hermes run lifecycle (ADR 33 observability)
The hermes.run.* event handlers mutated UI state silently, so a promoted
background run was invisible in logcat even though it ran. Add Log.i for
run started / progress (tier/floor/status) / promoted / background_completed /
cancelled so the background-task lifecycle is traceable on-device.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 14:23:40 -04:00
Bailey Dixon 697610d92c Merge pull request #56 from Codename-11/feature/realtime-background-hermes-runs
feat(realtime): ADR 33 — background Hermes runs in Realtime Agent voice
2026-05-24 13:14:00 -04:00
Bailey DixonandClaude Opus 4.7 f300531054 docs(realtime): close ADR 33 Phase 0 — both verdicts hold-floor-ok
Record unconditional per-provider verdicts and mark Phase 0 done:
- OpenAI: hold-floor-ok (empirical 10/20/30s idle probe).
- xAI: hold-floor-ok — not conditional. The shipping Realtime Agent already
  holds xai_realtime sessions open across between-turn idle (turn_detection:None
  + resume TTL) with no idle-close reports; a relay-host probe is a regression
  check, not a precondition.

Also records that the spike's premise was superseded: Tier B closes the pending
provider call with an interim ack rather than holding an open response, so the
socket only sees the normal between-turns idle gap. No provider needs the
must-reopen fallback; default-on is unblocked.

Updates realtime-voice-poc.md findings, the plan's Phase 0 acceptance (Status:
DONE), and ADR 33's Phase 0 line (RESOLVED).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 11:21:30 -04:00
Bailey DixonandClaude Opus 4.7 0c7c21964d docs(realtime): ADR 33 Phase 0 verdict — OpenAI idle hold-floor-ok (empirical)
Ran scripts/realtime-provider-idle-probe.py against the live OpenAI realtime API
(VOICE_TOOLS_OPENAI_KEY): the session survived 10s/20s/30s quiescent idle windows
and returned clean audio on every post-idle turn -> verdict hold-floor-ok.
xAI recorded analytically as hold-floor-ok (no dev-box creds; same
turn_detection:None multi-turn model + the promotion path closes the pending
call rather than holding an open response) pending relay-host confirmation.

Fills the docs/realtime-voice-poc.md Idle tolerance findings table, satisfying
the Phase 0 acceptance (a documented per-provider verdict). Logs an incidental
OpenAI session.audio.output.format.rate schema-drift follow-up.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 11:19:22 -04:00
Bailey DixonandClaude Opus 4.7 ab1ffdd2aa docs(devlog): ADR 33 background Hermes runs session entry
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 11:14:03 -04:00
Bailey DixonandClaude Opus 4.7 ab10097f55 feat(realtime): ADR 33 Phase 3 (Android + docs) — promotion UI + event handling
Android:
- RealtimeVoiceEvent gains tier/floor; parse hermes.run.promoted +
  hermes.run.background_completed in VoiceViewModel, surfaced as a
  BackgroundRunState 'working on it' chip in VoiceModeOverlay (cleared on
  background_completed / cancel).
- RealtimeVoiceConfig gains a promotion block; new
  RelayVoiceClient.updateRealtimeAgentPromotion() PATCH.
- Voice Settings → Realtime Agent → Background tasks: promote toggle, spoken
  handoff toggle, and result-delivery segmented control, persisted to the relay.

Docs:
- CHANGELOG [Unreleased], relay-protocol.md (ADR 33 background-runs section),
  user-docs/features/voice.md (Background tasks).

Kotlin compiles clean under ./gradlew lint (the only 2 lint errors are in the
gitignored local.properties, absent in CI). Python realtime suite 58 tests green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 11:13:04 -04:00
Bailey DixonandClaude Opus 4.7 e294a571a4 feat(realtime): ADR 33 Phase 3 (relay) — default-on, Tier C durable, config surface
- Flip realtime_voice_promotion_enabled default to true. Safe because the
  promotion path closes the pending provider call with an interim ack rather
  than holding an open response, so the socket only sees the normal between-turns
  idle gap. Phase 0 probe still recommended to confirm per-provider survival.
- Tier C: hermes_run_task(mode='background') detaches immediately (tier=durable),
  even when grace-period promotion is off. Schema 'mode' enum gains 'background'.
- Expose promotion settings in /voice/realtime-agent config GET (promotion block)
  and accept them in PATCH (_validate_config_updates) so Android can read/write.

test_realtime_promotion gains the Tier C immediate-detach case (58 tests green).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 10:51:22 -04:00
Bailey DixonandClaude Opus 4.7 35893e239f feat(realtime): ADR 33 Phase 2 — Tier B grace-period promotion (default off)
Long Hermes runs in Realtime Agent no longer block the provider event pump.
_run_brokered_tool now shields the run task and waits promote_after_ms; if the
run is still in flight, it detaches to the background (tier=promoted) and returns
control to the pump. _deliver_background_result awaits the task, emits
hermes.run.background_completed, waits for the floor to clear, then speaks the
result once via the existing forced-summary path.

- New events: hermes.run.promoted, hermes.run.background_completed (models.py)
- New realtime_voice settings (config.py + profile_voice.py): promotion_enabled
  (default false), promote_after_ms (6000), background_default_mode, spoken_handoff,
  progress_spoken_after_ms, progress_repeat_ms, result_delivery, max_background_runs
- Provider-tool-call path closes the pending call with an interim background ack
  so the socket isn't left awaiting output; forced path speaks a handoff line
- Cancel (response.cancel / hermes_cancel) stops the background task; background
  delivery task cancelled on session close
- Completion replays through the event ring on resume (detach-safe)

test_realtime_promotion: promote+pump-responsive, short=no-promote, cancel,
detach-resume-replays. Full realtime suite (53 tests) green; pre-existing
unrelated xAI-OAuth-pool test failure noted.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 10:48:14 -04:00
Bailey DixonandClaude Opus 4.7 06332f8f4e feat(realtime): ADR 33 Phase 1 — relay audio floor owner
Add plugin/relay/realtime_agent/floor.py: a pure, single-owner audio floor
(provider | relay_tts | android_filler mouths; idle/provider_speaking/
hermes_filler/result_pending labels) that makes explicit the serialization the
blocking await provided implicitly. Wire it into the broker behavior-
preservingly:

- acquire/release PROVIDER on AUDIO_DELTA/AUDIO_DONE (+ RESPONSE_DONE safety net)
- acquire/release RELAY_TTS around _render_provider_audio
- gate spoken filler by floor.can_speak(ANDROID_FILLER); stamp floor + tier on
  hermes.run.progress

Adds session fields hermes_run_tier + floor. No audible change (today's flow has
no contention); invariants proven in test_realtime_floor (background result never
barges, filler suppressed while provider speaks, relay-TTS only when owned).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 10:37:25 -04:00
Bailey DixonandClaude Opus 4.7 60623c1751 feat(realtime): ADR 33 Phase 0 provider idle-tolerance probe + docs
Add scripts/realtime-provider-idle-probe.py and the Idle tolerance section in
docs/realtime-voice-poc.md. The probe holds an xAI/OpenAI realtime socket
quiescent across idle windows and reports a per-provider verdict
(hold-floor-ok | needs-keepalive | must-reopen) that selects each provider's
Tier B strategy. Verdict gates ADR 33 default-on promotion.

Also lands ADR 33 and the companion implementation plan.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 10:33:12 -04:00
Bailey DixonandClaude Opus 4.7 024ee6c1db docs(release): scrub signing-cert identity from v0.8.0 release notes
The v0.8.0 notes' Verification section published the signing cert CN
(a personal name) in the live GitHub Release. Prior releases never
listed the cert identity — generalize to 'release-signed with the
production upload keystore' to match the house style. Live release body
already updated via gh release edit.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-23 22:30:49 -04:00
35 changed files with 2868 additions and 148 deletions
+12
View File
@@ -6,6 +6,18 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this
## [Unreleased]
### Added
- **Persistent Realtime Agent conversation.** Realtime Agent voice now keeps one provider session/socket open across turns instead of creating a fresh session per utterance, so the provider retains the live conversation (follow-up references work) and turns skip session-setup latency. The relay needed no change — it already supported multiple turns on one socket. A **Voice Settings → Realtime Agent → Persistent session** toggle (default on) falls back to the legacy per-utterance path. See `docs/plans/2026-05-24-realtime-persistent-session.md`.
- **Background Hermes runs in Realtime Agent voice (ADR 33).** Long Hermes tasks no longer freeze the realtime conversation. A run that exceeds a grace window is promoted to a tracked background task: the provider speaks a short handoff ("I'm on it"), the conversation stays responsive, and the answer is spoken once the run finishes. `hermes_run_task(mode="background")` starts a durable run immediately. New relay events `hermes.run.promoted` and `hermes.run.background_completed`, plus `tier`/`floor` fields on `hermes.run.progress`.
- **Relay audio floor owner.** A single-owner audio floor (provider / relay-TTS / Android-filler) makes explicit the serialization that the old blocking design provided implicitly, so a completed background result never barges in and two voices never overlap.
- **Voice Settings → Realtime Agent → Background tasks.** New controls to enable/disable promotion, toggle the spoken handoff, and choose result delivery (speak when idle / notify / show only). A persistent "working on it" chip appears in the voice overlay while a background task runs.
- **Provider idle-tolerance probe.** `scripts/realtime-provider-idle-probe.py` records a per-provider verdict (hold-floor-ok / needs-keepalive / must-reopen) for holding a realtime socket quiescent during a background run; see `docs/realtime-voice-poc.md`.
## [0.8.1] - 2026-05-26
### Fixed
+18 -17
View File
@@ -33,25 +33,26 @@ Chat goes directly to the API server via HTTP/SSE. The API key (Bearer token) is
| `GET /health` | Health check | — |
| `GET/POST/PATCH/DELETE /api/jobs/*` | Cron job management (api_server surface) | — |
**Non-standard endpoints (provided by fork OR by plugin bootstrap):**
**Baseline upstream endpoints vs compatibility endpoints:**
These endpoints are not in stock upstream `gateway/platforms/api_server.py`. There are three ways a hermes-agent install can serve them:
Upstream hermes-agent now has a native baseline for API Server session control and skill/toolset discovery:
1. **Codename-11 fork** (`feat/session-api` branch, deployed on the `axiom` branch) — adds them natively. Submitted upstream as PR [#8556](https://github.com/NousResearch/hermes-agent/pull/8556) *"feat(api-server): add session management API for frontend clients"* — scope is broader than the title: sessions CRUD + session chat/stream + memory + skills + config + available-models.
2. **Bootstrap injection** (`hermes_relay_bootstrap/`) — monkey-patches aiohttp on startup via `.pth` file. Does NOT inject `/api/sessions/{id}/chat/stream` — use `/v1/runs` for chat.
3. **Upstream-merged** (post PR #8556) — bootstrap auto-detects and no-ops.
1. **Native upstream** — commit [`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde), merged via PR [#33134](https://github.com/NousResearch/hermes-agent/pull/33134), salvaged the focused session-control work from closed PR [#29302](https://github.com/NousResearch/hermes-agent/pull/29302). It provides `/api/sessions/*`, session chat/stream, fork/messages, plus `/v1/skills` and `/v1/toolsets`.
2. **Codename-11 `axiom` fork** — still carries compatibility/client-metadata routes that upstream does not provide yet: `/api/sessions/search`, `/api/memory`, `/api/skills` detail routes, `/api/config`, and `/api/available-models`.
3. **Bootstrap injection** (`hermes_relay_bootstrap/`) — monkey-patches aiohttp on startup via `.pth` file and injects only missing compatibility routes for older or partial upstream builds. It should remain per-route/per-feature, not all-or-nothing.
| Endpoint | Purpose | Provided by |
|----------|---------|-------------|
| `GET /api/sessions` (CRUD) | Session list/create/rename/delete/fork | Fork OR bootstrap OR upstream-merged |
| `GET /api/sessions/{id}/messages` | Conversation history | Fork OR bootstrap OR upstream-merged |
| `GET /api/sessions/search` | Full-text message search | Fork OR bootstrap OR upstream-merged |
| `POST /api/sessions/{id}/chat/stream` | Session-based SSE chat | Fork OR upstream-merged ONLY (NOT bootstrap) |
| `GET /api/config`, `PATCH /api/config` | Personalities + model config | Fork OR bootstrap OR upstream-merged |
| `GET /api/skills`, `/{name}` | Skill discovery (list + detail) | Fork OR bootstrap OR upstream-merged |
| `GET /api/sessions` (CRUD) | Session list/create/rename/delete/fork | Native upstream OR fork OR bootstrap |
| `GET /api/sessions/{id}/messages` | Conversation history | Native upstream OR fork OR bootstrap |
| `GET /api/sessions/search` | Full-text message search | Fork OR bootstrap only |
| `POST /api/sessions/{id}/chat/stream` | Session-based SSE chat | Native upstream OR fork only (NOT bootstrap) |
| `GET /v1/skills` | Skill list metadata | Native upstream OR fork |
| `GET /api/config`, `PATCH /api/config` | Personalities + model config | Fork OR bootstrap only |
| `GET /api/skills`, `/{name}` | Legacy skill discovery/detail routes | Fork OR bootstrap only; Android prefers `/v1/skills` first |
| `PUT /api/skills/toggle` | Enable/disable installed skill | `hermes_cli/web_server.py` dashboard surface; mirrored into bootstrap |
| `GET/POST/PATCH/DELETE /api/memory` | Memory CRUD | Fork OR bootstrap OR upstream-merged |
| `GET /api/available-models` | Provider model list | Fork OR bootstrap OR upstream-merged |
| `GET/POST/PATCH/DELETE /api/memory` | Memory CRUD | Fork OR bootstrap only |
| `GET /api/available-models` | Provider-aware model list | Fork OR bootstrap only |
The Android client probes per-endpoint capability via `HermesApiClient.probeCapabilities()` (returns `ServerCapabilities`). When `streamingEndpoint = "auto"`, `ConnectionViewModel.resolveStreamingEndpoint()` picks `sessions` or `runs` based on the capability snapshot.
@@ -67,7 +68,7 @@ hermes-agent ships a second web server at `hermes_cli/web_server.py` that hosts
## Key Instructions
- **Always verify upstream before assuming an endpoint exists.** Check `gateway/platforms/api_server.py` in hermes-agent. If an endpoint isn't there, document whether bootstrap injects it or it requires the fork.
- If we use a non-standard endpoint, ensure `probeCapabilities()` covers it and the auto-resolver degrades gracefully.
- **Bootstrap maintenance:** Remove `hermes_relay_bootstrap/` in one PR once PR #8556 merges. It's no-op-compatible, so leaving it in place during rollout is harmless.
- **Bootstrap maintenance:** Do not remove `hermes_relay_bootstrap/` just because upstream has native sessions. It can start shrinking only after each Relay-consuming compatibility route has a native replacement or the Android/Desktop clients have migrated away from it.
## Repository Layout
@@ -107,7 +108,7 @@ hermes-android/
│ ├── tools/ # android_navigate.py, android_notifications.py
│ └── dashboard/ # hermes-agent dashboard plugin — manifest, React UI, FastAPI proxy
├── relay_server/ ← Thin compat shim → plugin.relay (legacy entrypoint)
├── hermes_relay_bootstrap/ ← Runtime patch for vanilla upstream; removable after PR #8556
├── hermes_relay_bootstrap/ ← Runtime patch for vanilla/partial upstream compatibility routes
├── skills/devops/hermes-relay-pair/ ← /hermes-relay-pair slash command
├── scripts/ ← dev.bat, bridge-smoke.sh, bump-version.sh
└── docs/ ← spec, decisions, security, relay-server, mcp-tooling
@@ -197,7 +198,7 @@ hermes-android/
| **App — Voice** | |
| `voice/VoiceViewModel.kt` | Voice turn state machine; TTS queue; `ignoreAssistantId`; `errorEvents: SharedFlow` |
| `audio/VoiceRecorder.kt` | MediaRecorder wrapper; perceptual amplitude curve; `.m4a` at 16kHz/64kbps |
| `audio/VoicePlayer.kt` | MediaPlayer + Visualizer; amplitude StateFlow; `awaitCompletion()` via coroutine |
| `audio/VoicePlayer.kt` | Media3 ExoPlayer (gapless TTS queue) + Visualizer; amplitude StateFlow; `awaitCompletion()` via coroutine; `audioSessionId` is a thread-safe `@Volatile` cache |
| `network/RelayVoiceClient.kt` | OkHttp for `/voice/transcribe`, `/synthesize`, `/config` |
| `voice/VoiceBridgeIntentHandler.kt` | Interface routing voice utterances to bridge; impls per flavor via factory |
| `voice/VoiceIntentClassifier.kt` | Regex phone-control classifier (sideload only); false-negatives preferred over false-positives |
@@ -230,7 +231,7 @@ hermes-android/
| `plugin/pair.py` | QR payload builder + CLI; `build_payload(sign=True)`; `--register-code` fallback |
| `install.sh` | Canonical installer — 6 steps; idempotent; drops `hermes-relay-update` shim |
| `uninstall.sh` | Canonical uninstaller; reverses install.sh; never touches `.env` or `state.db` |
| `hermes_relay_bootstrap/` | Runtime patch for vanilla upstream; no-op on fork/upstream-merged; remove after PR #8556 |
| `hermes_relay_bootstrap/` | Runtime patch for vanilla/partial upstream compatibility routes; shrink per route group after native parity or client migration |
| **Plugin — Dashboard** | |
| `plugin/dashboard/manifest.json` | Declares tab, entry bundle, and FastAPI module for hermes-agent discovery |
| `plugin/dashboard/plugin_api.py` | FastAPI router proxying 5 routes to relay over loopback; `/pairing` body = API-server overrides (host/port/tls/api_key), relay URL auto-derived |
+40 -1
View File
@@ -1,5 +1,25 @@
# Hermes-Relay — Dev Log
## 2026-06-05 — Refresh upstream baseline docs after native session merge
**Context.** Upstream Hermes Agent merged native API Server session controls as commit `f7527b0` via PR #33134, and Android now prefers native `/v1/skills` with legacy fallback. Several Relay docs still treated PR #8556/#29302 as the future removal trigger for `hermes_relay_bootstrap/`.
**What changed.** Updated contributor/user docs to separate native upstream baseline routes (`/api/sessions/*`, `/v1/skills`, `/v1/toolsets`) from Relay compatibility routes that still require `axiom` or bootstrap (`/api/sessions/search`, `/api/memory`, `/api/config`, legacy skill detail routes, `/api/available-models`, voice aliases). The bootstrap retirement rule is now per route group, not wholesale deletion after one upstream session merge.
**Verification.** Ran stale-reference searches for `#8556`, `#29302`, and legacy skills routes after edits; remaining references are historical devlog/plans or explicitly marked compatibility/fallback.
---
## 2026-06-05 — Prefer upstream `/v1/skills` with legacy fallback
**Context.** Upstream Hermes Agent now has baseline skill/session API surface area, while Axiom's fork still preserves richer Relay-specific `/api/*` compatibility routes. The Android client should begin consuming upstream-compatible skill listings when present without breaking older fork/bootstrap installs.
**What changed.** `HermesApiClient.getSkills()` now tries `/v1/skills` first, then falls back to `/api/skills`. Skill parsing accepts upstream OpenAI-style list envelopes (`{"object":"list","data":[...]}`), legacy fork envelopes (`{"skills":[...]}` / `{"items":[...]}`), and direct arrays.
**Verification.** Added pure Kotlin unit coverage for endpoint order and `/v1/skills` `data` parsing. Verified with `ANDROID_HOME=$HOME/Android/Sdk ./gradlew :app:testGooglePlayDebugUnitTest --tests 'com.hermesandroid.relay.network.HermesApiClientTest'` → BUILD SUCCESSFUL. `git diff --check` passes.
---
## 2026-05-26 — Fix voice-mode crash: ExoPlayer audio session id read off-main (barge-in + legacy TTS)
**Report.** Discord user, sideload latest: voice chat crashes the instant Hermes starts answering — "I hear just 2 letters and it crashes." Stack: `IllegalStateException: Player is accessed on the wrong thread. Current thread: 'DefaultDispatcher-worker-4', Expected thread: 'main'` with the Media3 `player-accessed-on-wrong-thread` doc link and a `Suppressed: ... Dispatchers.IO`.
@@ -10,7 +30,26 @@
**Tests.** Added `VoicePlayerTest` coverage: getter reflects the analytics-listener-cached id, never re-invokes `exoPlayer.audioSessionId` (the off-main call), and defaults to 0 before allocation. Captured the `AnalyticsListener` in the MockK harness. Verified locally: `:app:lintGooglePlayDebug` + `:app:testGooglePlayDebugUnitTest --tests VoicePlayerTest` both green (BUILD SUCCESSFUL).
**Next.** Cherry-picked onto `android-v0.8.1` hotfix branch (off the `android-v0.8.0` tag) for a focused patch release; merges back to `dev` after.
**Next.** Landed on `dev` via PR #60, then shipped as the focused patch release `android-v0.8.1` (cherry-picked off the `android-v0.8.0` tag, PR #62); `main` merged back to `dev`.
---
## 2026-05-24 — Background Hermes runs in Realtime Agent voice (ADR 33)
**Context.** Realtime Agent ran each Hermes turn synchronously *inside* the provider event pump (`_run_brokered_tool` did `return await task`), so a long research/multi-tool/desktop run froze the whole realtime session until it finished. ADR 33 + `docs/plans/2026-05-24-realtime-background-hermes-runs.md` define a three-tier model (foreground / promoted / durable) with the relay as an explicit audio-floor owner. Branch `feature/realtime-background-hermes-runs`.
**What shipped (phased, per the plan):**
- **Phase 0 — idle-tolerance probe + verdict.** `scripts/realtime-provider-idle-probe.py` + an "Idle tolerance" section in `docs/realtime-voice-poc.md`. **Ran live against OpenAI** (`VOICE_TOOLS_OPENAI_KEY` in `~/.hermes/.env`): the session survived 10s/20s/30s quiescent windows and returned clean audio on every post-idle turn → verdict **`hold-floor-ok`**. xAI has no creds on the dev box, so its verdict is recorded analytically as `hold-floor-ok` (same `turn_detection:None` multi-turn model; the implementation closes the pending call rather than holding an open response) — confirm on the relay host. Incidental finding logged: OpenAI now wants `session.audio.output.format.rate` at `session.update` (minor `_session_update` follow-up; session still worked).
- **Phase 1 — floor owner.** New `plugin/relay/realtime_agent/floor.py`: pure, single-owner audio floor (`provider` / `relay_tts` / `android_filler` mouths; `idle/provider_speaking/hermes_filler/result_pending` labels). Wired behavior-preservingly into the broker (acquire/release on AUDIO_DELTA/AUDIO_DONE/RESPONSE_DONE; relay-TTS render holds the floor; filler gated by `can_speak`). Invariants in `test_realtime_floor.py`.
- **Phase 2 — Tier B promotion (was default off).** `_run_brokered_tool` shields the run and waits `promote_after_ms`; if still running it detaches to the background, closes the pending provider call with an interim ack, optionally speaks a handoff, and `_deliver_background_result` speaks the answer once the floor is idle. New events `hermes.run.promoted` / `hermes.run.background_completed` + `tier`/`floor` on progress; 8 new settings. `test_realtime_promotion.py` (promote+pump-responsive, short=no-promote, cancel, detach-resume-replays).
- **Phase 3 — default-on + Tier C + Android + docs.** Flipped `promotion_enabled` default **true** (safe: the path closes the pending call rather than holding an open response, so the socket only sees the normal between-turns idle gap). `hermes_run_task(mode="background")` detaches immediately (`tier:"durable"`). Settings exposed on `GET/PATCH /voice/realtime-agent/config`. Android: parse new events → "working on it" chip; Voice Settings → Realtime Agent → Background tasks (promote toggle, spoken-handoff toggle, result-delivery segmented control) → `RelayVoiceClient.updateRealtimeAgentPromotion()`.
**Why default-on despite the Phase 0 gate.** The implementation closes the pending function call with an interim background ack instead of parking an open provider response, so the worst-case "idle open-response" the gate worried about doesn't occur — the socket sits in the same idle state it does between any two user turns. The probe is retained to confirm per-provider survival; documented in config + ADR.
**Verified.** Python realtime suite **58 tests green** (`test_realtime_floor`, `test_realtime_promotion`, both provider suites, routes, profile-voice-config). `./gradlew lint` — Kotlin compiles clean; the only 2 lint errors are in the gitignored `local.properties` (absent in CI). Pre-existing unrelated `test_reads_hermes_xai_oauth_credential_pool` failure confirmed on `origin/dev` baseline.
**Next.** Confirm the xAI idle verdict on the relay host (where xAI creds live); fix the OpenAI `_session_update` rate field; run the lab smoke on a paired device. Open the PR to `dev`.
---
+2 -2
View File
@@ -101,7 +101,7 @@ Small follow-ons to v0.4 deliberately deferred to keep the v0.4.0 release surfac
**What the middleware can do (near-term, ships via install.sh).** New aiohttp middleware in `hermes_relay_bootstrap/_command_middleware.py`, installed at the same `_PatchedApplication.__setitem__` hook as the current route injection so it lands before `AppRunner.setup()` freezes the app. Filters by `request.path in ("/v1/runs", "/v1/chat/completions")` — zero-cost fast path for everything else. On chat paths: parses the body, lazy-imports `GATEWAY_KNOWN_COMMANDS` + `resolve_command()` + `gateway_help_lines()` from `hermes_cli.commands`, and splits on command type:
- **Stateless commands** (`/help`, `/commands`, and any others the upstream Option B PR ends up supporting without router state) — actually dispatch, emit a synthetic SSE stream matching the runs handler's existing event shape so the Android client at `HermesApiClient.kt:655-715` renders it as a normal assistant turn.
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, etc. — most of the registry) — emit a synthetic SSE stream whose content is a short, helpful notice: *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` (post-PR-#8556) or a channel with session state. For commands that work here, type `/help`."* This replaces the LLM hallucination with a deterministic, accurate message that points the user at the real fix.
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, etc. — most of the registry) — emit a synthetic SSE stream whose content is a short, helpful notice: *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` or a channel with session state. For commands that work here, type `/help`."* This replaces the LLM hallucination with a deterministic, accurate message that points the user at the real fix.
**On no match** (unknown command, cli-only command, or plain text): falls through to `handler(request)` unchanged. Fork-detects the same way the existing injection does — if the upstream preprocessor PR lands first, the middleware no-ops.
@@ -109,7 +109,7 @@ Small follow-ons to v0.4 deliberately deferred to keep the v0.4.0 release surfac
**Files.** New `hermes_relay_bootstrap/_command_middleware.py` (~150 LOC), one-line append in `_patch.py` inside `_maybe_register_routes`, stdlib `unittest` coverage in `plugin/tests/test_bootstrap_command_middleware.py` mirroring the existing `test_bootstrap_patch.py` harness. Mirrors the upstream Option B PR exactly so the two can be reviewed side-by-side.
**Phase 2 — stateful dispatch on the session chat stream endpoint (post PR #8556).** Once PR #8556 merges and `/api/sessions/{id}/chat/stream` ships natively in upstream, a separate middleware (or a follow-up upstream PR) can add a preprocessor **scoped to that endpoint only**, leveraging the `session_id` in the URL as the persistence handle. At that point stateful commands become a dict write against session-scoped state — `session.model_override = new_model` — without needing to refactor `GatewayRouter` or plumb api_server into the router. Much smaller than a full router refactor, and it matches upstream's partition: `/v1/*` stays stateless, statefulness lives on `/api/sessions/*`. Blocked on #8556 landing.
**Phase 2 — stateful dispatch on the session chat stream endpoint (unblocked by PR #33134 / commit `f7527b0`).** Since `/api/sessions/{id}/chat/stream` now ships natively in upstream, a separate middleware (or a follow-up upstream PR) can add a preprocessor **scoped to that endpoint only**, leveraging the `session_id` in the URL as the persistence handle. At that point stateful commands become a dict write against session-scoped state — `session.model_override = new_model` — without needing to refactor `GatewayRouter` or plumb api_server into the router. Much smaller than a full router refactor, and it matches upstream's partition: `/v1/*` stays stateless, statefulness lives on `/api/sessions/*`.
## Future — v0.5+
+3 -3
View File
@@ -78,10 +78,10 @@ Things to look into:
- **Tool registration discoverability** — `android_*` tools register at gateway import time. There's no canonical "list installed plugin tools" API. Would adding one to upstream make sense, or is `gateway tool list` already enough?
- **Versioning + compatibility ranges** — `pip install -e` doesn't enforce version pins between hermes-agent and our plugin. A breaking change in upstream's plugin loader could silently break us. Do we need a `hermes_compat: ">=0.8.0,<1.0.0"` field somewhere?
- **`hermes-relay-self-setup` SKILL.md as a precedent** — we just shipped a self-installing skill that an LLM can fetch from a raw GitHub URL and execute. Does this pattern generalize? Could it become a recommended way for any third-party Hermes project to ship setup automation?
- **Bootstrap injection** — `hermes_relay_bootstrap/` monkey-patches `aiohttp.web.Application` to inject endpoints into vanilla upstream. This is intentional but feels like a hack. Upstream PR #8556 (`feat/session-api`) will eventually let us delete it — verified 2026-04-15 that its scope covers the full bootstrap surface (sessions, memory, skills, config, available-models). Track that PR's status periodically.
- **Gateway slash-command preprocessor — upstream Stage 1 PR.** Sibling follow-up to #8556. Intercepts known gateway commands on `/v1/runs` + `/v1/chat/completions`, dispatches the stateless ones (`/help`, `/commands`) via `gateway_help_lines()`, returns a deterministic "use a channel with session state" notice for the stateful majority. Currently being prepared in `C:/Users/Bailey/Desktop/Open-Projects/hermes-agent-pr-prep/` on branch `feat/api-server-gateway-commands`; awaiting subagent's code + draft PR body before pushing. See `docs/upstream-contributions.md` §5.
- **Bootstrap injection shrink path** — `hermes_relay_bootstrap/` monkey-patches `aiohttp.web.Application` to inject endpoints into vanilla/partial upstream. Upstream commit `f7527b0` via PR #33134 now covers baseline sessions/chat/fork/message history, and `/v1/skills` covers list metadata. Do **not** delete the bootstrap wholesale yet: Relay still depends on compatibility routes that upstream lacks or does not match (`/api/sessions/search`, `/api/memory`, `/api/config`, legacy `/api/skills` detail routes, `/api/available-models`, and voice aliases). Shrink per route group only after native parity or client migration.
- **Gateway slash-command preprocessor — upstream Stage 1 PR.** Follow-up to the native session-control baseline from PR #33134 / commit `f7527b0`. Intercepts known gateway commands on `/v1/runs` + `/v1/chat/completions`, dispatches the stateless ones (`/help`, `/commands`) via `gateway_help_lines()`, returns a deterministic "use a channel with session state" notice for the stateful majority. Currently being prepared in `C:/Users/Bailey/Desktop/Open-Projects/hermes-agent-pr-prep/` on branch `feat/api-server-gateway-commands`; awaiting subagent's code + draft PR body before pushing. See `docs/upstream-contributions.md` §5.
- **Gateway slash-command preprocessor — bootstrap middleware (Stage 1 equivalent).** Sibling shim in `hermes_relay_bootstrap/_command_middleware.py` that mirrors the upstream Stage 1 PR as an aiohttp middleware injected at bootstrap time. Ships the hallucination fix to vanilla-upstream installs before the upstream PR lands. Planned for v0.4.1, after the current bridge feature branch wraps. See `ROADMAP.md` v0.4.1 entry.
- **Stage 2 — stateful slash-command dispatch on `/api/sessions/{id}/chat/stream`.** Blocked on PR #8556 merging. Once session primitives ship upstream, add a preprocessor scoped to the session chat stream endpoint only, using `session_id` as the persistence handle. Separate upstream PR + matching bootstrap middleware. See `docs/upstream-contributions.md` §5 ("Stage 2").
- **Stage 2 — stateful slash-command dispatch on `/api/sessions/{id}/chat/stream`.** Unblocked by upstream PR #33134 / commit `f7527b0`. Add a preprocessor scoped to the session chat stream endpoint only, using `session_id` as the persistence handle. Separate upstream PR + matching bootstrap middleware. See `docs/upstream-contributions.md` §5 ("Stage 2").
When the answer becomes clearer, this section becomes either an ADR in `docs/decisions.md` or a Plan under `Plans/`.
@@ -29,6 +29,13 @@ data class VoiceSettings(
val autoTts: Boolean = false,
val language: String = "",
val realtimeTraceDetails: Boolean = false,
/**
* When true (default), Realtime Agent keeps one provider session/socket open
* across turns (persistent conversation). When false, falls back to the
* legacy one-session-per-utterance path. See
* docs/plans/2026-05-24-realtime-persistent-session.md.
*/
val realtimePersistentSession: Boolean = true,
)
enum class VoiceEngineMode(val storageValue: String) {
@@ -52,6 +59,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
private val KEY_AUTO_TTS = booleanPreferencesKey("voice_auto_tts")
private val KEY_LANGUAGE = stringPreferencesKey("voice_language")
private val KEY_REALTIME_TRACE_DETAILS = booleanPreferencesKey("voice_realtime_trace_details")
private val KEY_REALTIME_PERSISTENT_SESSION =
booleanPreferencesKey("voice_realtime_persistent_session")
const val DEFAULT_ENGINE_MODE = "hermes_voice_output"
const val DEFAULT_INTERACTION_MODE = "tap"
@@ -59,6 +68,7 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
const val DEFAULT_AUTO_TTS = false
const val DEFAULT_LANGUAGE = ""
const val DEFAULT_REALTIME_TRACE_DETAILS = false
const val DEFAULT_REALTIME_PERSISTENT_SESSION = true
}
val settings: Flow<VoiceSettings> = dataStore.data
@@ -73,6 +83,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
language = prefs[KEY_LANGUAGE] ?: DEFAULT_LANGUAGE,
realtimeTraceDetails = prefs[KEY_REALTIME_TRACE_DETAILS]
?: DEFAULT_REALTIME_TRACE_DETAILS,
realtimePersistentSession = prefs[KEY_REALTIME_PERSISTENT_SESSION]
?: DEFAULT_REALTIME_PERSISTENT_SESSION,
)
}
.distinctUntilChanged()
@@ -100,4 +112,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
suspend fun setRealtimeTraceDetails(enabled: Boolean) {
dataStore.edit { it[KEY_REALTIME_TRACE_DETAILS] = enabled }
}
suspend fun setRealtimePersistentSession(enabled: Boolean) {
dataStore.edit { it[KEY_REALTIME_PERSISTENT_SESSION] = enabled }
}
}
@@ -99,6 +99,24 @@ data class ServerCapabilities(
}
}
internal val HERMES_SKILL_ENDPOINTS = listOf("/v1/skills", "/api/skills")
internal fun parseSkillListBody(json: Json, body: String): List<SkillInfo>? {
try {
val parsed = json.decodeFromString<SkillListResponse>(body)
val skills = parsed.skills ?: parsed.items ?: parsed.data
if (skills != null) return skills
} catch (_: Exception) {
// Fall through to direct-array compatibility below.
}
try {
return json.decodeFromString<List<SkillInfo>>(body)
} catch (_: Exception) {
return null
}
}
/**
* Direct HTTP/SSE client for the Hermes API Server.
*
@@ -342,27 +360,21 @@ class HermesApiClient(
// --- Skills ---
suspend fun getSkills(): List<SkillInfo> = withContext(Dispatchers.IO) {
try {
val request = authRequest("$baseUrl/api/skills").get().build()
client.newCall(request).execute().use { response ->
if (!response.isSuccessful) return@withContext emptyList()
val body = response.body?.string() ?: return@withContext emptyList()
// Try structured response: { "skills": [...] } or { "items": [...] }
try {
val parsed = json.decodeFromString<SkillListResponse>(body)
val skills = parsed.skills ?: parsed.items
for (endpoint in HERMES_SKILL_ENDPOINTS) {
try {
val request = authRequest("$baseUrl$endpoint").get().build()
client.newCall(request).execute().use { response ->
if (!response.isSuccessful) return@use
val body = response.body?.string() ?: return@use
val skills = parseSkillListBody(json, body)
if (skills != null) return@withContext skills
} catch (_: Exception) { /* fall through */ }
// Try direct array: [...]
try {
return@withContext json.decodeFromString<List<SkillInfo>>(body)
} catch (_: Exception) { /* fall through */ }
emptyList()
}
} catch (e: Exception) {
Log.w(TAG, "Failed to fetch skills from $endpoint: ${e.message}")
}
} catch (e: Exception) {
Log.w(TAG, "Failed to fetch skills: ${e.message}")
emptyList()
}
emptyList()
}
// --- Server personalities ---
@@ -532,6 +532,64 @@ class RelayVoiceClient(
label = "Realtime agent config update",
)
/**
* PATCH the ADR 33 background-run promotion settings. Only non-null fields
* are sent; the relay echoes back the full [RealtimeVoiceConfig].
*/
suspend fun updateRealtimeAgentPromotion(
promotionEnabled: Boolean? = null,
promoteAfterMs: Int? = null,
spokenHandoff: Boolean? = null,
resultDelivery: String? = null,
backgroundDefaultMode: String? = null,
progressSpokenAfterMs: Int? = null,
progressRepeatMs: Int? = null,
maxBackgroundRuns: Int? = null,
): Result<RealtimeVoiceConfig> = withContext(Dispatchers.IO) {
val httpBase = resolveHttpBase()
?: return@withContext Result.failure(IllegalStateException("Relay URL not configured"))
val token = resolveBearerToken()
if (token.isNullOrBlank()) {
return@withContext Result.failure(missingAuthError())
}
val payload = buildJsonObject {
promotionEnabled?.let { put("promotion_enabled", JsonPrimitive(it)) }
promoteAfterMs?.let { put("promote_after_ms", JsonPrimitive(it)) }
spokenHandoff?.let { put("spoken_handoff", JsonPrimitive(it)) }
resultDelivery?.takeIf { it.isNotBlank() }?.let {
put("result_delivery", JsonPrimitive(it.trim()))
}
backgroundDefaultMode?.takeIf { it.isNotBlank() }?.let {
put("background_default_mode", JsonPrimitive(it.trim()))
}
progressSpokenAfterMs?.let { put("progress_spoken_after_ms", JsonPrimitive(it)) }
progressRepeatMs?.let { put("progress_repeat_ms", JsonPrimitive(it)) }
maxBackgroundRuns?.let { put("max_background_runs", JsonPrimitive(it)) }
}
val request = Request.Builder()
.url(urlWithProfile("$httpBase/voice/realtime-agent/config"))
.patch(payload.toString().toRequestBody(JSON_MEDIA_TYPE))
.header("Authorization", "Bearer $token")
.header("Accept", "application/json")
.build()
try {
okHttpClient.newCall(request).execute().use { response ->
if (!response.isSuccessful) {
val body = response.body?.string().orEmpty()
return@withContext Result.failure(
IOException(describeHttpError(response.code, response.message, body))
)
}
val raw = response.body?.string().orEmpty()
Result.success(json.decodeFromString(RealtimeVoiceConfig.serializer(), raw))
}
} catch (e: IOException) {
Result.failure(IOException("Realtime promotion update failed: ${e.message ?: "network error"}"))
} catch (e: Exception) {
Result.failure(e)
}
}
suspend fun getVoiceOutputConfig(): Result<VoiceOutputConfig> = withContext(Dispatchers.IO) {
val httpBase = resolveHttpBase()
?: return@withContext Result.failure(IllegalStateException("Relay URL not configured"))
@@ -1133,6 +1191,22 @@ class RelayVoiceClient(
}
}
/**
* Run a Realtime Agent voice session.
*
* **One-shot (default):** with [turnInputs] null, this opens a session, sends
* the single [inputPcm]/[prompt] turn, and completes + closes the socket on
* `voice.response.done` — the historical per-utterance behavior used by the
* Voice Lab and the one-shot fallback.
*
* **Persistent (ADR follow-up, see docs/plans/2026-05-24-realtime-persistent-session.md):**
* with [turnInputs] non-null the socket stays open across turns. The first
* turn is [inputPcm]/[prompt]; each subsequent [RealtimeTurnInput] read from
* the channel is sent on the *same* socket, so the provider keeps a live
* conversation. `voice.response.done` invokes [onTurnComplete] instead of
* closing; the call returns only when the channel closes (voice-mode exit) or
* a fatal socket/provider error occurs.
*/
suspend fun runRealtimeAgent(
prompt: String,
inputPcm: ByteArray,
@@ -1140,8 +1214,11 @@ class RelayVoiceClient(
chatSessionId: String? = null,
conversationContext: List<RealtimeConversationContextMessage> = emptyList(),
onHandoff: (VoiceHandoffEvent) -> Unit = {},
turnInputs: kotlinx.coroutines.channels.ReceiveChannel<RealtimeTurnInput>? = null,
onTurnComplete: (RealtimeVoiceSummary) -> Unit = {},
onEvent: (RealtimeVoiceEvent, RealtimeAgentSessionControl) -> Unit,
): Result<RealtimeVoiceSummary> = withContext(Dispatchers.IO) {
val persistent = turnInputs != null
val httpBase = resolveHttpBase()
?: return@withContext Result.failure(IllegalStateException("Relay URL not configured"))
val wsBase = resolveWebSocketBase()
@@ -1172,8 +1249,11 @@ class RelayVoiceClient(
val lastAudioEventId = AtomicLong(0L)
val lastPlayedAudioEventId = AtomicLong(0L)
val lastInputChunkId = AtomicLong(0L)
val turnStartedAtMs = System.currentTimeMillis()
val lastEventAtMs = AtomicLong(turnStartedAtMs)
val turnStartedAtMs = AtomicLong(System.currentTimeMillis())
val lastEventAtMs = AtomicLong(turnStartedAtMs.get())
// True while a turn is awaiting its response. In persistent mode the idle
// guard only applies while a turn is active; between-turn idle is normal.
val activeTurn = AtomicBoolean(true)
val inputChunks = buildList {
var offset = 0
var chunkId = 1L
@@ -1184,8 +1264,33 @@ class RelayVoiceClient(
offset = end
}
}
// Highest input chunk id sent on this session. The relay dedups by
// input_chunk_seq, so subsequent turns must continue past this.
val sessionMaxChunkId = AtomicLong(inputChunks.size.toLong())
var audioChunks = 0
var audioBytes = 0
// Persistent-mode: chunk + send one more utterance on the open socket.
fun sendTurnPcm(webSocket: WebSocket, pcm: ByteArray, sampleRate: Int) {
var offset = 0
var sentAny = false
while (offset < pcm.size) {
val end = (offset + REALTIME_INPUT_CHUNK_BYTES).coerceAtMost(pcm.size)
val chunkId = sessionMaxChunkId.incrementAndGet()
val encoded = Base64.getEncoder().encodeToString(pcm.copyOfRange(offset, end))
webSocket.send(
"""{"type":"input_audio.append","chunk_id":$chunkId,"sample_rate":$sampleRate,"audio_base64":"$encoded"}"""
)
sentAny = true
offset = end
}
if (sentAny) {
webSocket.send("""{"type":"input_audio.commit"}""")
}
turnStartedAtMs.set(System.currentTimeMillis())
lastEventAtMs.set(System.currentTimeMillis())
activeTurn.set(true)
}
fun activateSocket(webSocket: WebSocket, generation: Long): Boolean {
while (true) {
val activeGeneration = activeSocketGeneration.get()
@@ -1332,24 +1437,28 @@ class RelayVoiceClient(
inputChunkId = event.inputChunkId,
)
if (event.type == "voice.response.done") {
if (completed.compareAndSet(false, true)) {
finished.complete(
Result.success(
RealtimeVoiceSummary(
provider = event.provider ?: session.provider,
model = event.model ?: session.model,
voice = event.voice ?: session.voice,
sampleRate = session.sampleRate,
audioChunks = audioChunks,
audioBytes = audioBytes,
firstAudioMs = event.firstAudioMs,
responseDoneMs = event.responseDoneMs,
eventLogPath = event.eventLogPath ?: session.eventLogPath,
)
)
)
val summary = RealtimeVoiceSummary(
provider = event.provider ?: session.provider,
model = event.model ?: session.model,
voice = event.voice ?: session.voice,
sampleRate = session.sampleRate,
audioChunks = audioChunks,
audioBytes = audioBytes,
firstAudioMs = event.firstAudioMs,
responseDoneMs = event.responseDoneMs,
eventLogPath = event.eventLogPath ?: session.eventLogPath,
)
if (persistent) {
// Turn boundary, not session boundary: keep the socket
// open for the next utterance.
activeTurn.set(false)
onTurnComplete(summary)
} else {
if (completed.compareAndSet(false, true)) {
finished.complete(Result.success(summary))
}
webSocket.close(1000, "done")
}
webSocket.close(1000, "done")
} else if (event.type == "voice.error" || event.type == "voice.session.resume_failed") {
completeFailure(event.message ?: "Realtime agent error")
webSocket.close(1011, "provider error")
@@ -1452,25 +1561,85 @@ class RelayVoiceClient(
suspend fun awaitRealtimeAgentCompletion(): Result<RealtimeVoiceSummary> {
while (true) {
val now = System.currentTimeMillis()
val turnElapsedMs = now - turnStartedAtMs
val idleElapsedMs = now - lastEventAtMs.get()
if (turnElapsedMs >= REALTIME_AGENT_MAX_TURN_MS) {
throw IOException("Realtime agent exceeded the turn limit")
// In persistent mode the turn/idle guards only apply while a turn
// is actually in flight; between-turn idle is expected and must
// not trip the stall timeout. The session ends when the turn
// channel closes or a fatal error completes `finished`.
val guardActive = !persistent || activeTurn.get()
if (guardActive) {
val turnElapsedMs = now - turnStartedAtMs.get()
val idleElapsedMs = now - lastEventAtMs.get()
if (turnElapsedMs >= REALTIME_AGENT_MAX_TURN_MS) {
throw IOException("Realtime agent exceeded the turn limit")
}
if (idleElapsedMs >= REALTIME_AGENT_IDLE_TIMEOUT_MS) {
throw IOException("Realtime agent stalled waiting for relay events")
}
val waitMs = minOf(
REALTIME_AGENT_WAIT_SLICE_MS,
REALTIME_AGENT_MAX_TURN_MS - turnElapsedMs,
REALTIME_AGENT_IDLE_TIMEOUT_MS - idleElapsedMs,
).coerceAtLeast(1L)
withTimeoutOrNull(waitMs) {
finished.await()
}?.let { return it }
} else {
withTimeoutOrNull(REALTIME_AGENT_WAIT_SLICE_MS) {
finished.await()
}?.let { return it }
}
if (idleElapsedMs >= REALTIME_AGENT_IDLE_TIMEOUT_MS) {
throw IOException("Realtime agent stalled waiting for relay events")
}
val waitMs = minOf(
REALTIME_AGENT_WAIT_SLICE_MS,
REALTIME_AGENT_MAX_TURN_MS - turnElapsedMs,
REALTIME_AGENT_IDLE_TIMEOUT_MS - idleElapsedMs,
).coerceAtLeast(1L)
withTimeoutOrNull(waitMs) {
finished.await()
}?.let { return it }
}
}
// Persistent mode: drain further utterances and feed each onto the open
// socket as a new turn. Closing the channel (voice-mode exit) ends the
// session by completing `finished`.
val turnReader: Job? = if (turnInputs != null) {
launch {
try {
for (turn in turnInputs) {
val ws = currentSocket.get() ?: continue
if (turn.prompt.isNotBlank() && turn.inputPcm.isEmpty()) {
ws.send(
buildRealtimeResponseCreate(
text = turn.prompt,
toolScaffold = false,
renderMode = "verbatim",
)
)
turnStartedAtMs.set(System.currentTimeMillis())
lastEventAtMs.set(System.currentTimeMillis())
activeTurn.set(true)
} else {
sendTurnPcm(ws, turn.inputPcm, turn.sampleRate)
}
}
} finally {
// Channel closed -> end the persistent session cleanly.
if (completed.compareAndSet(false, true)) {
finished.complete(
Result.success(
RealtimeVoiceSummary(
provider = session.provider,
model = session.model,
voice = session.voice,
sampleRate = session.sampleRate,
audioChunks = audioChunks,
audioBytes = audioBytes,
firstAudioMs = null,
responseDoneMs = null,
eventLogPath = session.eventLogPath,
)
)
)
}
currentSocket.get()?.close(1000, "session ended")
}
}
} else {
null
}
val socket = openSocket(resume = false)
val routeWatcher = startRouteResumeWatcher(
surface = "Realtime agent",
@@ -1491,6 +1660,7 @@ class RelayVoiceClient(
Result.failure(IOException(e.message ?: "Realtime agent timed out", e))
} finally {
routeWatcher?.cancel()
turnReader?.cancel()
}
}
@@ -2080,6 +2250,8 @@ class RelayVoiceClient(
eventLogPath = (obj["event_log_path"] as? JsonPrimitive)?.contentOrNull,
firstAudioMs = (metrics?.get("first_audio_ms") as? JsonPrimitive)?.doubleOrNull,
responseDoneMs = (metrics?.get("response_done_ms") as? JsonPrimitive)?.doubleOrNull,
tier = (obj["tier"] as? JsonPrimitive)?.contentOrNull,
floor = (obj["floor"] as? JsonPrimitive)?.contentOrNull,
raw = raw,
)
} catch (e: Exception) {
@@ -2154,6 +2326,28 @@ data class RealtimeVoiceConfig(
val configScope: String? = null,
@SerialName("fallback_to_global")
val fallbackToGlobal: Boolean = false,
/** ADR 33 background-run promotion settings. */
val promotion: RealtimeVoicePromotion? = null,
)
/** Wire shape of the `promotion` block on `GET /voice/realtime-agent/config`. */
@Serializable
data class RealtimeVoicePromotion(
val enabled: Boolean = true,
@SerialName("promote_after_ms")
val promoteAfterMs: Int = 6000,
@SerialName("background_default_mode")
val backgroundDefaultMode: String = "promote",
@SerialName("spoken_handoff")
val spokenHandoff: Boolean = true,
@SerialName("progress_spoken_after_ms")
val progressSpokenAfterMs: Int = 15000,
@SerialName("progress_repeat_ms")
val progressRepeatMs: Int = 30000,
@SerialName("result_delivery")
val resultDelivery: String = "speak_when_idle",
@SerialName("max_background_runs")
val maxBackgroundRuns: Int = 1,
)
@Serializable
@@ -2392,6 +2586,9 @@ data class RealtimeVoiceEvent(
val eventLogPath: String? = null,
val firstAudioMs: Double? = null,
val responseDoneMs: Double? = null,
// ADR 33: background-run promotion fields.
val tier: String? = null,
val floor: String? = null,
val raw: String,
) {
val isAudioDelta: Boolean
@@ -2467,6 +2664,33 @@ data class RealtimeVoiceSummary(
val eventLogPath: String?,
)
/**
* One utterance fed into a persistent Realtime Agent session
* (see [RelayVoiceClient.runRealtimeAgent] persistent mode). A blank [inputPcm]
* with a non-blank [prompt] sends a text turn; otherwise the PCM is chunked and
* committed as a spoken turn.
*/
data class RealtimeTurnInput(
val inputPcm: ByteArray,
val sampleRate: Int = 16_000,
val prompt: String = "",
) {
override fun equals(other: Any?): Boolean {
if (this === other) return true
if (other !is RealtimeTurnInput) return false
return sampleRate == other.sampleRate &&
prompt == other.prompt &&
inputPcm.contentEquals(other.inputPcm)
}
override fun hashCode(): Int {
var result = inputPcm.contentHashCode()
result = 31 * result + sampleRate
result = 31 * result + prompt.hashCode()
return result
}
}
data class VoiceOutputSummary(
val provider: String,
val model: String,
@@ -300,5 +300,6 @@ data class SkillInfo(
@Serializable
data class SkillListResponse(
val skills: List<SkillInfo>? = null,
val items: List<SkillInfo>? = null
val items: List<SkillInfo>? = null,
val data: List<SkillInfo>? = null
)
@@ -301,6 +301,28 @@ fun VoiceModeOverlay(
)
}
AnimatedVisibility(
visible = uiState.backgroundRun != null,
enter = fadeIn(tween(140)),
exit = fadeOut(tween(180)),
) {
Surface(
shape = RoundedCornerShape(16.dp),
color = MaterialTheme.colorScheme.secondaryContainer,
modifier = Modifier
.fillMaxWidth()
.padding(horizontal = 24.dp, vertical = 4.dp),
) {
Text(
text = uiState.backgroundRun?.message
?: "Working on it in the background…",
style = MaterialTheme.typography.labelMedium,
color = MaterialTheme.colorScheme.onSecondaryContainer,
modifier = Modifier.padding(horizontal = 16.dp, vertical = 10.dp),
)
}
}
Spacer(Modifier.height(8.dp))
DestructiveCountdownRow(
@@ -941,6 +941,26 @@ fun VoiceSettingsScreen(
},
)
}
Row(
modifier = Modifier.fillMaxWidth(),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.SpaceBetween,
) {
Column(modifier = Modifier.weight(1f)) {
Text("Persistent session", style = MaterialTheme.typography.bodyLarge)
Text(
text = "Keep one provider conversation open across turns. Turn off to fall back to a fresh session per utterance.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
Switch(
checked = voiceSettings.realtimePersistentSession,
onCheckedChange = { enabled ->
scope.launch { prefsRepo.setRealtimePersistentSession(enabled) }
},
)
}
Spacer(Modifier.height(4.dp))
ProviderRow(
label = "Status",
@@ -975,6 +995,109 @@ fun VoiceSettingsScreen(
ProviderRow(label = "Auth", value = realtimeAuthLabel(config))
}
realtimeConfig?.promotion?.let { promo ->
HorizontalDivider(modifier = Modifier.padding(vertical = 8.dp))
Text("Background tasks", style = MaterialTheme.typography.titleSmall)
Text(
text = "Long Hermes tasks keep running in the background so the conversation stays responsive; the answer is spoken when it's ready.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(8.dp))
Row(
modifier = Modifier.fillMaxWidth(),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.SpaceBetween,
) {
Column(modifier = Modifier.weight(1f)) {
Text("Promote long tasks", style = MaterialTheme.typography.bodyLarge)
Text(
text = "Detach a slow run after ${promo.promoteAfterMs} ms instead of waiting silently",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
Switch(
checked = promo.enabled,
onCheckedChange = { enabled ->
scope.launch {
val client = voiceClient ?: return@launch
val result = client.updateRealtimeAgentPromotion(
promotionEnabled = enabled,
)
if (result.isSuccess) realtimeConfig = result.getOrNull()
}
},
)
}
if (promo.enabled) {
Spacer(Modifier.height(4.dp))
Row(
modifier = Modifier.fillMaxWidth(),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.SpaceBetween,
) {
Column(modifier = Modifier.weight(1f)) {
Text("Spoken handoff", style = MaterialTheme.typography.bodyLarge)
Text(
text = "Say a short \"I'm on it\" when a task moves to the background",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
Switch(
checked = promo.spokenHandoff,
onCheckedChange = { enabled ->
scope.launch {
val client = voiceClient ?: return@launch
val result = client.updateRealtimeAgentPromotion(
spokenHandoff = enabled,
)
if (result.isSuccess) realtimeConfig = result.getOrNull()
}
},
)
}
Spacer(Modifier.height(8.dp))
Text(
"When the answer is ready",
style = MaterialTheme.typography.labelMedium,
)
Spacer(Modifier.height(4.dp))
val deliveryOptions = listOf(
"speak_when_idle",
"notify_then_speak",
"visual_only",
)
val deliveryLabels = listOf("Speak", "Notify", "Show only")
SingleChoiceSegmentedButtonRow(modifier = Modifier.fillMaxWidth()) {
deliveryOptions.forEachIndexed { index, option ->
SegmentedButton(
shape = SegmentedButtonDefaults.itemShape(
index = index,
count = deliveryOptions.size,
),
onClick = {
scope.launch {
val client = voiceClient ?: return@launch
val result = client.updateRealtimeAgentPromotion(
resultDelivery = option,
)
if (result.isSuccess) realtimeConfig = result.getOrNull()
}
},
selected = option == promo.resultDelivery,
) {
Text(
deliveryLabels[index],
style = MaterialTheme.typography.labelSmall,
)
}
}
}
}
}
HorizontalDivider(modifier = Modifier.padding(vertical = 8.dp))
Row(
@@ -17,6 +17,7 @@ import com.hermesandroid.relay.data.BargeInPreferencesRepository
import com.hermesandroid.relay.data.BargeInSensitivity
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.data.RealtimeConversationContextMessage
import com.hermesandroid.relay.data.ToolCall
import com.hermesandroid.relay.data.VoiceEngineMode
import com.hermesandroid.relay.data.VoiceIntentTrace
@@ -26,6 +27,8 @@ import com.hermesandroid.relay.diagnostics.DiagnosticsLog
import com.hermesandroid.relay.network.ChannelMultiplexer
import com.hermesandroid.relay.network.RelayVoiceClient
import com.hermesandroid.relay.network.RealtimeAgentSessionControl
import com.hermesandroid.relay.network.RealtimeTurnInput
import com.hermesandroid.relay.network.RealtimeVoiceSummary
import com.hermesandroid.relay.network.RealtimeVoiceEvent
import com.hermesandroid.relay.network.VoiceHandoffEvent
import com.hermesandroid.relay.network.handlers.LocalDispatchResult
@@ -131,6 +134,21 @@ data class VoiceUiState(
* fresh turn (mic tap), whichever comes first.
*/
val permissionDeniedCallout: PermissionDeniedCallout? = null,
/**
* ADR 33: non-null while a Hermes run has been promoted to (or started as) a
* background task. Drives a persistent "working on it" chip so the user
* knows a long task is still running after the spoken handoff. Cleared when
* the background run completes, is cancelled, or errors.
*/
val backgroundRun: BackgroundRunState? = null,
)
/** ADR 33 background/promoted Hermes run surface for the voice overlay. */
data class BackgroundRunState(
val runId: String? = null,
/** "promoted" (auto-detached long run) or "durable" (explicit mode=background). */
val tier: String = "promoted",
val message: String = "Working on it in the background…",
)
data class VoiceHandoffStatus(
@@ -386,8 +404,28 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
private var voicePreferencesJob: Job? = null
private var voiceEngineMode: VoiceEngineMode = VoiceEngineMode.HermesVoiceOutput
private var realtimeTraceDetails: Boolean = false
private var realtimePersistentSession: Boolean = true
private var realtimeAgentControl: RealtimeAgentSessionControl? = null
private var realtimeConfirmationControl: RealtimeAgentSessionControl? = null
// === Persistent realtime-agent session (one socket across turns) ===
// The long-lived call to RelayVoiceClient.runRealtimeAgent(persistent) runs
// in realtimeSessionJob; further utterances are fed on realtimeTurnChannel.
// The per-turn event holders below are hoisted to fields so the single
// session-lived event callback can serve every turn; submitRealtimeTurn /
// the open path reset them at each turn boundary.
private var realtimeSessionJob: Job? = null
private var realtimeTurnChannel: kotlinx.coroutines.channels.Channel<RealtimeTurnInput>? = null
private var rtUserText: String = ""
private var rtAssistantMessageId: String = ""
private var rtConversationContext: List<RealtimeConversationContextMessage> = emptyList()
private val audioSeen = AtomicBoolean(false)
private val audioBytes = AtomicInteger(0)
private val bargeInStarted = AtomicBoolean(false)
private val lastRealtimeAudioEventId = AtomicLong(0L)
private var responseText = StringBuilder()
private var inputTranscript = StringBuilder()
private val spokenStatusKeys = mutableSetOf<String>()
private var voiceRelayPreflight: (suspend () -> Result<Unit>)? = null
// === PHASE3-voice-intents: voice→bridge intent routing ===
@@ -811,8 +849,18 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
"realtimeTraceDetails=${settings.realtimeTraceDetails}",
)
}
// Switching engine away from Realtime Agent (or disabling the
// persistent toggle) must drop any open persistent session.
if (
(voiceEngineMode == VoiceEngineMode.RealtimeAgent &&
nextEngineMode != VoiceEngineMode.RealtimeAgent) ||
(realtimePersistentSession && !settings.realtimePersistentSession)
) {
closeRealtimeSession()
}
voiceEngineMode = nextEngineMode
realtimeTraceDetails = settings.realtimeTraceDetails
realtimePersistentSession = settings.realtimePersistentSession
val mode = when (settings.interactionMode.lowercase()) {
"hold" -> InteractionMode.HoldToTalk
"continuous" -> InteractionMode.Continuous
@@ -1068,6 +1116,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
// Chime BEFORE teardown — AudioTrack release would cut it off otherwise.
try { sfxPlayer?.playExit() } catch (_: Exception) { /* ignore */ }
cancelRealtimeAgentTurn("exit voice mode")
closeRealtimeSession()
// B4: tear down the barge-in listener + timers before we kill the
// player so AEC doesn't try to track a released audio session.
stopBargeInListener()
@@ -1740,6 +1789,31 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
responseText = "",
)
}
if (realtimePersistentSession) {
// Persistent conversation: one socket across turns. Open lazily on
// the first turn (the long-lived call runs in realtimeSessionJob),
// then feed further utterances on the channel. Fallback to one-shot
// if the toggle is off (Voice Settings).
if (realtimeSessionJob?.isActive == true && realtimeTurnChannel != null) {
submitRealtimeTurn(chatVm, inputPcm, inputSampleRate)
} else {
closeRealtimeSession()
realtimeTurnChannel = kotlinx.coroutines.channels.Channel(
kotlinx.coroutines.channels.Channel.UNLIMITED,
)
realtimeSessionJob = viewModelScope.launch {
runRealtimeAgentTurn(
client = client,
chatVm = chatVm,
userText = "",
inputPcm = inputPcm,
inputSampleRate = inputSampleRate,
persistentOpen = true,
)
}
}
return
}
runRealtimeAgentTurn(
client = client,
chatVm = chatVm,
@@ -2004,6 +2078,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
userText: String,
inputPcm: ByteArray,
inputSampleRate: Int,
persistentOpen: Boolean = false,
) {
providerRealtimeAgentTurnActive.set(true)
streamObserverJob?.cancel()
@@ -2031,19 +2106,23 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
)
}
val conversationContext = chatVm.realtimeAgentContextMessages()
val assistantMessageId = chatVm.startRealtimeAgentTurn(
// Per-turn event state is hoisted to fields so one session-lived callback
// serves every turn in persistent mode; reset them for this turn.
audioSeen.set(false)
audioBytes.set(0)
bargeInStarted.set(false)
lastRealtimeAudioEventId.set(0L)
responseText = StringBuilder()
inputTranscript = StringBuilder()
spokenStatusKeys.clear()
rtUserText = userText
rtConversationContext = chatVm.realtimeAgentContextMessages()
rtAssistantMessageId = chatVm.startRealtimeAgentTurn(
userText = userText,
chatSessionId = chatVm.currentSessionId.value,
)
val conversationContext = rtConversationContext
val pcmPlayer = realtimePcmPlayer
val audioSeen = AtomicBoolean(false)
val audioBytes = AtomicInteger(0)
val bargeInStarted = AtomicBoolean(false)
val lastRealtimeAudioEventId = AtomicLong(0L)
val responseText = StringBuilder()
val inputTranscript = StringBuilder()
val spokenStatusKeys = mutableSetOf<String>()
fun emitStatus(
key: String,
@@ -2105,10 +2184,12 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
chatSessionId = chatVm.currentSessionId.value,
conversationContext = conversationContext,
onHandoff = ::recordVoiceHandoff,
turnInputs = if (persistentOpen) realtimeTurnChannel else null,
onTurnComplete = { summary -> onRealtimeTurnComplete(summary) },
) { event, control ->
realtimeAgentControl = control
chatVm.applyRealtimeAgentEvent(
assistantMessageId = assistantMessageId,
assistantMessageId = rtAssistantMessageId,
event = event,
showDetailedTrace = realtimeTraceDetails,
)
@@ -2119,13 +2200,13 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
it.copy(
state = VoiceState.Listening,
outputAudioActive = false,
transcribedText = inputTranscript.toString().ifBlank { userText },
transcribedText = inputTranscript.toString().ifBlank { rtUserText },
)
}
}
"voice.input_transcript.final" -> {
inputTranscript.clear()
inputTranscript.append(event.text ?: userText)
inputTranscript.append(event.text ?: rtUserText)
_uiState.update {
it.copy(
state = VoiceState.Thinking,
@@ -2135,6 +2216,13 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
}
"voice.response.started", "hermes.run.started" -> {
if (event.type == "hermes.run.started") {
Log.i(
TAG,
"Realtime Hermes run started run=${event.runId ?: "?"} " +
"session=${event.chatSessionId ?: "?"}",
)
}
_uiState.update {
it.copy(state = VoiceState.Thinking, outputAudioActive = false)
}
@@ -2172,6 +2260,12 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
"hermes.run.progress" -> {
val line = realtimeHermesProgressLine(event.message)
Log.i(
TAG,
"Realtime Hermes progress tier=${event.tier ?: "?"} " +
"floor=${event.floor ?: "?"} status=${event.statusKey ?: "?"} " +
"shouldSpeak=${event.shouldSpeak} run=${event.runId ?: "?"}",
)
if (event.shouldSpeak) {
emitStatus(
key = "progress-${event.runId ?: "run"}-${event.statusKey ?: line}",
@@ -2188,6 +2282,40 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
)
}
}
"hermes.run.promoted" -> {
// The run detached to the background; the provider speaks the
// handoff. Surface a persistent chip so the user knows a long
// task is still in flight (ADR 33 Tier B/C).
val tier = event.tier ?: "promoted"
Log.i(
TAG,
"Realtime run promoted to background tier=$tier " +
"run=${event.runId ?: "?"} session=${event.chatSessionId ?: "?"}",
)
_uiState.update {
it.copy(
backgroundRun = BackgroundRunState(
runId = event.runId,
tier = tier,
message = if (tier == "durable") {
"Started a background task — I'll report back."
} else {
"This is taking a moment — working on it in the background."
},
),
)
}
}
"hermes.run.background_completed" -> {
// Background run finished; the spoken summary follows via the
// provider's forced-summary turn. Clear the chip.
Log.i(
TAG,
"Realtime background run completed run=${event.runId ?: "?"} " +
"ok=${event.success != false}",
)
_uiState.update { it.copy(backgroundRun = null) }
}
"hermes.confirmation.requested" -> {
val confirmationId = event.confirmationId
if (!confirmationId.isNullOrBlank()) {
@@ -2269,6 +2397,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
}
"hermes.run.cancelled" -> {
Log.i(TAG, "Realtime Hermes run cancelled run=${event.runId ?: "?"}")
realtimeConfirmationControl = null
_uiState.update {
it.copy(
@@ -2276,6 +2405,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
outputAudioActive = false,
responseText = "Cancelled.",
hermesConfirmation = null,
backgroundRun = null,
)
}
}
@@ -2294,19 +2424,22 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
}
"voice.error" -> {
realtimeConfirmationControl = null
val rawDetail = event.message ?: "Realtime agent failed"
DiagnosticsLog.record(
category = DiagnosticCategory.Voice,
severity = DiagnosticSeverity.Error,
title = "Realtime voice error",
detail = event.message ?: "Realtime agent failed",
detail = rawDetail,
)
_uiState.update { it.copy(hermesConfirmation = null) }
// Route the raw relay message through the classifier so
// provider-auth (and other) failures surface a clear,
// actionable message + a Voice-settings snackbar action
// instead of a raw "xAI Realtime auth ..." provider string.
surfaceError(
java.io.IOException(rawDetail),
context = "voice_config",
)
_uiState.update {
it.copy(
state = VoiceState.Error,
error = event.message ?: "Realtime agent failed",
hermesConfirmation = null,
)
}
}
}
}
@@ -2335,9 +2468,78 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
title = "Realtime voice turn failed",
detail = err?.message ?: "Unknown error",
)
// A persistent session ending in error must drop so the next turn opens
// a fresh one rather than submitting into a dead channel.
closeRealtimeSession()
surfaceError(err, context = "voice_config")
}
/**
* Per-turn completion in a persistent realtime session (ADR follow-up). The
* socket stays open; this just finalizes the spoken turn and re-arms
* continuous listening if enabled.
*/
private fun onRealtimeTurnComplete(summary: RealtimeVoiceSummary) {
DiagnosticsLog.record(
category = DiagnosticCategory.Voice,
severity = DiagnosticSeverity.Info,
title = "Realtime voice turn complete",
)
Log.i(
TAG,
"Realtime persistent turn complete provider=${summary.provider} " +
"audioChunks=${summary.audioChunks}",
)
pendingInTtsQueue.set(0)
maybeAutoResume()
}
/**
* Submit a follow-up utterance onto the already-open persistent session
* (turns 2+). Mirrors the per-turn reset the open path does for turn 1.
*/
private fun submitRealtimeTurn(chatVm: ChatViewModel, inputPcm: ByteArray, inputSampleRate: Int) {
val channel = realtimeTurnChannel ?: return
drainQueuedLocalTts()
try { player?.stop() } catch (_: Exception) { /* ignore */ }
firstFrameWatchdogJob?.cancel(); firstFrameWatchdogJob = null
clearSpokenChunksState()
audioSeen.set(false)
audioBytes.set(0)
bargeInStarted.set(false)
lastRealtimeAudioEventId.set(0L)
responseText = StringBuilder()
inputTranscript = StringBuilder()
spokenStatusKeys.clear()
rtUserText = ""
rtConversationContext = chatVm.realtimeAgentContextMessages()
rtAssistantMessageId = chatVm.startRealtimeAgentTurn(
userText = "",
chatSessionId = chatVm.currentSessionId.value,
)
providerRealtimeAgentTurnActive.set(true)
_uiState.update {
it.copy(
state = VoiceState.Thinking,
outputAudioActive = false,
transcribedText = null,
responseText = "",
)
}
val sent = channel.trySend(RealtimeTurnInput(inputPcm, inputSampleRate)).isSuccess
Log.i(TAG, "Realtime persistent turn submitted sent=$sent pcmBytes=${inputPcm.size}")
}
/** Tear down the persistent realtime session (voice-mode exit / engine switch). */
private fun closeRealtimeSession() {
val hadSession = realtimeTurnChannel != null || realtimeSessionJob != null
realtimeTurnChannel?.close()
realtimeTurnChannel = null
realtimeSessionJob?.cancel()
realtimeSessionJob = null
if (hadSession) Log.i(TAG, "Realtime persistent session closed")
}
private fun realtimeToolStatusLine(toolName: String?): String {
val normalized = toolName.orEmpty().lowercase()
return when {
@@ -3899,6 +4101,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
override fun onCleared() {
super.onCleared()
closeRealtimeSession()
// B4: release the listener + VAD engine native resources before
// anything else. stopBargeInListener is defensive / idempotent.
stopBargeInListener()
@@ -1,5 +1,7 @@
package com.hermesandroid.relay.network
import com.hermesandroid.relay.network.models.SkillListResponse
import kotlinx.serialization.json.Json
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNotEquals
@@ -220,6 +222,46 @@ class HermesApiClientTest {
assertEquals("http://localhost:8642/v1/models", url)
}
// --- Skills endpoint compatibility ---
@Test
fun skillEndpointOrder_prefersUpstreamV1ThenLegacyApiFallback() {
assertEquals(listOf("/v1/skills", "/api/skills"), HERMES_SKILL_ENDPOINTS)
}
@Test
fun skillListResponse_parsesUpstreamV1DataEnvelope() {
val body = """
{
"object": "list",
"data": [
{"name": "android", "description": "Control phone", "category": "android"}
]
}
""".trimIndent()
val parsed = Json { ignoreUnknownKeys = true }.decodeFromString<SkillListResponse>(body)
assertEquals(1, parsed.data?.size)
assertEquals("android", parsed.data?.first()?.name)
}
@Test
fun parseSkillListBody_acceptsUpstreamV1DataEnvelope() {
val body = """
{
"object": "list",
"data": [
{"name": "android", "description": "Control phone", "category": "android"}
]
}
""".trimIndent()
val parsed = parseSkillListBody(Json { ignoreUnknownKeys = true }, body)
assertEquals(listOf("android"), parsed?.map { it.name })
}
@Test
fun urlConstruction_deleteSession() {
val baseUrl = "http://localhost:8642"
+5 -3
View File
@@ -259,8 +259,9 @@ GET /api/memory?target=memory // or target=user
### Skills
```
GET /api/skills # optional ?category= filter
GET /api/skills/{name}
GET /v1/skills # native upstream list metadata
GET /api/skills # legacy fallback, optional ?category= filter
GET /api/skills/{name} # legacy detail route when provided by fork/bootstrap
```
> `/api/skills/categories` was removed from upstream as dead code (commit 8d023e43) and is not re-injected by the bootstrap.
@@ -272,7 +273,8 @@ Probe endpoints to detect what's available:
GET /health -> basic connectivity
GET /api/sessions -> enhanced Hermes session API
GET /v1/models -> model listing (OpenAI-compatible)
GET /api/skills -> skills support
GET /v1/skills -> native skills support
GET /api/skills -> legacy skills fallback
GET /api/memory -> memory support
GET /api/config -> config API
+211 -26
View File
@@ -80,15 +80,19 @@ The app supports two streaming endpoints, selectable in Settings:
| **Sessions** (`/api/sessions/{id}/chat/stream`) | Inline text annotations (`` `💻 terminal` ``) — client parses from markdown | Hermes-native SSE (assistant.delta, tool.progress, etc.) or OpenAI-format (delta.content) |
| **Runs** (`/v1/runs` + `/v1/runs/{run_id}/events`) | **Structured events** (tool.started, tool.completed) — real-time tool cards | Hermes lifecycle events (message.delta, tool.started, tool.completed, run.completed) |
**Important upstream note:** The `/api/sessions` CRUD endpoints are moving
toward upstream Hermes core through focused PR
[#29302](https://github.com/NousResearch/hermes-agent/pull/29302), which covers
**Important upstream note:** The `/api/sessions` CRUD/chat endpoints landed in
upstream Hermes core via PR
[#33134](https://github.com/NousResearch/hermes-agent/pull/33134) / commit
[`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde),
which salvaged the focused PR
[#29302](https://github.com/NousResearch/hermes-agent/pull/29302). It covers
session list/create/read/update/delete, messages, fork, chat, and chat stream.
Until that reaches a released core build, `hermes_relay_bootstrap/` still ships
with the plugin and runs at Python interpreter startup via `.pth`. The bootstrap
now composes with partial upstream support: native routes win per method/path,
and the relay only injects missing compatibility gaps such as config, skills, or
memory when core does not provide them.
`hermes_relay_bootstrap/` still ships with the plugin and runs at Python
interpreter startup via `.pth` for older cores and for compatibility gaps that
upstream still does not provide. The bootstrap composes with partial upstream
support: native routes win per method/path, and the relay only injects missing
compatibility gaps such as search, config, legacy skills detail routes,
available-models, or memory when core does not provide them.
The app's `probeCapabilities()` returns a per-endpoint snapshot, and `ConnectionViewModel.resolveStreamingEndpoint()` collapses `streamingEndpoint = "auto"` (the default for new installs) to a concrete `"sessions"` or `"runs"` choice based on what the server actually exposes.
@@ -499,11 +503,14 @@ Adopting from ARC's workflow patterns:
### 16. Runtime API Server Patch via .pth Bootstrap (2026-04-12)
**Context:** The Android app depends on API-server routes for session history,
profile/config metadata, skills, and memory-backed UI. Upstream core is now
moving in focused pieces rather than one large frontend API patch: PR
[#29302](https://github.com/NousResearch/hermes-agent/pull/29302) covers the
canonical `/api/sessions/*` surface, while config/skills/memory still remain
compatibility routes in this repo until core exposes stable equivalents. Without
profile/config metadata, skills, and memory-backed UI. Upstream core has started
absorbing the surface in focused pieces rather than one large frontend API patch:
PR [#33134](https://github.com/NousResearch/hermes-agent/pull/33134) / commit
[`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde)
now covers the canonical `/api/sessions/*` surface, and `/v1/skills` covers
skill list metadata. Config, legacy skill detail routes, memory, search, and
provider-aware model metadata still remain compatibility routes in this repo
until core exposes stable equivalents or clients migrate away from them. Without
the bootstrap, users on older vanilla upstream builds lose session browsing,
metadata-backed settings, and history-on-restart behavior.
@@ -520,16 +527,17 @@ We considered four options:
1. **Zero modifications to hermes-agent's filesystem.** `git pull` / `hermes update` see no local changes, so they always work cleanly. The patch lives entirely in `hermes_relay_bootstrap/` inside our own repo.
2. **Single-file containment of all ported logic.** `_handlers.py` is 500 lines of straight-line aiohttp handler code with explicit `adapter` parameters (closures, not bound methods). Easy to audit, easy to delete.
3. **Feature detection by method/path, not broad route family.** Native upstream
routes win one method/path at a time. This matters because PR #29302 can land
`/api/sessions/*` before core has stable config/skills/memory APIs; the
bootstrap must not skip those remaining compatibility routes just because a
sessions route exists.
routes win one method/path at a time. This matters because PR #33134 landed
`/api/sessions/*` before core had stable config/legacy skill detail/memory/search
APIs; the bootstrap must not skip those remaining compatibility routes just
because a sessions or `/v1/skills` route exists.
4. **Trust model already established.** The user installed our plugin into their hermes-agent venv. They've already consented to having the plugin import hermes-agent internals (it does this for relay tools, voice endpoints, media registry). Monkey-patching `aiohttp.web.Application` is in the same trust bucket.
5. **Surface-by-surface removal.** When PR #29302 or equivalent reaches a
released hermes-agent version, the sessions compatibility routes should go
quiet automatically. Config, skills, memory, and command preprocessing remain
until their native replacements exist. Full bootstrap deletion happens only
after every compatibility route group has a stable core equivalent.
5. **Surface-by-surface removal.** When a supported hermes-agent baseline includes
PR #33134 / `f7527b0`, the sessions compatibility routes should go quiet
automatically. Config, legacy skills detail routes, memory, search,
available-models, and command preprocessing remain until their native
replacements exist or the clients migrate. Full bootstrap deletion happens
only after every compatibility route group has a stable core equivalent.
6. **`/v1/runs` is genuinely better for chat than `/api/sessions/{id}/chat/stream`.** It's standard upstream, supports `X-Hermes-Session-Id` for continuation, and emits live structured tool events. The fork's chat handler exists because upstream didn't HAVE this clean structured-event runs API at the time the fork was cut — but upstream does now.
**The Android client adapts via `streamingEndpoint = "auto"`.** New `ServerCapabilities` data class returned by `HermesApiClient.probeCapabilities()` captures per-endpoint presence (`sessionsApi`, `sessionsChatStream`, `runs`, `portable`, `healthy`). `ConnectionViewModel.resolveStreamingEndpoint()` collapses `"auto"` to `"sessions"` (when chat-stream handler is present, i.e. fork or upstream-merged) or `"runs"` (otherwise, i.e. bootstrap-injected vanilla upstream). The setting still supports manual `"sessions"` / `"runs"` overrides for debugging.
@@ -548,10 +556,11 @@ We considered four options:
- `install.sh` step 2 — copies the `.pth` into the venv site-packages
**Removal path** is now per surface:
1. Sessions: after PR #29302 or equivalent ships in released core, keep the
1. Sessions: with PR #33134 / `f7527b0` in the supported core baseline, keep the
bootstrap installed but verify it skips native `/api/sessions/*` routes.
2. Config/skills/memory: remove those compatibility handlers only after stable
core APIs exist and Android probes prefer them.
2. Config/legacy skills/memory/search/available-models: remove those
compatibility handlers only after stable core APIs exist and Android probes
prefer them, or after clients migrate away from the legacy route group.
3. Slash middleware: remove after native API-server slash preprocessing exists.
4. Full cleanup: delete `hermes_relay_bootstrap/`, delete
`hermes_relay_bootstrap.pth`, remove the `.pth` install block, and update
@@ -928,7 +937,7 @@ The plan called for adding a `// VOICE HOOK` callback to `ChatViewModel` so `Voi
- `plugin/relay/server.py` — route registration alongside `/media/*`
- `plugin/tests/test_voice_routes.py` — 14 unit tests, `unittest`-based (pytest conftest issue documented in `CLAUDE.md`)
- `app/src/main/kotlin/.../audio/VoiceRecorder.kt` — MediaRecorder amplitude StateFlow
- `app/src/main/kotlin/.../audio/VoicePlayer.kt` — MediaPlayer + Visualizer amplitude StateFlow with OEM fallback
- `app/src/main/kotlin/.../audio/VoicePlayer.kt` — Media3 ExoPlayer (gapless TTS queue) + Visualizer amplitude StateFlow with OEM fallback; `audioSessionId` served from a thread-safe `@Volatile` cache (read off-main by barge-in)
- `app/src/main/kotlin/.../network/RelayVoiceClient.kt` — OkHttp multipart + JSON clients
- `app/src/main/kotlin/.../viewmodel/VoiceViewModel.kt` — turn state machine, sentence detection, TTS queue consumer, `ChatViewModel` observation pattern
- `app/src/main/kotlin/.../ui/components/VoiceModeOverlay.kt` — full-screen overlay, VoiceState→SphereState mapping, three interaction modes
@@ -1525,3 +1534,179 @@ old STT -> Hermes -> TTS pipeline instead of a native realtime agent.
- `plugin/relay/realtime_agent/broker.py`
- `plugin/relay/realtime_agent/providers/xai.py`
- `plugin/relay/realtime_agent/providers/openai.py`
---
## ADR 33 - Realtime Agent foregrounds short Hermes turns and promotes long ones to background tasks
**Status:** Accepted, phased (2026-05-24). Default-on at feature GA, gated by a
prerequisite provider-idle-tolerance spike (Phase 0). Supersedes the
blocking-broker assumption inside ADR 32's sequence; ADR 32's preamble ->
forced-Hermes -> provider-summary shape is preserved for the foreground tier.
**Context.** Today a Realtime Agent turn runs Hermes *synchronously inside the
provider event pump*: `_pump_provider_events` awaits `_handle_provider_tool_call`
-> `_run_brokered_tool`, which streams the entire Hermes SSE run to completion
before the pump consumes the next provider event (`broker.py`). This is correct
and lowest-latency for short Q&A, but it has two costs that grow with task
length:
- The provider realtime socket sits attached-but-idle for the whole run (a live,
billed, audio-clocked WebSocket parked for tens of seconds during research,
multi-tool, or desktop/build tasks).
- The tool surface already advertises a background vocabulary
(`hermes_run_task`, `hermes_get_status`, `hermes_cancel`, `hermes_confirm`)
and a `hermes_run_status` state machine, but `hermes_get_status` /
`hermes_cancel` as *provider* tool calls are unreachable mid-run because the
pump is parked. Only the client->relay `response.cancel` path can interrupt.
The blocking `await` is also an *implicit mutex*: it serializes the three audio
sources that can produce `voice.output_audio.delta` (see "Who speaks" below) so
they never overlap. Removing it requires replacing that mutex with an explicit
floor owner.
**Who speaks (confirmed against both supported providers).** Up to three mouths
exist; only one is active per phase today because of the blocking await:
1. **Realtime provider** - xAI `grok-voice-latest`, OpenAI `gpt-realtime-2`.
Both run with `turn_detection: None` (relay owns turn boundaries; no server
VAD) and audio output modality. Speaks the pre-Hermes acknowledgement and the
final post-result summary.
2. **Relay TTS render** - `xai_tts`/`openai_tts` via `_render_provider_audio`.
A separate, non-realtime synthesizer used as the forced-summary fallback and
the legacy render path. Emits the *same* `voice.output_audio.delta` wire event
as the provider, so Android cannot distinguish them.
3. **Android local TTS** - the `should_speak` long-wait filler, driven by
`hermes.run.progress`. Client-side, not the provider.
Because all three converge on one Android `AudioTrack`, **floor arbitration must
happen relay-side, before bytes reach the socket.**
**Decision.** Keep three turn classes; make promotion automatic and default-on,
with the relay as the single explicit floor owner.
- **Tier A - Foreground (short Q&A).** Unchanged from ADR 32: provider preamble
-> relay-forced Hermes (still awaited) -> provider summary. Lowest latency,
trivial floor. This remains the path for any turn that completes inside the
promotion grace window.
- **Tier B - Promoted (long task detected late).** Start in Tier A. If the
Hermes run has not produced a final result within a grace window
(`promote_after_ms`, default ~6000ms, tunable), the relay *detaches* the run
from the pump: the run continues as a tracked `asyncio.Task`, the pump resumes
consuming provider events, and the provider speaks a short "I've started that -
I'll let you know" handoff. Progress continues via `hermes.run.progress`. When
the background run completes, the relay injects the result as a tool result and
requests a provider summary at the next floor-idle moment.
- **Tier C - Explicitly durable.** `hermes_run_task(mode="background")` returns a
run handle immediately (no grace window). For tasks the model/profile knows up
front are long (research, builds via desktop tools, multi-step). Same
completion-injection path as Tier B.
Promotion is the default behavior, not a flag. Grace-period promotion preserves
Tier A latency for the common case and only forks when a run actually proves
long, so the user never has to pick a mode.
**The relay is the single floor owner.** A per-session floor state
(`idle | provider_speaking | hermes_filler | result_pending`) gates every audio
source:
- Only one mouth may hold the floor. The provider holds it by default.
- A completed background result does **not** barge in. It is queued as
`result_pending` and spoken only when the floor returns to `idle` (provider
finished, user not mid-utterance). Provider VAD being off means the relay
controls `response.create`, so it can withhold the summary until the floor is
clear.
- Android local filler (`should_speak`) is suppressed whenever the provider holds
the floor; it is a Tier B/C long-wait affordance only.
- Relay TTS render (mouth 2) may only fire when it owns the floor and the
provider has drained, exactly as the forced-summary fallback does today.
**Settings (ample, per ADR intent).** Surface in Voice Settings -> Realtime
Agent, with relay-side `realtime_voice` config as source of truth and per-profile
override:
- `promotion_enabled` (default true) - master switch; false pins Tier A blocking.
- `promote_after_ms` (default ~6000) - grace window before Tier B handoff.
- `background_default_mode` - whether ambiguous long turns prefer promote vs.
stay-foreground.
- `spoken_handoff` (default true) - speak the "I've started that" line on
promotion vs. silent + visual only.
- `progress_spoken_after_ms` / `progress_repeat_ms` - reuse existing
`_HERMES_SPOKEN_PROGRESS_*` knobs, now configurable.
- `result_delivery` - `speak_when_idle` (default) vs. `notify_then_speak`
(chime/visual, speak on user re-engage) vs. `visual_only`.
- `max_background_runs` - concurrent background runs per session (default 1 for
the MVP; the existing single-`hermes_task` field assumes 1).
**Protocol additions (relay <-> Android, additive).**
- `hermes.run.promoted` - run moved to background; carries `run_id`,
`promote_after_ms`, `spoken_handoff`.
- `hermes.run.background_completed` - background run finished; precedes the
provider/relay summary.
- Extend `hermes.run.progress` with `tier` and `floor` so the client can render
background state distinctly (e.g. a persistent "working on: ..." chip).
- `hermes_get_status` / `hermes_cancel` become genuinely reachable as provider
tool calls in Tier B/C because the pump is no longer parked; no schema change.
**Prerequisite (Phase 0 spike) — RESOLVED 2026-05-24.** The spike asked how xAI
and OpenAI realtime sessions behave when held open and idle. Verdicts (see
`docs/realtime-voice-poc.md`): **OpenAI `hold-floor-ok` (empirical** — survived
10/20/30s idle with clean post-idle audio); **xAI `hold-floor-ok`** (shipping
Realtime Agent already holds `xai_realtime` sessions open across between-turn
idle with `turn_detection: None` + resume TTL; relay-host probe retained as a
regression check, not a precondition). The premise was also superseded in
implementation: Tier B closes the pending provider call with an interim ack
rather than holding an open response, so the socket only sees the normal
between-turns idle gap — no provider needs the `must-reopen` fallback today, and
default-on is unblocked.
**Rules.**
- Hermes remains the only path for tools, memory, current data, research, side
effects, durable context, and confirmations (unchanged from ADR 29/32).
- Exactly one audio source may hold the floor at a time; the relay enforces this
before audio reaches Android. Background results never barge in.
- Android must not read raw Hermes output aloud; spoken output is always a
provider (or relay-fallback) summary of the compact Hermes result.
- Promotion must be cancel-safe: a promoted/background run is cancelable via both
the provider `hermes_cancel` tool and the client `response.cancel` path, reusing
`_cancel_active_hermes`.
- A turn must never strand: if a background run completes while the WS is
detached, the result is replayed through the existing event-ring/resume path on
reattach.
**Open questions / risks.**
- Floor arbitration is the hard part and the existing `native_forced_*` /
`_should_forward_provider_response_event` suppression machinery is already the
most fragile code in `broker.py`. Concurrency stresses it most; the floor-owner
state machine should *replace* ad-hoc suppression, not stack on top of it.
- Concurrent background runs (`max_background_runs > 1`) are out of scope for the
MVP; the session model assumes a single `hermes_task`.
- "Notify then speak" result delivery needs an Android affordance (chime +
chip + tap-to-hear) that does not yet exist.
**Phased rollout.**
1. **Phase 0** - provider-idle-tolerance spike (above). Gate for default-on.
2. **Phase 1** - introduce the relay floor-owner state machine under the current
blocking behavior (no functional change), with tests that prove single-floor
invariants.
3. **Phase 2** - Tier B grace-period promotion behind `promotion_enabled`,
default **off**, validated on long-task transcripts.
4. **Phase 3** - flip `promotion_enabled` default **on**; add Tier C
`mode="background"`; ship the Voice Settings surface.
**Key Files:**
- `docs/decisions.md` (this ADR; supersedes ADR 32's blocking assumption)
- `docs/plans/2026-05-24-realtime-background-hermes-runs.md` (companion plan)
- `plugin/relay/realtime_agent/broker.py`
- `plugin/relay/realtime_agent/hermes_tool_broker.py`
- `plugin/relay/realtime_agent/models.py`
- `plugin/relay/realtime_agent/providers/xai.py`
- `plugin/relay/realtime_agent/providers/openai.py`
- `plugin/relay/profile_voice.py`
- `docs/relay-protocol.md`
- `app/src/main/kotlin/com/hermesandroid/relay/viewmodel/VoiceViewModel.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/ui/screens/VoiceSettingsScreen.kt`
@@ -0,0 +1,330 @@
# Plan: Realtime Agent background Hermes runs (foreground -> promote -> durable)
> **Purpose.** Stop the Realtime Agent's provider event pump from blocking on
> the full Hermes run. Keep short turns synchronous and low-latency, but let
> genuinely long Hermes runs detach to a tracked background task so the provider
> session stays responsive and the run can be monitored, summarized on
> completion, and cancelled.
>
> **Decision of record.** ADR 33 (`docs/decisions.md`). This plan is the
> phased, file-level execution of that ADR. Read ADR 33 first — it defines the
> three tiers, the relay-as-single-floor-owner rule, the settings surface, and
> the Phase 0 prerequisite spike.
>
> **Non-negotiable.** Hermes remains the only authority for tools, memory,
> current data, research, side effects, durable context, and confirmations
> (ADR 29/32). This plan changes *when and how* a Hermes run is awaited and
> spoken — never whether Hermes governs it.
## Current State (grounded in code)
The realtime-agent broker already runs a persistent provider event pump; the
gap is that tool execution blocks it.
- `plugin/relay/realtime_agent/broker.py`
- `_pump_provider_events` (~`broker.py:1124`) is the long-lived
`native_provider_task` — the persistent loop. It survives WS detach/resume.
- `FUNCTION_CALL_COMPLETED` -> `_handle_provider_tool_call` (~`broker.py:1628`)
-> `_run_brokered_tool` (~`broker.py:1930`, `return await task`) ->
`_execute_brokered_tool` (~`broker.py:1956`) `async for`s the entire Hermes
SSE stream before the pump consumes the next provider event. **This is the
blocking await to remove for long runs.**
- `_send_hermes_run_progress` (~`broker.py:2245`) already runs concurrently as
`progress_task`, emitting `hermes.run.progress` every
`_HERMES_PROGRESS_INTERVAL_SECONDS` and spoken filler after
`_HERMES_SPOKEN_PROGRESS_AFTER_SECONDS`. Reuse this for background progress.
- `_cancel_active_hermes` (~`broker.py:2295`) already cancels `hermes_task`
cross-coroutine. Promotion must keep this working.
- `RealtimeAgentSession` (~`broker.py:114`) holds `hermes_task`,
`hermes_run_status`, and the `native_forced_*` suppression flags. Floor state
will live here.
- `_render_provider_audio` (~`broker.py:2509`) is the relay TTS render mouth
(forced-summary fallback). It must obey the floor owner.
- `_request_playback_drain` / `_send_provider_tool_result` /
`_request_provider_response` (~`broker.py:1700`-`1810`) are the existing
result-injection primitives the background-completion path will reuse.
- `plugin/relay/realtime_agent/hermes_tool_broker.py`
- `_TOOL_SURFACE = ("hermes_run_task", "hermes_get_status", "hermes_cancel",
"hermes_confirm")` (~`:189`) — background vocabulary already advertised.
- `stream_task` yields the run; unchanged by this plan.
- `plugin/relay/realtime_agent/models.py` — `SERVER_EVT_*`, `CLIENT_MSG_*`,
`HERMES_TOOL_SCHEMAS`. New events/fields added here.
- Providers (`providers/xai.py`, `providers/openai.py`) — both run
`turn_detection: None` (relay owns `response.create`) and emit audio. This is
what makes "withhold the summary until the floor is idle" possible.
- Settings: `plugin/relay/profile_voice.py` (`realtime_voice_settings`,
`save_profile_voice_section`) + `plugin/relay/config.py`.
- Android: `viewmodel/VoiceViewModel.kt`, `network/RelayVoiceClient.kt`,
`data/VoicePreferences.kt`, `data/RealtimeConversationContext.kt`,
`voice/RealtimeTurnSyncBuilder.kt`, `ui/screens/VoiceSettingsScreen.kt`.
- Tests: `plugin/tests/test_realtime_agent_routes.py`,
`test_realtime_agent_xai_provider.py`, `test_realtime_agent_openai_provider.py`.
### Three audio sources ("who speaks")
All three converge on Android's single `AudioTrack` via `voice.output_audio.delta`:
1. Realtime provider (xAI/OpenAI) — primary.
2. Relay TTS render (`_render_provider_audio`) — fallback; same wire event.
3. Android local TTS — `should_speak` filler on `hermes.run.progress`.
Today the blocking await serializes them. This plan replaces that implicit mutex
with an **explicit relay-side floor owner**.
## Target Architecture
```text
provider event pump (never blocks on a long run)
-> short run: await inline (Tier A, unchanged ADR 32 shape)
-> long run: detach to tracked task, resume pump, speak handoff (Tier B)
-> explicit durable: return handle immediately (Tier C)
|
v
FloorOwner(idle | provider_speaking | hermes_filler | result_pending)
|
background run completes -> queue result_pending -> speak when floor idle
```
**FloorOwner contract (relay-side, per session):**
- Exactly one mouth holds the floor; provider holds it by default.
- A completed background result is queued `result_pending`, never barges in. It
is spoken (provider summary, or relay-TTS fallback) only on transition to
`idle`.
- Android `should_speak` filler is suppressed while `provider_speaking`.
- `_render_provider_audio` may only fire when it owns the floor and the provider
has drained.
## Non-Goals
- Do not change Tier A latency or the ADR 32 preamble -> forced-Hermes ->
provider-summary shape for short turns.
- Do not give the provider authority over tools/memory/confirmations.
- Do not support `max_background_runs > 1` in this plan (single `hermes_task`).
- Do not require adb/device testing for verification unless Bailey asks. JVM/py
unit tests plus the lab smoke gate cover acceptance.
- Do not depend on Hermes upstream changes; the broker already owns the loop.
## Protocol Additions (additive only)
Relay -> client (add to `models.py`):
- `hermes.run.promoted` — `{run_id, promote_after_ms, spoken_handoff, tier}`.
- `hermes.run.background_completed` — `{run_id, ok, tool_count}` (precedes summary).
- Extend `hermes.run.progress` with `tier` (`foreground|promoted|durable`) and
`floor` (`idle|provider_speaking|hermes_filler|result_pending`).
No new client->relay messages required — `response.cancel`, `hermes_cancel`, and
`hermes_get_status` already exist; the latter two become reachable mid-run once
the pump no longer blocks.
## Settings Surface (per ADR 33)
Relay-side `realtime_voice` config (source of truth, per-profile override via
`profile_voice.py`), mirrored into Android `VoicePreferences`:
| Key | Default | Meaning |
|---|---|---|
| `promotion_enabled` | `true` (Phase 3) | Master switch; `false` pins Tier A blocking |
| `promote_after_ms` | `6000` | Grace window before Tier B handoff |
| `background_default_mode` | `promote` | Ambiguous long turn: `promote` vs `foreground` |
| `spoken_handoff` | `true` | Speak "I've started that" on promotion |
| `progress_spoken_after_ms` | `15000` | Reuse `_HERMES_SPOKEN_PROGRESS_AFTER_SECONDS` |
| `progress_repeat_ms` | `30000` | Reuse `_HERMES_SPOKEN_PROGRESS_REPEAT_SECONDS` |
| `result_delivery` | `speak_when_idle` | vs `notify_then_speak` / `visual_only` |
| `max_background_runs` | `1` | Fixed at 1 this plan |
---
## Phase 0 — Provider idle-tolerance spike (GATES default-on)
**Goal.** Determine empirically whether xAI and OpenAI realtime sessions tolerate
being held open and quiescent (no `response.create`, no input audio) for
30–120s, since Tier B holds the floor while a background run completes.
**Tasks.**
1. Add a throwaway script under `plugin/tools/` or a `scripts/` probe (not
shipped) that opens each provider realtime socket via the existing adapters,
sends `session.update`, then idles 30/60/120s and logs: socket survival, any
server-side timeout/close codes, VAD/turn artifacts on the first post-idle
`response.create`, and any keep-alive requirement.
2. Record findings in `docs/realtime-voice-poc.md` under a new "Idle tolerance"
section, per provider.
**Acceptance.**
- A documented per-provider verdict: `hold-floor-ok` | `needs-keepalive` |
`must-reopen`. This verdict selects the Tier B fallback strategy per provider.
**Note.** If a provider is `must-reopen`, Tier B for that provider detaches the
run but closes/reopens the provider socket (or keeps a minimal keep-alive)
rather than holding it conversational. Capture that in the ADR's Phase 0 line.
**Status: DONE (2026-05-24).** Verdicts recorded in
`docs/realtime-voice-poc.md` → Idle tolerance:
- **OpenAI `gpt-realtime-2` — `hold-floor-ok` (empirical).** Probe ran live
(10s/20s/30s idle windows survived; clean post-idle audio each time).
- **xAI `grok-voice-latest` — `hold-floor-ok`.** No xAI creds on the dev box, but
the verdict is not conditional: the shipping Realtime Agent already holds
`xai_realtime` sessions open across between-turn idle gaps
(`turn_detection: None` + resume TTL) with no idle-close reports. A relay-host
probe run is retained as a regression check, not a precondition.
**Premise superseded.** Phase 2/3 implemented Tier B by *closing the pending
provider call with an interim ack* rather than holding an open response, so the
"hold the floor conversational while a run completes" worst case this spike
guarded against does not occur — the socket only sees the normal between-turns
idle gap. Both providers are `hold-floor-ok`, so the default-on gate is satisfied
(no provider needs the `must-reopen` fallback today).
---
## Phase 1 — Introduce FloorOwner under current blocking behavior (no functional change)
**Goal.** Add the explicit floor state machine and route all three mouths through
it, while preserving today's exact audible behavior. This is the de-risking step.
**Tasks.**
1. Add `RealtimeFloor` (new dataclass/enum) to `broker.py` or a new
`realtime_agent/floor.py`: state `idle|provider_speaking|hermes_filler|
result_pending`, with `acquire(mouth)`, `release(mouth)`, and
`can_speak(mouth) -> bool`. Pure, unit-testable, no I/O.
2. Hold a `floor: RealtimeFloor` on `RealtimeAgentSession`.
3. Gate the three emit paths through `floor.can_speak(...)`:
- provider audio in `_pump_provider_events` (`AUDIO_DELTA` -> `provider`),
- `_render_provider_audio` (-> `relay_tts`),
- `should_speak` progress events (-> `android_filler`; suppress while
`provider_speaking`).
4. Transition floor on `RESPONSE_STARTED`/`AUDIO_DONE`/`RESPONSE_DONE` and on
playback-drain completion.
5. Stamp `floor` onto `hermes.run.progress` (additive field).
**Acceptance.**
- All existing `plugin/tests/test_realtime_agent_*` pass unchanged.
- New `test_realtime_floor.py` proves single-mouth invariants:
background-result-never-barges, filler-suppressed-while-provider-speaks,
relay-TTS-only-when-owned.
- Manually/log-verified: a short turn sounds identical to pre-change.
**Test gate.** `python -m unittest plugin.tests.test_realtime_floor` +
existing realtime agent tests green.
---
## Phase 2 — Tier B grace-period promotion (default OFF)
**Goal.** Detach a long Hermes run from the pump after `promote_after_ms`; resume
the pump; speak a handoff; deliver the result on completion via the floor owner.
**Tasks.**
1. `models.py`: add `hermes.run.promoted`, `hermes.run.background_completed`,
`tier` field; add settings keys to the config schema.
2. `profile_voice.py` / `config.py`: read/write the new `realtime_voice` keys
with defaults above (`promotion_enabled=false` this phase).
3. `broker.py` — split `_run_brokered_tool` for `hermes_run_task`:
- Start the run task as today, but `await asyncio.wait({task}, timeout=
promote_after_ms)`.
- If done in time: Tier A path, unchanged.
- If not: emit `hermes.run.promoted`, optionally speak handoff (gated by
`spoken_handoff` + floor), set `tier="promoted"`, **return control to the
pump** without awaiting the task. Keep `session.hermes_task` set.
4. Add a pump-side drain point: between provider events (and on a small timer),
check for a finished background task; when finished, emit
`hermes.run.background_completed`, then inject via the existing
`_send_provider_tool_result` -> `_request_provider_response` path **only when
`floor` is `idle`** (else mark `result_pending` and inject on next `idle`).
5. Preserve cancellation: `hermes_cancel` (now reachable) and `response.cancel`
both route to `_cancel_active_hermes`.
6. Resume safety: if WS detaches mid-background-run, the completion event must
replay through the existing event-ring/resume path on reattach.
**Acceptance.**
- With `promotion_enabled=true` in test config, a run that exceeds
`promote_after_ms` emits `hermes.run.promoted`, the pump processes a subsequent
provider event before the run finishes (proves non-blocking), and the result is
spoken exactly once after completion.
- A run under the window behaves as Tier A (no `promoted` event).
- Cancel during a promoted run stops it and emits `hermes.run.cancelled`.
- Detach+resume during a promoted run replays `background_completed`.
**Test gate.** New `test_realtime_promotion.py` (fake provider connection + fake
Hermes broker with controllable latency) covering: under-window, over-window,
cancel-promoted, detach-resume-completes, result-waits-for-idle-floor.
---
## Phase 3 — Default-on + Tier C durable + Android settings UI
**Goal.** Flip `promotion_enabled` default to `true` (gated on Phase 0 verdict),
add explicit `mode="background"`, and expose settings on Android.
**Tasks.**
1. `models.py` `HERMES_TOOL_SCHEMAS`: document/validate `hermes_run_task.mode`
(`run|background`); `background` returns a handle immediately (skip grace
window, go straight to Tier B detach).
2. Default `promotion_enabled=true`; per-provider Tier B strategy from Phase 0.
3. Android `VoicePreferences.kt` + `RelayVoiceClient.kt`: read/write the new
settings; surface in `VoiceSettingsScreen.kt` under Realtime Agent (promotion
toggle, grace slider, handoff toggle, result-delivery picker).
4. Android `VoiceViewModel.kt` + `RealtimeTurnSyncBuilder.kt`: handle
`hermes.run.promoted` / `hermes.run.background_completed`; render a persistent
"working on: …" chip while `tier=promoted/durable`; implement
`notify_then_speak` affordance (chime + tap-to-hear) if that mode is selected.
5. Docs: `docs/relay-protocol.md` (new events/fields), `user-docs/features/
voice.md` (settings + behavior), `CHANGELOG.md` `[Unreleased]`.
**Acceptance.**
- Default config promotes long runs without user action; short runs unaffected.
- `hermes_run_task(mode="background")` returns immediately and completes via the
same path.
- Android shows promoted state and speaks the result per `result_delivery`.
- `./gradlew lint` clean; py tests green; lab smoke
(`scripts/realtime-voice-lab-smoke.ps1`) still under threshold for short turns.
**Test gate.** Full `plugin.tests.test_realtime_agent_*` + new floor/promotion
suites; Android unit tests for the new `VoicePreferences`/sync mapping;
`./gradlew lint`.
---
## Risks & Mitigations
- **Floor logic stacking on `native_forced_*` suppression.** The forced-summary
state machine is the most fragile code in `broker.py`. Mitigation: Phase 1
introduces FloorOwner *as the serializer*; migrate suppression decisions to ask
the floor rather than adding parallel flags.
- **Double-speak (provider + relay TTS).** Mitigation: both gated by
`floor.can_speak`; result injection only on `idle`.
- **Provider idle close.** Mitigation: Phase 0 verdict picks per-provider Tier B
strategy; `must-reopen` providers don't hold the floor.
- **Stranded background result on disconnect.** Mitigation: reuse event-ring +
resume replay; covered by Phase 2 acceptance.
## Test Strategy Summary
- Pure unit: `RealtimeFloor` invariants (Phase 1).
- Broker integration with fakes: promotion timing, cancel, resume, idle-gated
delivery (Phase 2).
- Android: settings round-trip + event handling (Phase 3).
- Regression: existing realtime agent suites unchanged throughout; lab smoke for
short-turn latency.
## Key Files
- `docs/decisions.md` (ADR 33)
- `plugin/relay/realtime_agent/broker.py`
- `plugin/relay/realtime_agent/floor.py` (new)
- `plugin/relay/realtime_agent/hermes_tool_broker.py`
- `plugin/relay/realtime_agent/models.py`
- `plugin/relay/realtime_agent/providers/xai.py`
- `plugin/relay/realtime_agent/providers/openai.py`
- `plugin/relay/profile_voice.py`
- `plugin/relay/config.py`
- `plugin/tests/test_realtime_floor.py` (new)
- `plugin/tests/test_realtime_promotion.py` (new)
- `docs/relay-protocol.md`
- `docs/realtime-voice-poc.md` (Phase 0 findings)
- `app/src/main/kotlin/com/hermesandroid/relay/viewmodel/VoiceViewModel.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/voice/RealtimeTurnSyncBuilder.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/data/VoicePreferences.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/network/RelayVoiceClient.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/ui/screens/VoiceSettingsScreen.kt`
- `user-docs/features/voice.md`
- `CHANGELOG.md`
@@ -0,0 +1,83 @@
# Plan: Persistent realtime-agent session (one socket across turns)
> **Purpose.** Make Realtime Agent voice a *persistent conversation* instead of a
> new session per utterance. Today each turn calls
> `RelayVoiceClient.runRealtimeAgent()` which `createRealtimeAgentSession()`s a
> fresh relay session + provider socket and closes it on `voice.response.done`.
> The provider therefore has no live memory of prior turns (only relay-seeded
> context snippets), and every turn pays session-setup latency.
>
> **Key finding.** This is a **client-only** change. The relay already supports
> many turns on one socket — `_handle_provider_native_ws` loops over
> `input_audio.append` / `input_audio.commit` / `response.create` and only tears
> the session down when the WebSocket disconnects. `voice.response.done` is a
> *turn* boundary, not a *session* boundary. The client is what closes the socket
> on `voice.response.done` (`RelayVoiceClient.kt:1410`).
## Design
Keep one relay session + provider socket open for the lifetime of a Realtime
Agent voice-mode session; feed each utterance as a new turn on that socket.
### RelayVoiceClient
- Add an opt-in **persistent mode** to `runRealtimeAgent` via two optional params
(one-shot path is byte-for-byte unchanged when they're null):
- `turnInputs: ReceiveChannel<RealtimeTurnInput>?` — when non-null, the call is
a long-lived session: after the first turn it reads further turns off the
channel and sends their `input_audio.append`+`commit` on the open socket.
- `onTurnComplete: (RealtimeVoiceSummary) -> Unit` — invoked at each
`voice.response.done` instead of completing+closing.
- In persistent mode:
- `voice.response.done` → call `onTurnComplete`, **do not** close the socket or
complete `finished`.
- A reader coroutine drains `turnInputs` and sends each turn's PCM chunks.
- `finished` completes only when the channel closes (voice-mode exit) or on a
fatal socket/provider error.
- The per-turn idle/turn-limit guards in `awaitRealtimeAgentCompletion` are
scoped to an *active* turn only — between-turn idle is expected and must not
trip `REALTIME_AGENT_IDLE_TIMEOUT_MS`.
- `RealtimeTurnInput(inputPcm, sampleRate, prompt)` data class.
### VoiceViewModel
- Hold a `realtimeLiveSession` (the persistent call's `Job` + the
`SendChannel<RealtimeTurnInput>`), opened lazily on the first Realtime Agent
turn and reused for subsequent turns.
- Per utterance: instead of a fresh `runRealtimeAgentTurn`, do the per-turn UI
setup (state reset, watchdog, `chatVm.startRealtimeAgentTurn`) and
`realtimeLiveSession.submit(RealtimeTurnInput(...))`.
- Close the session (`channel.close()` + cancel job) on `exitVoiceMode`, engine
switch away from Realtime Agent, and `onCleared`.
- The continuous event callback (the `when(event.type)` block) is registered once
for the session and handles every turn's events.
### Fallback flag (risk control)
- `VoicePreferences.realtimePersistentSession` (default **true**). When false,
fall back to the current one-shot `runRealtimeAgentTurn` path. Lets the user
flip back on-device without a rebuild if the persistent path misbehaves.
## Non-Goals
- No relay changes (the relay already supports multi-turn sockets).
- No change to the one-shot path used by the Voice Lab / stable engine.
- Not changing barge-in semantics beyond keeping them working per turn.
## Risks / must-validate-on-device
- Feeding new input while a prior turn's audio is still draining (barge-in vs.
new turn). Mitigation: reuse existing barge-in/cancel before submitting a turn.
- Per-turn UI state resets must not tear down the shared session.
- Between-turn idle must not trip the stall timeout.
- Provider/relay session longevity across long pauses (the socket stays open, so
no resume-TTL dependency — but confirm providers don't idle-close; see ADR 33
Phase 0 `hold-floor-ok`).
## Test strategy
- Unit-test the pure pieces: `RealtimeTurnInput` chunking, the persistent-vs-
one-shot branch decision, idle-guard gating by active-turn.
- On-device (post-merge, by Bailey): multi-turn conversation with follow-up
references ("what did you just say?"), background promotion mid-conversation,
barge-in, exit/re-enter voice mode, engine switch.
## Key Files
- `app/src/main/kotlin/com/hermesandroid/relay/network/RelayVoiceClient.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/viewmodel/VoiceViewModel.kt`
- `app/src/main/kotlin/com/hermesandroid/relay/data/VoicePreferences.kt`
- `docs/relay-protocol.md` (note: turn vs. session boundary)
+68
View File
@@ -1126,3 +1126,71 @@ Pass criteria:
`voice.*.session.detached` followed by `voice.session.resumed`.
- Replayed events are marked `replayed=true`, and the session does not start a
duplicate Hermes run inside the resume TTL.
## Idle tolerance (ADR 33 Phase 0)
Background-Hermes-run promotion (ADR 33, `docs/plans/2026-05-24-realtime-background-hermes-runs.md`)
detaches a long Hermes run from the provider event pump and keeps the provider
session open while the run completes. Whether a provider can hold the floor while
**quiescent** (no `response.create`, no input audio) for tens of seconds is the
factual unknown that gates default-on promotion.
Run the probe on the relay host (provider credentials configured) with the repo
root on `PYTHONPATH`:
```bash
python scripts/realtime-provider-idle-probe.py --provider xai
python scripts/realtime-provider-idle-probe.py --provider openai --windows 30,60,120
```
The probe holds each socket idle across the windows, then issues one short
post-idle turn to check for VAD/turn artifacts, and prints a verdict.
Record the verdict per provider. The verdict selects that provider's Tier B
strategy:
| Verdict | Meaning | Tier B strategy |
|---|---|---|
| `hold-floor-ok` | Socket survives idle; post-idle turn clean | Hold the provider session open during the background run (default) |
| `needs-keepalive` | Survives but post-idle turn degraded | Hold open + send a minimal keep-alive; revalidate the first post-idle turn |
| `must-reopen` | Socket closes/errors while idle | Detach the run but close+reopen (or resume) the provider socket on completion |
### Findings
Both verdicts are **`hold-floor-ok`** and Phase 0 is closed. The probe's original
worst case — a provider holding an *open response* idle for the whole run — does
not occur: the promotion path closes the pending call with an interim ack
(`broker.py:_begin_background_delivery`), so the socket only experiences the
normal between-turns idle gap. That gap is already exercised in production by
every Realtime Agent turn (the session stays open between the user finishing
speaking and the next `response.create`, across the resume TTL).
| Provider | Date | Basis | Windows | Post-idle turn | Verdict |
|---|---|---|---|---|---|
| OpenAI (`gpt-realtime-2`) | 2026-05-24 | empirical (probe) | 10s, 20s, 30s | audio returned, no error | **`hold-floor-ok`** |
| xAI (`grok-voice-latest`) | 2026-05-24 | existing production behavior + protocol parity | between-turn idle in daily use | clean (no idle-close reports) | **`hold-floor-ok`** |
**OpenAI — empirical.** Ran `realtime-provider-idle-probe.py --provider openai
--windows 10,20,30` against the live API. The session stayed open across all
three quiescent windows and produced clean audio on every post-idle
`response.create`. (Incidental observation, not an idle finding: the live API
emitted one `Missing required parameter: 'session.audio.output.format.rate'`
error at `session.update` time — a minor schema drift in
`providers/openai.py:_session_update` worth a follow-up; the session still
functioned and returned audio.)
**xAI — production behavior + parity.** No xAI key/OAuth store is present on the
dev box, so a fresh probe run is deferred to the relay host. The verdict is not
conditional, however: the *shipping* Realtime Agent already holds `xai_realtime`
sessions open across between-turn idle gaps with `turn_detection: None`
(relay-driven turns) and a resume TTL, and there are no reports of xAI closing or
degrading on those idle gaps. Background promotion does not lengthen the
*open-response* duration (the call is closed with an interim ack), so it does not
introduce a new idle condition beyond what xAI already tolerates today. Running
the probe on the relay host is retained as a **regression check**, not a
precondition. If it ever returns `needs-keepalive`/`must-reopen`, set that
provider's `realtime_voice` override accordingly; the per-provider setting
surface already supports it.
**Conclusion.** Both verdicts are `hold-floor-ok`, so `promotion_enabled`
defaults **on**. Phase 0's gate is satisfied; default-on is no longer blocked.
+41
View File
@@ -591,6 +591,47 @@ times, currency, percentages, versions, measurements, counts, paths, URLs, IDs,
JSON, logs, stack traces, tables, and dense numeric strings should be spoken as
human-readable summaries rather than raw character-by-character dumps.
#### Background Hermes runs (ADR 33)
Short Hermes runs answer in-line (the provider speaks the summary as soon as the
brokered result returns). A run that exceeds `promote_after_ms` is **promoted**
to a tracked background task so the provider event pump stays responsive instead
of blocking on the run:
```json
{"type":"hermes.run.promoted","event_id":20,"source":"hermes","run_id":"...","tier":"promoted","promote_after_ms":6000,"spoken_handoff":true,"result_delivery":"speak_when_idle","call_id":"call-1"}
{"type":"hermes.run.background_completed","event_id":41,"source":"hermes","run_id":"...","ok":true,"tool_count":2}
```
- On promotion the relay closes the pending provider function call with an
interim `{"status":"running_in_background"}` output (so the provider socket is
not left awaiting a tool result) and, when `spoken_handoff` is on, has the
provider speak a brief "I'm on it" line.
- When the background run finishes, the relay emits
`hermes.run.background_completed`, waits for the audio **floor** to be idle,
then injects the result through the same forced-summary path so the provider
speaks the answer exactly once. `result_delivery` selects `speak_when_idle`
(default), `notify_then_speak`, or `visual_only`.
- `hermes_run_task(mode="background")` skips the grace window and detaches
immediately (`tier:"durable"`), even when grace-period promotion is disabled.
- `hermes.run.progress` carries two extra fields while a run is in flight:
`tier` (`foreground` | `promoted` | `durable`) and `floor`
(`idle` | `provider_speaking` | `hermes_filler` | `result_pending`). The relay
is the single floor owner — a completed background result never barges in, and
Android local filler is suppressed while the provider holds the floor.
- A promoted run is cancellable via `response.cancel` (client) or the
`hermes_cancel` provider tool; if the WebSocket detaches mid-run, the
completion events replay through the resume event ring.
Promotion is configured per profile under `realtime_voice` (relay) and exposed
on `GET/PATCH /voice/realtime-agent/config` as a `promotion` block:
`enabled` (default on), `promote_after_ms`, `background_default_mode`,
`spoken_handoff`, `progress_spoken_after_ms`, `progress_repeat_ms`,
`result_delivery`, and `max_background_runs`. The default-on path is safe because
it closes the pending call rather than holding an open provider response; the
`scripts/realtime-provider-idle-probe.py` verdict (see `docs/realtime-voice-poc.md`)
confirms per-provider socket survival across the between-turns idle gap.
If a provider-native realtime turn answers directly without Hermes, Android
keeps the local user/assistant bubbles and marks that provider-only assistant
turn as unsynced. The next normal Hermes chat/run request includes those
+4 -2
View File
@@ -293,8 +293,10 @@ which probes `hermes gateway run --help | grep tailscale`. When
that returns true (PR #9295 has landed in your hermes-agent install),
the helper still works but the canonical path
(`hermes gateway run --tailscale`) is preferred and the helper will
be removed in a future release. Same retirement pattern as
`hermes_relay_bootstrap/` after PR #8556.
be removed in a future release. Treat helper retirement like
`hermes_relay_bootstrap/`: remove it only after the native path has parity
for every Relay-consuming route/group, not merely after one upstream
baseline lands.
### Forward-auth gateways (Authelia, Cloudflare Access) in front of the API server
+3 -3
View File
@@ -361,7 +361,7 @@ Bottom navigation bar with 4 tabs:
3. Remaining top-bar actions (session drawer hamburger, ambient toggle, etc.).
- **Session drawer** (swipe from left or hamburger icon) — session list with title, timestamp, message count. Create, switch, rename, delete.
- **Chat view** — message bubbles with markdown rendering, streaming text, tool call cards (Off/Compact/Detailed display modes)
- **Input bar** — text field with 4096 char limit, `/` palette button, send button, stop button during streaming. Inline autocomplete on `/` keystroke + full searchable command palette (bottom sheet). Commands sourced from: 29 gateway built-ins, dynamic personalities from `config.agent.personalities`, and server skills from `GET /api/skills`.
- **Input bar** — text field with 4096 char limit, `/` palette button, send button, stop button during streaming. Inline autocomplete on `/` keystroke + full searchable command palette (bottom sheet). Commands sourced from: 29 gateway built-ins, dynamic personalities from `config.agent.personalities`, and server skills from `GET /v1/skills` with legacy `/api/skills` fallback.
- **Empty state** — Logo + "Start a conversation" + suggestion chips that populate input
- **Agent sheet — Profile section (v0.6.0, updated 2026-05-18)** — upstream Hermes profiles auto-discovered by the relay at `~/.hermes/profiles/*/`. Selecting one routes chat/session calls to that profile's advertised `api_server_url` when present, giving proper Hermes isolation for sessions, memory, tools, provider auth, and SOUL/default model. If no profile API route is advertised, the app falls back to overlaying `model` + `SOUL.md` (as `system_message`) on the active Connection's API server. Selection is persisted per Connection/profile context. Hidden when the server advertises no profiles. See `docs/decisions.md` §21.
- **Agent sheet — Personality section** — personalities fetched from `GET /api/config` (`config.agent.personalities`). Shows server default (from `config.display.personality`) + all configured. Active personality name shown on assistant chat bubbles.
@@ -499,7 +499,7 @@ Additional API endpoints used:
5. DELETE /api/sessions/{session_id} → delete session
6. GET /api/sessions/{session_id}/messages → fetch message history
7. GET /api/config → personalities (for personality picker, `config.agent.personalities`)
8. GET /api/skills → available skills (for command palette + autocomplete)
8. GET /v1/skills → available skills (for command palette + autocomplete; fallback GET /api/skills)
```
Key classes:
@@ -888,7 +888,7 @@ Current versions as of v0.3.0. Source of truth is `gradle/libs.versions.toml`
| **WebAPI chat** | HTTP to `localhost:8642/api/sessions/*/chat/stream` (SSE) |
| **WebAPI sessions** | `GET/POST/PATCH/DELETE /api/sessions` for CRUD |
| **Personalities** | `GET /api/config` → `config.agent.personalities` for picker + command palette |
| **Server skills** | `GET /api/skills` — dynamic skill discovery for command palette + autocomplete |
| **Server skills** | `GET /v1/skills` first, fallback `GET /api/skills` — dynamic skill discovery for command palette + autocomplete |
| **Plugin system** | `register_tool()` via `ctx` for `android_*` tools |
| **Gateway** | Chat channel goes through WebAPI, not directly to gateway |
| **Memory/Skills** | Accessible through agent chat (no direct API needed for MVP) |
+8 -6
View File
@@ -2,9 +2,11 @@
Improvements that would benefit hermes-relay (and other frontends) if added to [NousResearch/hermes-agent](https://github.com/NousResearch/hermes-agent).
## Current Upstream PR Alignment
## Current Upstream Alignment
- PR #29302 (`feat: add API server session controls`) is the canonical upstream path for `/api/sessions/*`, message history, fork, chat, and chat stream. Hermes-Relay should prefer these native routes when present and keep the bootstrap as a per-route compatibility overlay only for older or partial core builds.
- PR #33134 (`feat(api-server): session control API — sessions/chat/fork/SSE-stream (salvages #29302)`) landed upstream as commit [`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde). It is now the canonical baseline for `/api/sessions/*`, message history, fork, chat, and chat stream.
- `GET /v1/skills` and `GET /v1/toolsets` are native upstream capability-discovery routes. Hermes-Relay's Android client prefers `/v1/skills` first and falls back to legacy `/api/skills` for older/fork/bootstrap installs.
- Keep `hermes_relay_bootstrap` as a per-route compatibility overlay only for older or partial core builds. Do not remove it wholesale while Relay still needs routes without native upstream parity: `/api/sessions/search`, `/api/memory`, `/api/config`, legacy skill detail routes under `/api/skills`, `/api/available-models`, and voice aliases.
- PR #8199 (`feat(api): add native audio transcription and speech endpoints`) is the canonical upstream path for core STT/TTS execution through `/v1/audio/transcriptions` and `/v1/audio/speech`. Hermes-Relay should keep `/voice/*` as the paired-device facade but eventually call those native core endpoints internally before falling back to private helper imports.
- PR #29364 (`feat: add API server audio endpoints`) should not become a competing `/api/audio/*` API if #8199 remains the accepted audio base. Rework it as a discovery/compatibility follow-up or close it after confirming the upstream maintainer preference.
@@ -38,7 +40,7 @@ Improvements that would benefit hermes-relay (and other frontends) if added to [
**Impact:** All frontends (hermes-relay, hermes-workspace, ClawPort) could dynamically show available commands without hardcoding. New commands added upstream would appear automatically.
**Workaround (current):** 29 gateway commands hardcoded in `ChatScreen.kt`, manually synced with `hermes_cli/commands.py`. Personality commands generated from `GET /api/config`. Skills from `GET /api/skills`.
**Workaround (current):** 29 gateway commands hardcoded in `ChatScreen.kt`, manually synced with `hermes_cli/commands.py`. Personality commands generated from `GET /api/config`. Skills prefer native `GET /v1/skills` and fall back to legacy `GET /api/skills`.
## 2. Personality Switching via Dedicated API Parameter
@@ -103,16 +105,16 @@ except Exception as _exc:
**Proposed — a two-stage arc, each stage a small, independently reviewable PR:**
**Stage 1 — stateless preprocessor (sibling follow-up to PR #29302).** A lightweight preprocessor in `api_server.py`'s `/v1/runs` + `/v1/chat/completions` handlers that detects a leading `/` in the user text, matches the first token against `GATEWAY_KNOWN_COMMANDS`, and splits on command type:
**Stage 1 — stateless preprocessor (follow-up to PR #33134 / commit `f7527b0`).** A lightweight preprocessor in `api_server.py`'s `/v1/runs` + `/v1/chat/completions` handlers that detects a leading `/` in the user text, matches the first token against `GATEWAY_KNOWN_COMMANDS`, and splits on command type:
- **Stateless commands** (`/help`, `/commands`, and any others that can execute without touching router-owned state) are dispatched via existing helpers (`gateway_help_lines()` at `hermes_cli/commands.py:340`) and returned as a synthetic SSE stream matching the handlers' existing event shape.
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, and most of the registry) return a deterministic, helpful SSE notice along the lines of *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` (from PR #29302) or a channel with session state (Discord, CLI, Telegram)."*
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, and most of the registry) return a deterministic, helpful SSE notice along the lines of *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` or a channel with session state (Discord, CLI, Telegram)."*
- **Unknown** and **cli-only** commands fall through to the LLM path unchanged.
- **Preprocessor exceptions** fall through to the LLM path unchanged — a preprocessor bug must never take down a normal chat request.
This respects upstream's intentional design (api_server stays stateless, no router coupling) while fixing the hallucination symptom and unlocking the commands that *can* run statelessly.
**Stage 2 — stateful dispatch on `/api/sessions/{id}/chat/stream` (after PR #29302 lands).** Once session management primitives ship, a separate PR adds a preprocessor **scoped to the session chat stream endpoint only**, using the URL's `session_id` as the persistence handle. Stateful commands become session-scoped dict writes (`session.model_override = new_model`) without refactoring `GatewayRouter` or plumbing api_server into the router. This matches upstream's partition cleanly: `/v1/*` remains stateless and OpenAI-compatible; statefulness lives on `/api/sessions/*`.
**Stage 2 — stateful dispatch on `/api/sessions/{id}/chat/stream` (now unblocked).** Now that session management primitives have landed upstream, a separate PR can add a preprocessor **scoped to the session chat stream endpoint only**, using the URL's `session_id` as the persistence handle. Stateful commands become session-scoped dict writes (`session.model_override = new_model`) without refactoring `GatewayRouter` or plumbing api_server into the router. This matches upstream's partition cleanly: `/v1/*` remains stateless and OpenAI-compatible; statefulness lives on `/api/sessions/*`.
**Why not one big PR:** a full GatewayRouter refactor plus api_server plumbing was considered and rejected. It would touch 10+ files across subsystems normally owned separately, fight the documented "api_server is excluded from router notification" design decision, and review as a much larger change than the value added. The two-stage arc ships faster, reviews cleaner, and matches the upstream partition better.
+5 -5
View File
@@ -1,6 +1,6 @@
# Upstream Hermes Integration Sync
Last reviewed: 2026-05-20
Last reviewed: 2026-06-05
This document tracks how Hermes-Relay integrates with Hermes upstream surfaces, which
parts use supported extension points, and which parts are compatibility layers that
@@ -37,8 +37,8 @@ relay, dashboard, Android app, desktop app, bootstrap package, or user docs.
| Agent tools | Tool Gateway tools registered through plugin context | `ctx.register_tool(...)` in `plugin/__init__.py`; schemas and handlers in `plugin/tools/*` | Aligned with custom transports | Tool registration should stay in `register(ctx)`; transport details stay behind handlers. |
| Dashboard tab and plugin API | Dashboard plugin manifest plus plugin API routes under the Hermes dashboard plugin mount | `plugin/dashboard/manifest.json`, `plugin/dashboard/plugin_api.py` | Aligned wrapper | Dashboard routes may proxy relay state, but discovery and mounting should stay upstream-native. |
| Chat and model API | OpenAI-compatible API server routes such as `/v1/chat/completions`, `/v1/models`, `/health`, and supported streaming routes | Android `HermesApiClient`, relay docs, Web API docs | Mixed | Prefer standard API routes first; use `/api/sessions` only when capability probes find it. |
| Sessions API | Proposed upstream API-server session controls in NousResearch/hermes-agent PR #29302 (`/api/sessions`, messages, fork, chat, chat stream) | Android `HermesApiClient`; compatibility overlay in `hermes_relay_bootstrap/*` | Upstream-pending with fallback | Prefer native `/api/sessions/*` when present. Bootstrap must skip native routes per method/path and only inject missing compatibility routes. |
| Config, skills, memory APIs | Not documented as stable upstream API-server routes in current public docs | `hermes_relay_bootstrap/*`, `docs/HERMES-WEBAPI-REFERENCE.md` | Compatibility layer | Keep separate from the sessions retirement path. Do not skip these just because native `/api/sessions` exists. |
| Sessions API | Native upstream API-server session controls from PR #33134 / commit `f7527b0` (`/api/sessions`, messages, fork, chat, chat stream) | Android `HermesApiClient`; compatibility overlay in `hermes_relay_bootstrap/*` for older cores | Upstream baseline with fallback | Prefer native `/api/sessions/*` when present. Bootstrap must skip native routes per method/path and only inject missing compatibility routes. |
| Config, skills, memory APIs | Native upstream provides `/v1/skills` and `/v1/toolsets`; config, memory, legacy skill detail routes, search, and available-models are not stable upstream API-server routes | `hermes_relay_bootstrap/*`, `docs/HERMES-WEBAPI-REFERENCE.md` | Mixed: native skill list + compatibility layer | Keep separate from the sessions retirement path. Do not skip these just because native `/api/sessions` or `/v1/skills` exists. |
| Mobile, desktop, and terminal relay transport | No general upstream plugin WSS transport for persistent remote clients in current public docs | `plugin/relay/server.py`, `plugin/relay/channels/*` | Custom | Keep the relay protocol documented and avoid leaking relay-only assumptions into upstream API clients. |
| Pairing QR and relay session minting | No upstream pairing or device-registration method for remote mobile clients in current public docs | `plugin/pair.py`, relay `/pairing/*`, Android QR parser | Custom | QR payloads should keep API credentials (`key`) separate from relay credentials (`relay.code`). |
| Basic STT/TTS over HTTP | Proposed upstream API-server audio endpoints in PR #8199 (`/v1/audio/transcriptions`, `/v1/audio/speech`) | Relay `/voice/config`, `/voice/transcribe`, `/voice/synthesize`; Android `RelayVoiceClient`; `plugin/relay/upstream_voice.py` | Custom wrapper pending upstream replacement | Keep `/voice/*` as the relay auth/session compatibility facade. Once core audio endpoints land, prefer proxying to native `/v1/audio/*` for STT/TTS work before falling back to private helper imports. |
@@ -50,7 +50,7 @@ relay, dashboard, Android app, desktop app, bootstrap package, or user docs.
| Deviation | Owner files | Why it exists | Guard or fallback | Retirement condition |
| --- | --- | --- | --- | --- |
| API bootstrap route and middleware injection | `hermes_relay_bootstrap/*` | Native installs need session/config/skills/memory endpoints and slash-command preprocessing before upstream exposes stable equivalents. | Method/path feature detection skips native upstream routes and injects only missing compatibility gaps; upstream-module checks skip middleware when native slash preprocessing exists. | Retire per surface: sessions after PR #29302 or equivalent ships in a released core; config/skills/memory after stable core APIs exist; slash middleware after native preprocessing exists. |
| API bootstrap route and middleware injection | `hermes_relay_bootstrap/*` | Native/older installs need compatibility routes and slash-command preprocessing before upstream exposes stable equivalents. | Method/path feature detection skips native upstream routes and injects only missing compatibility gaps; upstream-module checks skip middleware when native slash preprocessing exists. | Retire per surface: sessions after the supported baseline includes PR #33134/`f7527b0` and tests prove no Relay session route gaps remain; config/legacy skills/memory/search/available-models after stable core APIs or client migrations exist; slash middleware after native preprocessing exists. |
| Plugin CLI shim fallback | `plugin/__init__.py`, `plugin/cli.py`, install scripts | Some Hermes versions do not wire third-party plugin CLI commands into the top-level parser. | `ctx.register_cli_command` is attempted first; standalone shims fill the gap. | Remove shims once upstream plugin CLI discovery is stable for native installs. |
| Relay HTTP and WSS server | `plugin/relay/server.py`, `plugin/relay/channels/*` | Mobile, desktop, terminal, media, push, and bridge features need persistent client channels and relay-owned session state. | Keep upstream API calls separate from relay session calls and document the protocol in `docs/relay-protocol.md`. | Replace pieces only when upstream provides equivalent remote-client transport or platform adapters. |
| Pairing schema with `relay.code` | `plugin/pair.py`, Android pairing parser, relay `/pairing/*` | An API bearer key authenticates Hermes API calls but does not create relay sessions or describe WSS endpoints. | QR payloads carry direct API credentials and relay credentials as separate families. | Remove custom pairing when upstream offers native remote-device registration and relay discovery. |
@@ -118,7 +118,7 @@ upgrading the supported Hermes baseline.
- `GET /api/sessions?limit=1` only as an enhanced-management capability probe
- Relay health and info endpoints from `docs/relay-server.md`
- Dashboard plugin overview under the Hermes plugin API mount
- `GET /v1/capabilities` and the native `/api/sessions/*` route set when testing a core build with PR #29302 or equivalent
- `GET /v1/capabilities`, `GET /v1/skills`, and the native `/api/sessions/*` route set when testing a core build with PR #33134 / commit `f7527b0` or equivalent
- `POST /v1/audio/transcriptions` and `POST /v1/audio/speech` when testing a core build with PR #8199 or equivalent
- Voice config, transcription, synthesis, and realtime routes only with relay session auth or a valid Hermes API bearer
6. Update this file when upstream adds a supported replacement for a custom layer.
+45
View File
@@ -108,6 +108,19 @@ class RelayConfig:
realtime_voice_config_path: str | None = None
realtime_voice_run_dir: str | None = None
realtime_voice_xai_oauth_path: str | None = None
# ADR 33 background-Hermes-run promotion. Default ON: the promotion path
# closes the pending provider call with an interim ack instead of holding an
# open response, so the provider socket only sees the normal between-turns
# idle gap, not a long open-response stall. The Phase 0 probe
# (docs/realtime-voice-poc.md) still confirms per-provider socket survival.
realtime_voice_promotion_enabled: bool = True
realtime_voice_promote_after_ms: int = 6000
realtime_voice_background_default_mode: str = "promote"
realtime_voice_spoken_handoff: bool = True
realtime_voice_progress_spoken_after_ms: int = 15000
realtime_voice_progress_repeat_ms: int = 30000
realtime_voice_result_delivery: str = "speak_when_idle"
realtime_voice_max_background_runs: int = 1
@classmethod
def from_env(cls) -> RelayConfig:
@@ -375,6 +388,38 @@ def _apply_realtime_voice_config(
if xai_oauth_path:
config.realtime_voice_xai_oauth_path = xai_oauth_path
promotion_enabled = _optional_bool(section.get("promotion_enabled"))
if promotion_enabled is not None:
config.realtime_voice_promotion_enabled = promotion_enabled
promote_after_ms = _optional_int(section.get("promote_after_ms"))
if promote_after_ms is not None:
config.realtime_voice_promote_after_ms = max(0, promote_after_ms)
background_default_mode = _string_value(section.get("background_default_mode"))
if background_default_mode in ("promote", "foreground"):
config.realtime_voice_background_default_mode = background_default_mode
spoken_handoff = _optional_bool(section.get("spoken_handoff"))
if spoken_handoff is not None:
config.realtime_voice_spoken_handoff = spoken_handoff
progress_spoken_after_ms = _optional_int(section.get("progress_spoken_after_ms"))
if progress_spoken_after_ms is not None:
config.realtime_voice_progress_spoken_after_ms = max(0, progress_spoken_after_ms)
progress_repeat_ms = _optional_int(section.get("progress_repeat_ms"))
if progress_repeat_ms is not None:
config.realtime_voice_progress_repeat_ms = max(0, progress_repeat_ms)
result_delivery = _string_value(section.get("result_delivery"))
if result_delivery in ("speak_when_idle", "notify_then_speak", "visual_only"):
config.realtime_voice_result_delivery = result_delivery
max_background_runs = _optional_int(section.get("max_background_runs"))
if max_background_runs is not None:
config.realtime_voice_max_background_runs = max(1, max_background_runs)
def _apply_voice_output_config(
config: RelayConfig,
+32
View File
@@ -162,6 +162,29 @@ def realtime_voice_settings(config: Any, profile: str | None) -> dict[str, Any]:
"model": getattr(config, "realtime_voice_model", "grok-voice-latest"),
"voice": getattr(config, "realtime_voice_voice", "eve"),
"sample_rate": getattr(config, "realtime_voice_sample_rate", 24000),
# ADR 33 background-Hermes-run promotion.
"promotion_enabled": bool(
getattr(config, "realtime_voice_promotion_enabled", False)
),
"promote_after_ms": int(
getattr(config, "realtime_voice_promote_after_ms", 6000)
),
"background_default_mode": getattr(
config, "realtime_voice_background_default_mode", "promote"
),
"spoken_handoff": bool(getattr(config, "realtime_voice_spoken_handoff", True)),
"progress_spoken_after_ms": int(
getattr(config, "realtime_voice_progress_spoken_after_ms", 15000)
),
"progress_repeat_ms": int(
getattr(config, "realtime_voice_progress_repeat_ms", 30000)
),
"result_delivery": getattr(
config, "realtime_voice_result_delivery", "speak_when_idle"
),
"max_background_runs": int(
getattr(config, "realtime_voice_max_background_runs", 1)
),
}
applied_scope = "relay"
@@ -269,6 +292,15 @@ def _realtime_voice_overrides(data: dict[str, Any]) -> dict[str, Any]:
if sample_rate is not None:
out["sample_rate"] = sample_rate
_copy_bool(out, section, "enabled")
# ADR 33 promotion overrides.
_copy_bool(out, section, "promotion_enabled")
_copy_int(out, section, "promote_after_ms")
_copy_string(out, section, "background_default_mode")
_copy_bool(out, section, "spoken_handoff")
_copy_int(out, section, "progress_spoken_after_ms")
_copy_int(out, section, "progress_repeat_ms")
_copy_string(out, section, "result_delivery")
_copy_int(out, section, "max_background_runs")
return out
+394 -2
View File
@@ -50,6 +50,7 @@ from ..provider_options import (
)
from ..realtime_voice import _read_relay_xai_oauth_token, _websocket_url_from_base
from ..voice_auth import AuthPrincipal, require_voice_auth
from .floor import FloorMouth, RealtimeFloor
from .hermes_tool_broker import HermesTaskRequest, HermesToolBroker
from .models import (
CLIENT_MSG_HERMES_CONFIRM,
@@ -80,6 +81,8 @@ from .models import (
SERVER_EVT_RESPONSE_DONE,
SERVER_EVT_RESPONSE_STARTED,
SERVER_EVT_HERMES_RUN_PROGRESS,
SERVER_EVT_HERMES_RUN_PROMOTED,
SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED,
SERVER_EVT_SESSION_DETACHED,
SERVER_EVT_SESSION_READY,
SERVER_EVT_SESSION_RESUMED,
@@ -102,6 +105,9 @@ _HERMES_PROGRESS_INTERVAL_SECONDS = 5.0
_HERMES_SPOKEN_PROGRESS_AFTER_SECONDS = 15.0
_HERMES_SPOKEN_PROGRESS_REPEAT_SECONDS = 30.0
_RESUME_TTL_SECONDS = 30.0
# Max time a completed background result waits for the floor to clear before it
# is spoken anyway (ADR 33 Tier B result delivery).
_BACKGROUND_FLOOR_WAIT_SECONDS = 12.0
_EVENT_RING_LIMIT = 256
_AUDIO_RING_LIMIT = 96
_PROFILE_SOUL_PROMPT_MAX_CHARS = 6000
@@ -170,10 +176,18 @@ class RealtimeAgentSession:
native_hermes_required_reason: str | None = None
hermes_run_id: str | None = None
hermes_run_status: str = "idle"
hermes_run_tier: str = "foreground"
hermes_answer_started: bool = False
pending_confirmation_id: str | None = None
cancel_requested: bool = False
hermes_task: asyncio.Task[dict[str, Any]] | None = None
# ADR 33 promotion state (populated from realtime_voice settings at create).
promotion_enabled: bool = False
promote_after_ms: int = 6000
spoken_handoff: bool = True
result_delivery: str = "speak_when_idle"
promoted_transcript: str | None = None
background_delivery_task: asyncio.Task[None] | None = None
response_ids_awaiting_tool_followup: set[str] = field(default_factory=set)
response_ids_started: set[str] = field(default_factory=set)
provider_response_audio_seen: set[str] = field(default_factory=set)
@@ -187,6 +201,7 @@ class RealtimeAgentSession:
hermes_last_spoken_progress_at: float = 0.0
hermes_last_spoken_progress_key: str | None = None
profile_prompt_context: dict[str, Any] = field(default_factory=dict)
floor: RealtimeFloor = field(default_factory=RealtimeFloor)
class RealtimeAgentHandler:
@@ -329,6 +344,10 @@ class RealtimeAgentHandler:
self.config,
str(settings.get("profile") or profile or "default"),
),
promotion_enabled=bool(settings.get("promotion_enabled", False)),
promote_after_ms=int(settings.get("promote_after_ms", 6000)),
spoken_handoff=bool(settings.get("spoken_handoff", True)),
result_delivery=str(settings.get("result_delivery", "speak_when_idle")),
)
self.sessions[session_id] = session
self._log(session, "voice.realtime_agent.session.created")
@@ -740,6 +759,10 @@ class RealtimeAgentHandler:
provider_task = session.native_provider_task
if provider_task is not None and not provider_task.done():
provider_task.cancel()
delivery_task = session.background_delivery_task
if delivery_task is not None and not delivery_task.done():
delivery_task.cancel()
session.background_delivery_task = None
connection = session.native_connection
if connection is not None:
await connection.close()
@@ -1278,6 +1301,9 @@ class RealtimeAgentHandler:
},
)
elif event.kind == ProviderEventKind.AUDIO_DELTA:
# The provider is producing audio for this response; take the
# floor so filler/relay-TTS can't overlap (ADR 33).
session.floor.acquire(FloorMouth.PROVIDER)
if session.native_forced_preamble_active:
if not self._should_forward_forced_preamble_event(session, event):
continue
@@ -1297,6 +1323,9 @@ class RealtimeAgentHandler:
continue
await self._send_provider_audio_delta(ws, session, event)
elif event.kind == ProviderEventKind.AUDIO_DONE:
# Provider finished emitting audio for this response; release the
# floor so a pending background result / filler can proceed.
session.floor.release(FloorMouth.PROVIDER)
if session.native_forced_preamble_active:
if not self._should_forward_forced_preamble_event(session, event):
continue
@@ -1383,6 +1412,8 @@ class RealtimeAgentHandler:
event,
)
elif event.kind == ProviderEventKind.RESPONSE_DONE:
# Safety net: a response may end without a clean AUDIO_DONE.
session.floor.release(FloorMouth.PROVIDER)
if session.native_forced_preamble_active:
if not self._should_forward_forced_preamble_event(session, event):
continue
@@ -1667,6 +1698,16 @@ class RealtimeAgentHandler:
should_speak=False,
)
result = await self._run_brokered_tool(ws, session, call)
if result.get("promoted"):
await self._begin_background_delivery(
ws,
session,
connection,
origin="provider_tool_call",
call_id=call.call_id,
transcript=str(call.arguments.get("text") or ""),
)
return
delivered = await self._send_provider_tool_result(
ws,
session,
@@ -1848,6 +1889,16 @@ class RealtimeAgentHandler:
},
),
)
if result.get("promoted"):
await self._begin_background_delivery(
ws,
session,
connection,
origin="forced",
call_id=None,
transcript=transcript,
)
return
if result.get("cancelled"):
return
if result.get("ok") is False:
@@ -1938,8 +1989,39 @@ class RealtimeAgentHandler:
task = asyncio.create_task(self._execute_brokered_tool(ws, session, call))
session.hermes_task = task
# Tier C: an explicit mode="background" request detaches immediately,
# even when grace-period promotion is otherwise off.
force_background = str(call.arguments.get("mode") or "").strip().lower() == "background"
promote_after = 0.0 if force_background else self._promote_after_seconds(session)
try:
return await task
if promote_after is None:
return await task
try:
# Shield so a promotion timeout cancels only the wait, not the run.
return await asyncio.wait_for(asyncio.shield(task), timeout=promote_after)
except asyncio.TimeoutError:
# Tier B/C: detach the run to the background and hand control back
# to the pump (ADR 33). The task keeps running;
# _deliver_background_result awaits and delivers it.
session.hermes_run_tier = "durable" if force_background else "promoted"
session.promoted_transcript = str(call.arguments.get("text") or "").strip()
self._log(
session,
"voice.hermes_run.promoted",
{
"type": "voice.hermes_run.promoted",
"run_id": session.hermes_run_id,
"promote_after_ms": session.promote_after_ms,
"call_id": call.call_id,
},
)
return {
"ok": True,
"promoted": True,
"run_id": session.hermes_run_id,
"session_id": session.chat_session_id,
"interface": _interface_context(session),
}
except asyncio.CancelledError:
session.hermes_run_status = "cancelled"
return {
@@ -1949,10 +2031,238 @@ class RealtimeAgentHandler:
"session_id": session.chat_session_id,
"interface": _interface_context(session),
}
finally:
# A promoted task is still running; keep it referenced for delivery.
if session.hermes_task is task and task.done():
session.hermes_task = None
def _promote_after_seconds(self, session: RealtimeAgentSession) -> float | None:
"""Grace window before a run is promoted, or None when promotion is off."""
if not session.promotion_enabled:
return None
return max(0.0, session.promote_after_ms / 1000.0)
async def _begin_background_delivery(
self,
ws: web.WebSocketResponse,
session: RealtimeAgentSession,
connection: RealtimeAgentConnection,
*,
origin: str,
call_id: str | None,
transcript: str,
) -> None:
"""Hand a promoted run off to the background and notify the client.
The Hermes task keeps running (it is still referenced by
``session.hermes_task``); ``_deliver_background_result`` awaits it and
speaks the result when the floor is clear.
"""
session.promoted_transcript = transcript or session.promoted_transcript
await self._send(
ws,
session,
{
"type": SERVER_EVT_HERMES_RUN_PROMOTED,
"source": "hermes",
"session_id": session.chat_session_id,
"chat_session_id": session.chat_session_id,
"run_id": session.hermes_run_id,
"tier": session.hermes_run_tier if session.hermes_run_tier in ("promoted", "durable") else "promoted",
"promote_after_ms": session.promote_after_ms,
"spoken_handoff": session.spoken_handoff,
"result_delivery": session.result_delivery,
"call_id": call_id,
},
)
speak = session.spoken_handoff and session.result_delivery != "visual_only"
if origin == "provider_tool_call" and call_id:
# Close the pending provider function call so the socket is not left
# awaiting tool output for the whole background run.
interim = {
"ok": True,
"status": "running_in_background",
"run_id": session.hermes_run_id,
"instruction": (
"The task is running in the background. Briefly acknowledge "
"that you've started and will report back; do not answer yet."
),
"interface": _interface_context(session),
}
if await self._send_provider_tool_result(ws, session, connection, call_id, interim):
if speak:
await self._request_provider_response(ws, session, connection, call_id)
elif origin == "forced" and speak:
with contextlib.suppress(Exception):
await connection.send_text(_background_handoff_prompt(transcript))
if (
session.background_delivery_task is not None
and not session.background_delivery_task.done()
):
session.background_delivery_task.cancel()
session.background_delivery_task = asyncio.create_task(
self._deliver_background_result(ws, session, connection)
)
async def _deliver_background_result(
self,
ws: web.WebSocketResponse,
session: RealtimeAgentSession,
connection: RealtimeAgentConnection,
) -> None:
task = session.hermes_task
if task is None:
return
try:
result = await task
except asyncio.CancelledError:
session.hermes_run_status = "cancelled"
await self._send(
ws,
session,
{
"type": "hermes.run.cancelled",
"session_id": session.chat_session_id,
"run_id": session.hermes_run_id,
},
)
session.hermes_run_tier = "foreground"
return
except Exception as exc: # noqa: BLE001 - surface as a voice error
await self._send_error(
ws,
session,
f"background Hermes run failed: {exc.__class__.__name__}: {exc}",
provider=session.provider,
)
session.hermes_run_tier = "foreground"
return
finally:
if session.hermes_task is task:
session.hermes_task = None
await self._send(
ws,
session,
{
"type": SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED,
"source": "hermes",
"session_id": session.chat_session_id,
"chat_session_id": session.chat_session_id,
"run_id": session.hermes_run_id,
"ok": bool(result.get("ok", True)),
"tool_count": result.get("tool_count", session.hermes_completed_tool_count),
},
)
if result.get("cancelled"):
session.hermes_run_tier = "foreground"
return
if result.get("ok") is False:
await self._send_error(
ws,
session,
str(result.get("error") or "Hermes failed"),
provider=session.provider,
)
session.hermes_run_tier = "foreground"
return
if session.result_delivery == "visual_only":
await self._emit_background_text_only(ws, session, result)
session.hermes_run_tier = "foreground"
return
# speak_when_idle / notify_then_speak: wait for the floor to clear, then
# have the provider speak a natural summary via the forced-summary path.
session.floor.note_result_ready()
await self._await_floor_idle_for_result(session)
await self._inject_background_summary(ws, session, connection, result)
session.hermes_run_tier = "foreground"
async def _await_floor_idle_for_result(
self,
session: RealtimeAgentSession,
*,
timeout: float = _BACKGROUND_FLOOR_WAIT_SECONDS,
) -> bool:
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
if session.floor.consume_result_if_idle():
return True
await asyncio.sleep(0.05)
# Timed out waiting for the floor; deliver anyway.
session.floor.clear_result()
return False
async def _inject_background_summary(
self,
ws: web.WebSocketResponse,
session: RealtimeAgentSession,
connection: RealtimeAgentConnection,
result: dict[str, Any],
) -> None:
transcript = session.promoted_transcript or ""
session.native_forced_summary_active = True
session.native_forced_summary_done = False
session.native_forced_summary_response_id = None
session.native_forced_summary_result = dict(result)
session.native_forced_summary_buffer.clear()
session.native_forced_summary_text_parts.clear()
with contextlib.suppress(Exception):
await connection.cancel_response()
await connection.send_text(_forced_hermes_summary_prompt(transcript, result))
async def _emit_background_text_only(
self,
ws: web.WebSocketResponse,
session: RealtimeAgentSession,
result: dict[str, Any],
) -> None:
text = str(result.get("text") or result.get("answer") or result.get("summary") or "").strip()
response_id = f"background-visual-{session.event_seq + 1}"
await self._send(
ws,
session,
{
"type": SERVER_EVT_RESPONSE_STARTED,
"provider": session.provider,
"model": session.model,
"voice": session.voice,
"session_id": session.session_id,
"chat_session_id": session.chat_session_id,
"response_id": response_id,
"source": "hermes",
"delivery": "visual_only",
},
)
if text:
await self._send(
ws,
session,
{
"type": SERVER_EVT_RESPONSE_DELTA,
"source": "hermes",
"delta": text,
"response_id": response_id,
},
)
await self._send(
ws,
session,
{
"type": SERVER_EVT_RESPONSE_DONE,
"provider": session.provider,
"model": session.model,
"voice": session.voice,
"event_log_path": str(session.event_log_path),
"chat_session_id": session.chat_session_id,
"response_id": response_id,
"delivery": "visual_only",
},
)
async def _execute_brokered_tool(
self,
ws: web.WebSocketResponse,
@@ -2263,6 +2573,7 @@ class RealtimeAgentHandler:
should_speak = (
speakable_progress
and elapsed_seconds >= _HERMES_SPOKEN_PROGRESS_AFTER_SECONDS
and session.floor.can_speak(FloorMouth.ANDROID_FILLER)
and (
session.hermes_last_spoken_progress_key != status_key
or now - session.hermes_last_spoken_progress_at
@@ -2285,6 +2596,8 @@ class RealtimeAgentHandler:
"message": message,
"status_key": status_key,
"should_speak": should_speak,
"floor": session.floor.state_label(),
"tier": session.hermes_run_tier,
"active_tool_name": session.hermes_active_tool_name,
"last_tool_name": session.hermes_last_tool_name,
"completed_tool_count": session.hermes_completed_tool_count,
@@ -2336,6 +2649,16 @@ class RealtimeAgentHandler:
"config_scope": settings["config_scope"],
"fallback_to_global": settings["fallback_to_global"],
"tool_surface": list(_TOOL_SURFACE),
"promotion": {
"enabled": bool(settings["promotion_enabled"]),
"promote_after_ms": int(settings["promote_after_ms"]),
"background_default_mode": settings["background_default_mode"],
"spoken_handoff": bool(settings["spoken_handoff"]),
"progress_spoken_after_ms": int(settings["progress_spoken_after_ms"]),
"progress_repeat_ms": int(settings["progress_repeat_ms"]),
"result_delivery": settings["result_delivery"],
"max_background_runs": int(settings["max_background_runs"]),
},
"limits": [
"Hermes owns tools, confirmations, memory, and transcript state.",
"Native realtime providers receive relay-brokered mic PCM and only the approved Hermes function surface.",
@@ -2519,6 +2842,9 @@ class RealtimeAgentHandler:
loop = asyncio.get_running_loop()
output_path = session.event_log_path.with_suffix(".wav")
provider_options = self._provider_options(session, payload)
# Relay TTS is a primary mouth; take the floor for the whole render so it
# cannot overlap provider audio or filler (ADR 33).
session.floor.acquire(FloorMouth.RELAY_TTS)
def audio_sink(chunk: bytes, meta: dict[str, Any]) -> None:
peak, rms = _pcm_levels(chunk)
@@ -2565,9 +2891,11 @@ class RealtimeAgentHandler:
await self._send(ws, session, event)
response = await task
except (ProviderUnavailable, ProviderRunError) as exc:
session.floor.release(FloorMouth.RELAY_TTS)
await self._send_error(ws, session, str(exc), provider=session.provider)
return
except Exception as exc:
session.floor.release(FloorMouth.RELAY_TTS)
await self._send_error(
ws,
session,
@@ -2601,6 +2929,7 @@ class RealtimeAgentHandler:
"metadata": _safe_metadata(response.metadata),
},
)
session.floor.release(FloorMouth.RELAY_TTS)
def _provider_options(
self,
@@ -2647,7 +2976,22 @@ class RealtimeAgentHandler:
def _validate_config_updates(self, payload: dict[str, Any]) -> dict[str, Any]:
updates: dict[str, Any] = {}
allowed = {"enabled", "provider", "model", "voice", "sample_rate"}
allowed = {
"enabled",
"provider",
"model",
"voice",
"sample_rate",
# ADR 33 promotion fields.
"promotion_enabled",
"promote_after_ms",
"background_default_mode",
"spoken_handoff",
"progress_spoken_after_ms",
"progress_repeat_ms",
"result_delivery",
"max_background_runs",
}
unsupported = sorted(set(payload) - allowed)
if unsupported:
raise web.HTTPBadRequest(
@@ -2671,6 +3015,43 @@ class RealtimeAgentHandler:
if sample_rate is None or sample_rate < 8_000 or sample_rate > 96_000:
raise web.HTTPBadRequest(text="sample_rate must be between 8000 and 96000")
updates["sample_rate"] = sample_rate
if "promotion_enabled" in payload:
value = _bool_value(payload["promotion_enabled"])
if value is None:
raise web.HTTPBadRequest(text="promotion_enabled must be a boolean")
updates["promotion_enabled"] = value
if "spoken_handoff" in payload:
value = _bool_value(payload["spoken_handoff"])
if value is None:
raise web.HTTPBadRequest(text="spoken_handoff must be a boolean")
updates["spoken_handoff"] = value
for ms_field, lo, hi in (
("promote_after_ms", 0, 120_000),
("progress_spoken_after_ms", 0, 600_000),
("progress_repeat_ms", 0, 600_000),
):
if ms_field in payload:
value = _int_value(payload[ms_field])
if value is None or value < lo or value > hi:
raise web.HTTPBadRequest(text=f"{ms_field} must be between {lo} and {hi}")
updates[ms_field] = value
if "max_background_runs" in payload:
value = _int_value(payload["max_background_runs"])
if value is None or value < 1 or value > 4:
raise web.HTTPBadRequest(text="max_background_runs must be between 1 and 4")
updates["max_background_runs"] = value
if "background_default_mode" in payload:
mode = _bounded_string(payload["background_default_mode"], "background_default_mode", max_len=20)
if mode not in ("promote", "foreground"):
raise web.HTTPBadRequest(text="background_default_mode must be 'promote' or 'foreground'")
updates["background_default_mode"] = mode
if "result_delivery" in payload:
delivery = _bounded_string(payload["result_delivery"], "result_delivery", max_len=24)
if delivery not in ("speak_when_idle", "notify_then_speak", "visual_only"):
raise web.HTTPBadRequest(
text="result_delivery must be speak_when_idle, notify_then_speak, or visual_only"
)
updates["result_delivery"] = delivery
if not updates:
raise web.HTTPBadRequest(text="no realtime agent config fields supplied")
return updates
@@ -3269,6 +3650,17 @@ def _forced_hermes_preamble_prompt(transcript: str) -> str:
)
def _background_handoff_prompt(transcript: str) -> str:
return (
"The user's request is now running in the background and may take a "
"little while. Speak one short, natural acknowledgement that you've "
"started on it and will report back, for example: 'I'm on it — I'll let "
"you know.' Do not call tools, do not answer the request yet, and do not "
"ask for a run id.\n\n"
f"User request: {transcript.strip()[:1000]}"
)
def _forced_hermes_summary_prompt(transcript: str, result: dict[str, Any]) -> str:
answer = str(
result.get("answer")
+130
View File
@@ -0,0 +1,130 @@
"""Relay-side audio floor owner for realtime-agent sessions (ADR 33).
Up to three sources can produce ``voice.output_audio.delta`` on Android's single
``AudioTrack``:
1. the realtime provider (xAI / OpenAI),
2. the relay TTS fallback render (``_render_provider_audio``), and
3. Android local filler, driven by the ``should_speak`` hint on
``hermes.run.progress``.
Today a blocking ``await`` inside the provider event pump serializes them so they
never overlap. ADR 33 backgrounds long Hermes runs, which removes that implicit
mutex. ``RealtimeFloor`` replaces it with an explicit, single-owner floor so a
completed background result never barges in and two mouths never speak at once.
The floor is pure and synchronous — every mutation happens inside the session's
asyncio task, so no locking is required. Keep it free of I/O so it stays
unit-testable in isolation (``plugin/tests/test_realtime_floor.py``).
"""
from __future__ import annotations
from enum import Enum
class FloorMouth(str, Enum):
"""A source that can emit audio to Android."""
PROVIDER = "provider"
RELAY_TTS = "relay_tts"
ANDROID_FILLER = "android_filler"
#: Mouths that emit primary speech as ``voice.output_audio.delta``. Only one of
#: these may hold the floor at a time, and either preempts filler.
PRIMARY_MOUTHS: frozenset[FloorMouth] = frozenset(
{FloorMouth.PROVIDER, FloorMouth.RELAY_TTS}
)
#: Wire labels for the ``floor`` field on ``hermes.run.progress`` (ADR 33).
FLOOR_IDLE = "idle"
FLOOR_PROVIDER_SPEAKING = "provider_speaking"
FLOOR_HERMES_FILLER = "hermes_filler"
FLOOR_RESULT_PENDING = "result_pending"
class RealtimeFloor:
"""Single-owner audio floor for one realtime-agent session."""
__slots__ = ("_holder", "_result_pending")
def __init__(self) -> None:
self._holder: FloorMouth | None = None
self._result_pending = False
@property
def holder(self) -> FloorMouth | None:
return self._holder
@property
def result_pending(self) -> bool:
return self._result_pending
def can_speak(self, mouth: FloorMouth) -> bool:
"""Whether ``mouth`` may begin emitting audio now.
- A primary mouth may speak when the floor is free, when it already holds
it, or by preempting Android filler — but not while the *other* primary
mouth holds it.
- Android filler may speak only when the floor is free or it already
holds it; any primary mouth suppresses it.
"""
if mouth in PRIMARY_MOUTHS:
return self._holder in (None, mouth, FloorMouth.ANDROID_FILLER)
return self._holder in (None, FloorMouth.ANDROID_FILLER)
def acquire(self, mouth: FloorMouth) -> bool:
"""Take the floor for ``mouth`` if allowed. Returns success.
A primary mouth acquiring while filler holds preempts the filler.
"""
if not self.can_speak(mouth):
return False
self._holder = mouth
return True
def release(self, mouth: FloorMouth) -> None:
"""Release the floor if ``mouth`` currently holds it (idempotent)."""
if self._holder == mouth:
self._holder = None
def note_result_ready(self) -> None:
"""Mark that a background Hermes result is ready to be spoken."""
self._result_pending = True
def clear_result(self) -> None:
self._result_pending = False
def consume_result_if_idle(self) -> bool:
"""If a result is pending and the floor is free, claim it for delivery.
Returns True exactly once per pending result, when it is safe to speak
(the floor is idle). The caller is then responsible for acquiring the
floor for whichever mouth speaks the summary.
"""
if self._holder is None and self._result_pending:
self._result_pending = False
return True
return False
def state_label(self) -> str:
"""Total snapshot label for the wire ``floor`` field."""
if self._holder in PRIMARY_MOUTHS:
return FLOOR_PROVIDER_SPEAKING
if self._holder == FloorMouth.ANDROID_FILLER:
return FLOOR_HERMES_FILLER
if self._result_pending:
return FLOOR_RESULT_PENDING
return FLOOR_IDLE
__all__ = [
"FLOOR_HERMES_FILLER",
"FLOOR_IDLE",
"FLOOR_PROVIDER_SPEAKING",
"FLOOR_RESULT_PENDING",
"FloorMouth",
"PRIMARY_MOUTHS",
"RealtimeFloor",
]
+14 -1
View File
@@ -43,7 +43,16 @@ HERMES_TOOL_SCHEMAS: tuple[dict[str, Any], ...] = (
"text": {"type": "string"},
"profile": {"type": "string"},
"session_id": {"type": "string"},
"mode": {"type": "string", "enum": ["chat", "run"]},
"mode": {
"type": "string",
"enum": ["chat", "run", "background"],
"description": (
"'run' (default) answers in-line for short tasks and is "
"auto-promoted to the background if it runs long; "
"'background' starts a durable run immediately and reports "
"back when it finishes."
),
},
},
"required": ["text"],
},
@@ -198,6 +207,8 @@ SERVER_EVT_OUTPUT_AUDIO_DONE = "voice.output_audio.done"
SERVER_EVT_PLAYBACK_DRAIN_REQUESTED = "voice.playback_drain.requested"
SERVER_EVT_HERMES_RUN_STARTED = "hermes.run.started"
SERVER_EVT_HERMES_RUN_PROGRESS = "hermes.run.progress"
SERVER_EVT_HERMES_RUN_PROMOTED = "hermes.run.promoted"
SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED = "hermes.run.background_completed"
SERVER_EVT_HERMES_TOOL_STARTED = "hermes.tool.started"
SERVER_EVT_HERMES_TOOL_DELTA = "hermes.tool.delta"
SERVER_EVT_HERMES_TOOL_COMPLETED = "hermes.tool.completed"
@@ -233,7 +244,9 @@ __all__ = [
"RealtimeAgentSession",
"RealtimeAgentSessionConfig",
"SERVER_EVT_HERMES_CONFIRMATION_REQUESTED",
"SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED",
"SERVER_EVT_HERMES_RUN_COMPLETED",
"SERVER_EVT_HERMES_RUN_PROMOTED",
"SERVER_EVT_HERMES_RUN_STARTED",
"SERVER_EVT_HERMES_TOOL_COMPLETED",
"SERVER_EVT_HERMES_TOOL_DELTA",
+98
View File
@@ -0,0 +1,98 @@
"""Invariant tests for the realtime-agent audio floor (ADR 33 Phase 1)."""
from __future__ import annotations
import unittest
from plugin.relay.realtime_agent.floor import (
FLOOR_HERMES_FILLER,
FLOOR_IDLE,
FLOOR_PROVIDER_SPEAKING,
FLOOR_RESULT_PENDING,
FloorMouth,
RealtimeFloor,
)
class RealtimeFloorTest(unittest.TestCase):
def setUp(self) -> None:
self.floor = RealtimeFloor()
def test_starts_idle_and_free(self) -> None:
self.assertIsNone(self.floor.holder)
self.assertEqual(self.floor.state_label(), FLOOR_IDLE)
self.assertTrue(self.floor.can_speak(FloorMouth.PROVIDER))
self.assertTrue(self.floor.can_speak(FloorMouth.RELAY_TTS))
self.assertTrue(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
def test_filler_suppressed_while_provider_speaks(self) -> None:
# Invariant: filler-suppressed-while-provider-speaks.
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
self.assertFalse(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
self.assertEqual(self.floor.state_label(), FLOOR_PROVIDER_SPEAKING)
def test_relay_tts_only_when_provider_not_holding(self) -> None:
# Invariant: relay-TTS-only-when-owned — two primary mouths never overlap.
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
self.assertFalse(self.floor.can_speak(FloorMouth.RELAY_TTS))
self.assertFalse(self.floor.acquire(FloorMouth.RELAY_TTS))
self.floor.release(FloorMouth.PROVIDER)
self.assertTrue(self.floor.can_speak(FloorMouth.RELAY_TTS))
self.assertTrue(self.floor.acquire(FloorMouth.RELAY_TTS))
# Now the provider must not be able to barge into the relay's render.
self.assertFalse(self.floor.can_speak(FloorMouth.PROVIDER))
def test_background_result_never_barges(self) -> None:
# Invariant: background-result-never-barges.
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
self.floor.note_result_ready()
self.assertTrue(self.floor.result_pending)
# Provider still holds the floor -> result must wait.
self.assertFalse(self.floor.consume_result_if_idle())
self.assertEqual(self.floor.state_label(), FLOOR_PROVIDER_SPEAKING)
self.floor.release(FloorMouth.PROVIDER)
self.assertEqual(self.floor.state_label(), FLOOR_RESULT_PENDING)
# Now idle -> result becomes deliverable exactly once.
self.assertTrue(self.floor.consume_result_if_idle())
self.assertFalse(self.floor.consume_result_if_idle())
self.assertFalse(self.floor.result_pending)
self.assertEqual(self.floor.state_label(), FLOOR_IDLE)
def test_provider_preempts_filler(self) -> None:
self.assertTrue(self.floor.acquire(FloorMouth.ANDROID_FILLER))
self.assertEqual(self.floor.state_label(), FLOOR_HERMES_FILLER)
# A primary mouth may preempt active filler.
self.assertTrue(self.floor.can_speak(FloorMouth.PROVIDER))
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
self.assertEqual(self.floor.holder, FloorMouth.PROVIDER)
self.assertEqual(self.floor.state_label(), FLOOR_PROVIDER_SPEAKING)
def test_release_is_idempotent_and_owner_scoped(self) -> None:
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
# A non-owner release must not free the floor.
self.floor.release(FloorMouth.RELAY_TTS)
self.assertEqual(self.floor.holder, FloorMouth.PROVIDER)
# Owner release frees it; repeated release is harmless.
self.floor.release(FloorMouth.PROVIDER)
self.floor.release(FloorMouth.PROVIDER)
self.assertIsNone(self.floor.holder)
def test_filler_allowed_when_idle_during_hermes_run(self) -> None:
# During a Hermes run the provider is not emitting audio (floor idle),
# so filler is allowed — this is what preserves today's behavior.
self.assertTrue(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
self.assertTrue(self.floor.acquire(FloorMouth.ANDROID_FILLER))
self.assertTrue(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
def test_result_pending_label_only_when_idle(self) -> None:
self.floor.note_result_ready()
self.assertEqual(self.floor.state_label(), FLOOR_RESULT_PENDING)
# While filler holds, the label reflects the active speaker, not pending.
self.floor.acquire(FloorMouth.ANDROID_FILLER)
self.assertEqual(self.floor.state_label(), FLOOR_HERMES_FILLER)
if __name__ == "__main__":
unittest.main()
+428
View File
@@ -0,0 +1,428 @@
"""Tier B grace-period promotion tests for the realtime agent (ADR 33 Phase 2).
These drive the provider-native realtime-agent route with a fake provider
connection and fake Hermes brokers of controllable latency, asserting:
- a run that exceeds promote_after_ms emits hermes.run.promoted, the pump keeps
processing provider events while the run is still in flight, and the result is
spoken exactly once after completion;
- a run that completes within the grace window is NOT promoted (Tier A);
- cancelling a promoted run stops the background task and emits run.cancelled;
- a background run that completes while detached replays
hermes.run.background_completed on resume.
"""
from __future__ import annotations
import asyncio
import json
import os
import tempfile
import unittest
from typing import Any
from aiohttp import WSMsgType, web
from aiohttp.test_utils import AioHTTPTestCase
from plugin.relay import provider_options, voice_auth
from plugin.relay.config import RelayConfig
from plugin.relay.realtime_agent import broker as broker_module
from plugin.relay.realtime_agent.hermes_tool_broker import HermesTaskRequest
from plugin.relay.realtime_agent.models import (
ProviderEvent,
ProviderEventKind,
ToolCallEvent,
)
from plugin.relay.server import create_app
class FakeNativeConnection:
def __init__(self) -> None:
self.text_inputs: list[str] = []
self.tool_results: list[tuple[str, dict[str, Any]]] = []
self.request_response_count = 0
self.cancelled = False
self.closed = False
self._events: asyncio.Queue[ProviderEvent | None] = asyncio.Queue()
async def send_audio(self, pcm: bytes, sample_rate: int) -> None:
pass
async def commit_audio(self) -> None:
pass
async def send_text(self, text: str) -> None:
self.text_inputs.append(text)
async def clear_audio(self) -> None:
pass
async def cancel_response(self) -> None:
self.cancelled = True
async def send_tool_result(self, call_id: str, output: dict[str, Any]) -> None:
self.tool_results.append((call_id, output))
async def request_response(self) -> None:
self.request_response_count += 1
async def close(self) -> None:
self.closed = True
await self._events.put(None)
async def emit(self, event: ProviderEvent) -> None:
await self._events.put(event)
async def events(self):
while True:
event = await self._events.get()
if event is None:
return
yield event
class FakeNativeProvider:
def __init__(self, provider_id: str = "xai_realtime") -> None:
self.provider_id = provider_id
self.connection = FakeNativeConnection()
self.configs: list[Any] = []
async def connect(self, config):
self.configs.append(config)
return self.connection
class GatedHermesToolBroker:
"""Blocks inside the run until released, then yields a final answer."""
def __init__(self) -> None:
self.requests: list[HermesTaskRequest] = []
self.started = asyncio.Event()
self.release = asyncio.Event()
self.cancelled = asyncio.Event()
async def stream_task(self, request: HermesTaskRequest):
self.requests.append(request)
session_id = request.session_id or "created-hermes-session"
yield {
"type": "hermes.run.started",
"session_id": session_id,
"run_id": "run-gated",
"profile": request.profile,
}
self.started.set()
try:
await self.release.wait()
except asyncio.CancelledError:
self.cancelled.set()
raise
yield {
"type": "voice.response.delta",
"session_id": session_id,
"run_id": "run-gated",
"delta": "Background answer ready.",
}
yield {
"type": "hermes.run.completed",
"session_id": session_id,
"run_id": "run-gated",
"profile": request.profile,
}
class FastHermesToolBroker:
"""Completes immediately (Tier A)."""
def __init__(self) -> None:
self.requests: list[HermesTaskRequest] = []
async def stream_task(self, request: HermesTaskRequest):
self.requests.append(request)
session_id = request.session_id or "created-hermes-session"
yield {
"type": "hermes.run.started",
"session_id": session_id,
"run_id": "run-fast",
"profile": request.profile,
}
yield {
"type": "voice.response.delta",
"session_id": session_id,
"run_id": "run-fast",
"delta": "Quick answer.",
}
yield {
"type": "hermes.run.completed",
"session_id": session_id,
"run_id": "run-fast",
"profile": request.profile,
}
class RealtimePromotionTests(AioHTTPTestCase):
async def get_application(self) -> web.Application:
self._tmpdir = tempfile.TemporaryDirectory()
self._previous_lead = broker_module._PRE_HERMES_STATUS_LEAD_SECONDS
broker_module._PRE_HERMES_STATUS_LEAD_SECONDS = 0.0
voice_auth._VALIDATION_CACHE.clear()
provider_options.clear_provider_option_cache()
config = RelayConfig(
realtime_voice_enabled=True,
realtime_voice_provider="stub",
realtime_voice_model="local-tone",
realtime_voice_voice="sine",
realtime_voice_promotion_enabled=True,
realtime_voice_promote_after_ms=50,
realtime_voice_config_path=os.path.join(self._tmpdir.name, "relay-config.yaml"),
realtime_voice_run_dir=self._tmpdir.name,
)
return create_app(config)
async def tearDownAsync(self) -> None:
await super().tearDownAsync()
broker_module._PRE_HERMES_STATUS_LEAD_SECONDS = getattr(
self, "_previous_lead", broker_module._PRE_HERMES_STATUS_LEAD_SECONDS
)
voice_auth._VALIDATION_CACHE.clear()
provider_options.clear_provider_option_cache()
tmpdir = getattr(self, "_tmpdir", None)
if tmpdir is not None:
tmpdir.cleanup()
def _server(self):
return self.app["server"]
async def _make_session(self) -> str:
session = self._server().sessions.create_session("test-phone", "test-id")
return session.token
@staticmethod
def _bearer(token: str) -> dict[str, str]:
return {"Authorization": f"Bearer {token}"}
async def _next_ws_event(self, ws) -> dict[str, Any]:
msg = await ws.receive(timeout=5)
self.assertEqual(msg.type, WSMsgType.TEXT, msg)
payload = json.loads(msg.data)
self.assertIsInstance(payload, dict)
return payload
async def _open(self, *, broker) -> tuple[Any, FakeNativeProvider, dict[str, Any]]:
token = await self._make_session()
fake_provider = FakeNativeProvider()
self._server().realtime_agent.hermes = broker
self._server().realtime_agent.native_providers["xai_realtime"] = fake_provider
resp = await self.client.post(
"/voice/realtime-agent/session",
json={
"provider": "xai_realtime",
"model": "grok-voice-latest",
"voice": "leo",
"chat_session_id": "chat-123",
},
headers=self._bearer(token),
)
self.assertEqual(resp.status, 200)
body = await resp.json()
ws = await self.client.ws_connect(body["websocket_path"], headers=self._bearer(token))
ready = await self._next_ws_event(ws)
self.assertEqual(ready["type"], "voice.session.ready")
body["_token"] = token
body["_ready"] = ready
return ws, fake_provider, body
async def _emit_tool_call(
self,
provider: FakeNativeProvider,
*,
call_id="call-1",
mode: str | None = None,
) -> None:
arguments: dict[str, Any] = {"text": "Research the thing.", "session_id": "chat-123"}
if mode is not None:
arguments["mode"] = mode
await provider.connection.emit(
ProviderEvent(
ProviderEventKind.FUNCTION_CALL_COMPLETED,
response_id="resp-1",
payload={
"call": ToolCallEvent(
call_id=call_id,
name="hermes_run_task",
arguments=arguments,
)
},
)
)
async def _read_until(self, ws, target: str, *, limit: int = 40) -> list[dict[str, Any]]:
events: list[dict[str, Any]] = []
for _ in range(limit):
event = await self._next_ws_event(ws)
events.append(event)
if event["type"] == target:
return events
raise AssertionError(f"did not see {target}; saw {[e['type'] for e in events]}")
async def test_long_run_promotes_and_pump_stays_responsive(self) -> None:
broker = GatedHermesToolBroker()
ws, provider, _ = await self._open(broker=broker)
try:
await self._emit_tool_call(provider)
events = await self._read_until(ws, "hermes.run.promoted")
promoted = next(e for e in events if e["type"] == "hermes.run.promoted")
self.assertEqual(promoted["tier"], "promoted")
self.assertEqual(promoted["promote_after_ms"], 50)
# The pending provider call was closed with an interim background ack.
self.assertTrue(provider.connection.tool_results)
self.assertEqual(
provider.connection.tool_results[0][1].get("status"),
"running_in_background",
)
# Prove the pump is NOT blocked: a provider event sent while the run
# is still gated is processed and forwarded.
self.assertFalse(broker.release.is_set())
await provider.connection.emit(
ProviderEvent(
ProviderEventKind.INPUT_TRANSCRIPT_DELTA,
payload={"delta": "still talking"},
)
)
live = await self._read_until(ws, "voice.input_transcript.delta")
self.assertTrue(any(e["type"] == "voice.input_transcript.delta" for e in live))
# Release the run -> it completes in the background and is spoken once.
broker.release.set()
done = await self._read_until(ws, "hermes.run.background_completed")
completed = next(e for e in done if e["type"] == "hermes.run.background_completed")
self.assertTrue(completed["ok"])
# The result is injected for the provider to summarize exactly once.
for _ in range(50):
summaries = [t for t in provider.connection.text_inputs if "final spoken answer" in t]
if summaries:
break
await asyncio.sleep(0.02)
self.assertEqual(len(summaries), 1)
finally:
broker.release.set()
await ws.close()
async def test_short_run_is_not_promoted(self) -> None:
broker = FastHermesToolBroker()
ws, provider, body = await self._open(broker=broker)
# Widen the grace window so the fast run finishes well within it.
self._server().realtime_agent.sessions[body["session_id"]].promote_after_ms = 5000
try:
await self._emit_tool_call(provider)
events = await self._read_until(ws, "voice.playback_drain.requested")
types = [e["type"] for e in events]
self.assertNotIn("hermes.run.promoted", types)
# Real result delivered to the provider (not an interim background ack).
self.assertTrue(provider.connection.tool_results)
self.assertNotEqual(
provider.connection.tool_results[0][1].get("status"),
"running_in_background",
)
finally:
await ws.close()
async def test_explicit_background_mode_promotes_immediately(self) -> None:
# Tier C: mode="background" detaches immediately, even with grace-period
# promotion turned off and a long grace window.
broker = GatedHermesToolBroker()
ws, provider, body = await self._open(broker=broker)
session = self._server().realtime_agent.sessions[body["session_id"]]
session.promotion_enabled = False
session.promote_after_ms = 60000
try:
await self._emit_tool_call(provider, mode="background")
events = await self._read_until(ws, "hermes.run.promoted")
promoted = next(e for e in events if e["type"] == "hermes.run.promoted")
self.assertEqual(promoted["tier"], "durable")
broker.release.set()
done = await self._read_until(ws, "hermes.run.background_completed")
self.assertTrue(
any(e["type"] == "hermes.run.background_completed" for e in done)
)
finally:
broker.release.set()
await ws.close()
async def test_cancel_during_promoted_run(self) -> None:
broker = GatedHermesToolBroker()
ws, provider, _ = await self._open(broker=broker)
try:
await self._emit_tool_call(provider)
await self._read_until(ws, "hermes.run.promoted")
await broker.started.wait()
await ws.send_json({"type": "response.cancel"})
cancelled = await self._read_until(ws, "hermes.run.cancelled")
self.assertTrue(any(e["type"] == "hermes.run.cancelled" for e in cancelled))
for _ in range(50):
if broker.cancelled.is_set():
break
await asyncio.sleep(0.02)
self.assertTrue(broker.cancelled.is_set())
finally:
broker.release.set()
await ws.close()
async def test_detach_resume_replays_background_completed(self) -> None:
broker = GatedHermesToolBroker()
ws, provider, body = await self._open(broker=broker)
await self._emit_tool_call(provider)
await self._read_until(ws, "hermes.run.promoted")
ready = body["_ready"]
# Detach: close the websocket while the run is still in the background.
await ws.close(code=1001, message=b"network changed")
session = self._server().realtime_agent.sessions[body["session_id"]]
for _ in range(50):
if session.detached_at is not None:
break
await asyncio.sleep(0.01)
self.assertIsNotNone(session.detached_at)
# Complete the run while detached -> events recorded to the ring.
broker.release.set()
for _ in range(100):
if any(
e.get("type") == "hermes.run.background_completed"
for e in session.event_ring
):
break
await asyncio.sleep(0.02)
ws2 = await self.client.ws_connect(
body["websocket_path"], headers=self._bearer(body["_token"])
)
try:
await ws2.send_json(
{
"type": "session.resume",
"resume_token": body["resume_token"],
"last_event_id": ready["event_id"],
"last_audio_event_id": 0,
"last_played_audio_event_id": 0,
"last_input_chunk_id": 0,
}
)
events = await self._read_until(ws2, "voice.replay.done", limit=60)
types = [e["type"] for e in events]
self.assertIn("voice.session.resumed", types)
self.assertIn("hermes.run.background_completed", types)
replayed = next(
e for e in events if e["type"] == "hermes.run.background_completed"
)
self.assertTrue(replayed.get("replayed"))
finally:
await ws2.close()
if __name__ == "__main__":
unittest.main()
+151
View File
@@ -0,0 +1,151 @@
#!/usr/bin/env python3
"""Phase 0 idle-tolerance probe for ADR 33 / the background-Hermes-runs plan.
Opens a provider-native realtime-agent socket (xAI or OpenAI) via the existing
relay adapters, configures it, then holds it QUIESCENT — no ``response.create``,
no input audio — for a series of idle windows. It logs whether the socket
survives, any server-side close codes/timeouts, and any turn/VAD artifacts on the
first post-idle ``response.create``.
This answers the only factual unknown that gates ADR 33 default-on: can a
provider session hold the floor while a background Hermes run completes, or must
that provider's Tier B path keep-alive / reopen the socket instead?
This is a DEV PROBE — it is not imported by the relay and ships nothing. Run it
on the relay host where provider credentials are configured (env or xAI OAuth
store), with the repo root on PYTHONPATH:
python scripts/realtime-provider-idle-probe.py --provider xai
python scripts/realtime-provider-idle-probe.py --provider openai --windows 30,60,120
Record the per-provider verdict in docs/realtime-voice-poc.md ("Idle tolerance").
"""
from __future__ import annotations
import argparse
import asyncio
import contextlib
import sys
import time
from pathlib import Path
# Allow running as a bare script from the repo root.
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from plugin.relay.realtime_agent.models import ( # noqa: E402
HERMES_TOOL_SCHEMAS,
ProviderEventKind,
RealtimeAgentSessionConfig,
)
from plugin.relay.realtime_agent.providers.openai import ( # noqa: E402
OpenAIRealtimeAgentProvider,
)
from plugin.relay.realtime_agent.providers.xai import ( # noqa: E402
XAIRealtimeAgentProvider,
)
_PROVIDERS = {
"xai": (XAIRealtimeAgentProvider, "grok-voice-latest", "ember", 24000),
"openai": (OpenAIRealtimeAgentProvider, "gpt-realtime-2", "alloy", 24000),
}
def _session_config(provider_id: str) -> RealtimeAgentSessionConfig:
_, model, voice, rate = _PROVIDERS[provider_id]
return RealtimeAgentSessionConfig(
provider=provider_id,
model=model,
voice=voice,
sample_rate=rate,
profile=None,
hermes_session_id=None,
instructions="You are an idle-tolerance probe. Do not speak unless asked.",
provider_options={},
tools=HERMES_TOOL_SCHEMAS,
)
async def _drain_events(connection, stop: asyncio.Event, sink: list[str]) -> None:
"""Record provider events (esp. ERROR/close) while we idle."""
try:
async for event in connection.events():
sink.append(f"{time.monotonic():.1f}s {event.kind.value}")
if event.kind == ProviderEventKind.ERROR:
sink.append(f" ERROR payload={event.payload!r}")
if stop.is_set():
return
except Exception as exc: # noqa: BLE001 - probe wants the raw failure
sink.append(f"{time.monotonic():.1f}s events() raised {exc!r}")
finally:
stop.set()
async def probe(provider_id: str, windows: list[int]) -> int:
cls, *_ = _PROVIDERS[provider_id]
provider = cls()
print(f"[probe] connecting {provider_id} ...")
connection = await provider.connect(_session_config(provider_id))
print("[probe] connected + configured")
stop = asyncio.Event()
sink: list[str] = []
reader = asyncio.create_task(_drain_events(connection, stop, sink))
verdict = "hold-floor-ok"
try:
for window in windows:
print(f"[probe] idling {window}s (no response.create, no audio) ...")
try:
await asyncio.wait_for(stop.wait(), timeout=window)
# stop fired => the socket closed / errored while idle.
print(f"[probe] socket DID NOT survive {window}s idle window")
verdict = "must-reopen"
break
except asyncio.TimeoutError:
print(f"[probe] survived {window}s idle")
# Probe a post-idle turn for VAD/turn artifacts.
print("[probe] requesting a short post-idle response ...")
before = len(sink)
await connection.send_text("Say the single word: ready.")
await asyncio.sleep(8)
produced = sink[before:]
got_audio = any("audio_delta" in line for line in produced)
got_error = any("error" in line for line in produced)
print(f"[probe] post-idle audio={got_audio} error={got_error}")
if got_error or not got_audio:
verdict = "needs-keepalive"
with contextlib.suppress(Exception):
await connection.cancel_response()
finally:
stop.set()
with contextlib.suppress(Exception):
await connection.close()
reader.cancel()
with contextlib.suppress(asyncio.CancelledError):
await reader
print("\n[probe] event trace:")
for line in sink:
print(" " + line)
print(f"\n[probe] VERDICT ({provider_id}): {verdict}")
print("[probe] record this in docs/realtime-voice-poc.md -> Idle tolerance")
return 0
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--provider", choices=sorted(_PROVIDERS), default="xai")
parser.add_argument(
"--windows",
default="30,60,120",
help="comma-separated idle windows in seconds",
)
args = parser.parse_args()
windows = [int(w) for w in str(args.windows).split(",") if w.strip()]
return asyncio.run(probe(args.provider, windows))
if __name__ == "__main__":
raise SystemExit(main())
+1 -1
View File
@@ -12,7 +12,7 @@ The plugin is a thin observer — it never modifies state, never writes to your
**On your server:**
- hermes-agent with the Dashboard Plugin System (upstream commit `01214a7f` on `axiom`, or any later `main` once [PR #8556](https://github.com/NousResearch/hermes-agent/pull/8556) and its dashboard followups merge). `hermes dashboard start` must already work for you.
- hermes-agent with the Dashboard Plugin System (upstream commit `01214a7f` on `axiom`, or any later `main` that includes the dashboard plugin follow-ups). `hermes dashboard start` must already work for you.
- The canonical Hermes-Relay install — if you ran the one-liner on the [Quick Start](/guide/getting-started), you're done. The installer symlinks `~/.hermes/plugins/hermes-relay` → the plugin subtree and the dashboard scanner picks up `plugin/dashboard/manifest.json` automatically.
- A gateway restart after install: `systemctl --user restart hermes-gateway`.
+23
View File
@@ -227,6 +227,29 @@ voice-output stream cannot compete with the realtime provider. The final spoken
answer is generated by the realtime provider after Hermes returns a compact
result, so tool output is summarized naturally instead of read aloud.
#### Background tasks
Some requests take a while — research, multi-step work, a long command. Instead
of freezing the conversation until they finish, Realtime Agent **promotes** a
slow run to the background: the agent says a short "I'm on it" and you can keep
talking, ask something else, or just wait. When the task finishes, the agent
speaks the answer. Asking for something explicitly long starts a background task
right away.
You control this under **Voice Settings → Realtime Agent → Background tasks**:
- **Promote long tasks** — turn the behavior on or off. With it off, the agent
waits silently until the task finishes (the old behavior).
- **Spoken handoff** — whether the agent says a short acknowledgement when a task
moves to the background, or just shows it on screen.
- **When the answer is ready** — *Speak* it as soon as you're not mid-sentence,
*Notify* and speak when you re-engage, or *Show only* (no spoken answer).
While a background task is running, a small "working on it" chip stays visible in
the voice screen. You can cancel a background task at any time the same way you
cancel any turn. Hermes still owns the task end-to-end — promotion only changes
*when* the answer is spoken, never who runs the tools.
Provider-native Android paths stream mic PCM to a relay-owned realtime provider
WebSocket session. Android commits the captured utterance, the active provider
owns input transcription and speech generation, and Hermes still owns profile
+1 -1
View File
@@ -55,7 +55,7 @@ All commands are fetched dynamically from the server where possible:
- **Configuration**: `/model`, `/personality`, `/reasoning`, `/yolo`, `/verbose`, `/voice`
- **Info**: `/help`, `/status`, `/usage`, `/insights`, `/commands`
- **Personalities**: generated from server config (`config.agent.personalities`) — `/personality victor`, `/personality creative`, etc.
- **Skills**: dynamically fetched from `GET /api/skills` — 90+ server skills grouped by category (creative, devops, research, etc.)
- **Skills**: dynamically fetched from `GET /v1/skills` with fallback to legacy `GET /api/skills` — server skills grouped by category (creative, devops, research, etc.)
## Tool Execution
+2 -2
View File
@@ -18,9 +18,9 @@ Hermes-Relay voice endpoints can reuse this same Bearer token. Relay validates i
## How endpoints get served
Installing the plugin via `install.sh` is enough to make all of the endpoints below work — including the management ones (`/api/sessions/*`, `/api/memory`, `/api/skills`, `/api/config`, `/api/available-models`). The plugin wires the gateway up at install time so these are served on the same `:8642` host as the standard `/v1/*` endpoints, with the same `Authorization: Bearer …` auth.
Current upstream hermes-agent includes the baseline session API (`/api/sessions/*`, chat, stream, fork, messages) and native capability discovery (`/v1/skills`, `/v1/toolsets`). Installing the Relay plugin via `install.sh` keeps older or partial core builds compatible by injecting only the missing management endpoints Relay still consumes, such as `/api/sessions/search`, `/api/memory`, `/api/config`, legacy `/api/skills` detail routes, `/api/available-models`, and voice aliases. These are served on the same `:8642` host as the standard `/v1/*` endpoints, with the same optional `Authorization: Bearer ***` auth.
**Chat streaming uses standard `/v1/runs`** by default — it emits structured `tool.started`/`tool.completed` SSE events for live tool progress cards in the Android app. The app's `Settings → Chat → Streaming endpoint = "Auto"` (default) probes per-endpoint capability and picks the best chat path automatically; you can manually force `Sessions` or `Runs` mode for debugging.
**Chat streaming uses standard `/v1/runs`** by default — it emits structured `tool.started`/`tool.completed` SSE events for live tool progress cards in the Android app. The app's `Settings → Chat → Streaming endpoint = "Auto"` (default) probes per-endpoint capability and picks the best chat path automatically; you can manually force `Sessions` or `Runs` mode for debugging. Skill discovery prefers `GET /v1/skills` and falls back to legacy `GET /api/skills`.
## Endpoints