Compare commits
24
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0a5649d016 | ||
|
|
7d667b9096 | ||
|
|
f7edffcd81 | ||
|
|
b53df95bbf | ||
|
|
f879fbbd51 | ||
|
|
a586f3dd60 | ||
|
|
d41e1b659b | ||
|
|
de9988322b | ||
|
|
ff09c6116b | ||
|
|
7c5b7729c1 | ||
|
|
304d8e26c8 | ||
|
|
742d899042 | ||
|
|
783fc34e01 | ||
|
|
ac51a95842 | ||
|
|
697610d92c | ||
|
|
f300531054 | ||
|
|
0c7c21964d | ||
|
|
ab1ffdd2aa | ||
|
|
ab10097f55 | ||
|
|
e294a571a4 | ||
|
|
35893e239f | ||
|
|
06332f8f4e | ||
|
|
60623c1751 | ||
|
|
024ee6c1db |
@@ -6,6 +6,18 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
|
||||
- **Persistent Realtime Agent conversation.** Realtime Agent voice now keeps one provider session/socket open across turns instead of creating a fresh session per utterance, so the provider retains the live conversation (follow-up references work) and turns skip session-setup latency. The relay needed no change — it already supported multiple turns on one socket. A **Voice Settings → Realtime Agent → Persistent session** toggle (default on) falls back to the legacy per-utterance path. See `docs/plans/2026-05-24-realtime-persistent-session.md`.
|
||||
|
||||
- **Background Hermes runs in Realtime Agent voice (ADR 33).** Long Hermes tasks no longer freeze the realtime conversation. A run that exceeds a grace window is promoted to a tracked background task: the provider speaks a short handoff ("I'm on it"), the conversation stays responsive, and the answer is spoken once the run finishes. `hermes_run_task(mode="background")` starts a durable run immediately. New relay events `hermes.run.promoted` and `hermes.run.background_completed`, plus `tier`/`floor` fields on `hermes.run.progress`.
|
||||
|
||||
- **Relay audio floor owner.** A single-owner audio floor (provider / relay-TTS / Android-filler) makes explicit the serialization that the old blocking design provided implicitly, so a completed background result never barges in and two voices never overlap.
|
||||
|
||||
- **Voice Settings → Realtime Agent → Background tasks.** New controls to enable/disable promotion, toggle the spoken handoff, and choose result delivery (speak when idle / notify / show only). A persistent "working on it" chip appears in the voice overlay while a background task runs.
|
||||
|
||||
- **Provider idle-tolerance probe.** `scripts/realtime-provider-idle-probe.py` records a per-provider verdict (hold-floor-ok / needs-keepalive / must-reopen) for holding a realtime socket quiescent during a background run; see `docs/realtime-voice-poc.md`.
|
||||
|
||||
## [0.8.1] - 2026-05-26
|
||||
|
||||
### Fixed
|
||||
|
||||
@@ -33,25 +33,26 @@ Chat goes directly to the API server via HTTP/SSE. The API key (Bearer token) is
|
||||
| `GET /health` | Health check | — |
|
||||
| `GET/POST/PATCH/DELETE /api/jobs/*` | Cron job management (api_server surface) | — |
|
||||
|
||||
**Non-standard endpoints (provided by fork OR by plugin bootstrap):**
|
||||
**Baseline upstream endpoints vs compatibility endpoints:**
|
||||
|
||||
These endpoints are not in stock upstream `gateway/platforms/api_server.py`. There are three ways a hermes-agent install can serve them:
|
||||
Upstream hermes-agent now has a native baseline for API Server session control and skill/toolset discovery:
|
||||
|
||||
1. **Codename-11 fork** (`feat/session-api` branch, deployed on the `axiom` branch) — adds them natively. Submitted upstream as PR [#8556](https://github.com/NousResearch/hermes-agent/pull/8556) *"feat(api-server): add session management API for frontend clients"* — scope is broader than the title: sessions CRUD + session chat/stream + memory + skills + config + available-models.
|
||||
2. **Bootstrap injection** (`hermes_relay_bootstrap/`) — monkey-patches aiohttp on startup via `.pth` file. Does NOT inject `/api/sessions/{id}/chat/stream` — use `/v1/runs` for chat.
|
||||
3. **Upstream-merged** (post PR #8556) — bootstrap auto-detects and no-ops.
|
||||
1. **Native upstream** — commit [`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde), merged via PR [#33134](https://github.com/NousResearch/hermes-agent/pull/33134), salvaged the focused session-control work from closed PR [#29302](https://github.com/NousResearch/hermes-agent/pull/29302). It provides `/api/sessions/*`, session chat/stream, fork/messages, plus `/v1/skills` and `/v1/toolsets`.
|
||||
2. **Codename-11 `axiom` fork** — still carries compatibility/client-metadata routes that upstream does not provide yet: `/api/sessions/search`, `/api/memory`, `/api/skills` detail routes, `/api/config`, and `/api/available-models`.
|
||||
3. **Bootstrap injection** (`hermes_relay_bootstrap/`) — monkey-patches aiohttp on startup via `.pth` file and injects only missing compatibility routes for older or partial upstream builds. It should remain per-route/per-feature, not all-or-nothing.
|
||||
|
||||
| Endpoint | Purpose | Provided by |
|
||||
|----------|---------|-------------|
|
||||
| `GET /api/sessions` (CRUD) | Session list/create/rename/delete/fork | Fork OR bootstrap OR upstream-merged |
|
||||
| `GET /api/sessions/{id}/messages` | Conversation history | Fork OR bootstrap OR upstream-merged |
|
||||
| `GET /api/sessions/search` | Full-text message search | Fork OR bootstrap OR upstream-merged |
|
||||
| `POST /api/sessions/{id}/chat/stream` | Session-based SSE chat | Fork OR upstream-merged ONLY (NOT bootstrap) |
|
||||
| `GET /api/config`, `PATCH /api/config` | Personalities + model config | Fork OR bootstrap OR upstream-merged |
|
||||
| `GET /api/skills`, `/{name}` | Skill discovery (list + detail) | Fork OR bootstrap OR upstream-merged |
|
||||
| `GET /api/sessions` (CRUD) | Session list/create/rename/delete/fork | Native upstream OR fork OR bootstrap |
|
||||
| `GET /api/sessions/{id}/messages` | Conversation history | Native upstream OR fork OR bootstrap |
|
||||
| `GET /api/sessions/search` | Full-text message search | Fork OR bootstrap only |
|
||||
| `POST /api/sessions/{id}/chat/stream` | Session-based SSE chat | Native upstream OR fork only (NOT bootstrap) |
|
||||
| `GET /v1/skills` | Skill list metadata | Native upstream OR fork |
|
||||
| `GET /api/config`, `PATCH /api/config` | Personalities + model config | Fork OR bootstrap only |
|
||||
| `GET /api/skills`, `/{name}` | Legacy skill discovery/detail routes | Fork OR bootstrap only; Android prefers `/v1/skills` first |
|
||||
| `PUT /api/skills/toggle` | Enable/disable installed skill | `hermes_cli/web_server.py` dashboard surface; mirrored into bootstrap |
|
||||
| `GET/POST/PATCH/DELETE /api/memory` | Memory CRUD | Fork OR bootstrap OR upstream-merged |
|
||||
| `GET /api/available-models` | Provider model list | Fork OR bootstrap OR upstream-merged |
|
||||
| `GET/POST/PATCH/DELETE /api/memory` | Memory CRUD | Fork OR bootstrap only |
|
||||
| `GET /api/available-models` | Provider-aware model list | Fork OR bootstrap only |
|
||||
|
||||
The Android client probes per-endpoint capability via `HermesApiClient.probeCapabilities()` (returns `ServerCapabilities`). When `streamingEndpoint = "auto"`, `ConnectionViewModel.resolveStreamingEndpoint()` picks `sessions` or `runs` based on the capability snapshot.
|
||||
|
||||
@@ -67,7 +68,7 @@ hermes-agent ships a second web server at `hermes_cli/web_server.py` that hosts
|
||||
## Key Instructions
|
||||
- **Always verify upstream before assuming an endpoint exists.** Check `gateway/platforms/api_server.py` in hermes-agent. If an endpoint isn't there, document whether bootstrap injects it or it requires the fork.
|
||||
- If we use a non-standard endpoint, ensure `probeCapabilities()` covers it and the auto-resolver degrades gracefully.
|
||||
- **Bootstrap maintenance:** Remove `hermes_relay_bootstrap/` in one PR once PR #8556 merges. It's no-op-compatible, so leaving it in place during rollout is harmless.
|
||||
- **Bootstrap maintenance:** Do not remove `hermes_relay_bootstrap/` just because upstream has native sessions. It can start shrinking only after each Relay-consuming compatibility route has a native replacement or the Android/Desktop clients have migrated away from it.
|
||||
|
||||
## Repository Layout
|
||||
|
||||
@@ -107,7 +108,7 @@ hermes-android/
|
||||
│ ├── tools/ # android_navigate.py, android_notifications.py
|
||||
│ └── dashboard/ # hermes-agent dashboard plugin — manifest, React UI, FastAPI proxy
|
||||
├── relay_server/ ← Thin compat shim → plugin.relay (legacy entrypoint)
|
||||
├── hermes_relay_bootstrap/ ← Runtime patch for vanilla upstream; removable after PR #8556
|
||||
├── hermes_relay_bootstrap/ ← Runtime patch for vanilla/partial upstream compatibility routes
|
||||
├── skills/devops/hermes-relay-pair/ ← /hermes-relay-pair slash command
|
||||
├── scripts/ ← dev.bat, bridge-smoke.sh, bump-version.sh
|
||||
└── docs/ ← spec, decisions, security, relay-server, mcp-tooling
|
||||
@@ -197,7 +198,7 @@ hermes-android/
|
||||
| **App — Voice** | |
|
||||
| `voice/VoiceViewModel.kt` | Voice turn state machine; TTS queue; `ignoreAssistantId`; `errorEvents: SharedFlow` |
|
||||
| `audio/VoiceRecorder.kt` | MediaRecorder wrapper; perceptual amplitude curve; `.m4a` at 16kHz/64kbps |
|
||||
| `audio/VoicePlayer.kt` | MediaPlayer + Visualizer; amplitude StateFlow; `awaitCompletion()` via coroutine |
|
||||
| `audio/VoicePlayer.kt` | Media3 ExoPlayer (gapless TTS queue) + Visualizer; amplitude StateFlow; `awaitCompletion()` via coroutine; `audioSessionId` is a thread-safe `@Volatile` cache |
|
||||
| `network/RelayVoiceClient.kt` | OkHttp for `/voice/transcribe`, `/synthesize`, `/config` |
|
||||
| `voice/VoiceBridgeIntentHandler.kt` | Interface routing voice utterances to bridge; impls per flavor via factory |
|
||||
| `voice/VoiceIntentClassifier.kt` | Regex phone-control classifier (sideload only); false-negatives preferred over false-positives |
|
||||
@@ -230,7 +231,7 @@ hermes-android/
|
||||
| `plugin/pair.py` | QR payload builder + CLI; `build_payload(sign=True)`; `--register-code` fallback |
|
||||
| `install.sh` | Canonical installer — 6 steps; idempotent; drops `hermes-relay-update` shim |
|
||||
| `uninstall.sh` | Canonical uninstaller; reverses install.sh; never touches `.env` or `state.db` |
|
||||
| `hermes_relay_bootstrap/` | Runtime patch for vanilla upstream; no-op on fork/upstream-merged; remove after PR #8556 |
|
||||
| `hermes_relay_bootstrap/` | Runtime patch for vanilla/partial upstream compatibility routes; shrink per route group after native parity or client migration |
|
||||
| **Plugin — Dashboard** | |
|
||||
| `plugin/dashboard/manifest.json` | Declares tab, entry bundle, and FastAPI module for hermes-agent discovery |
|
||||
| `plugin/dashboard/plugin_api.py` | FastAPI router proxying 5 routes to relay over loopback; `/pairing` body = API-server overrides (host/port/tls/api_key), relay URL auto-derived |
|
||||
|
||||
@@ -1,5 +1,25 @@
|
||||
# Hermes-Relay — Dev Log
|
||||
|
||||
## 2026-06-05 — Refresh upstream baseline docs after native session merge
|
||||
|
||||
**Context.** Upstream Hermes Agent merged native API Server session controls as commit `f7527b0` via PR #33134, and Android now prefers native `/v1/skills` with legacy fallback. Several Relay docs still treated PR #8556/#29302 as the future removal trigger for `hermes_relay_bootstrap/`.
|
||||
|
||||
**What changed.** Updated contributor/user docs to separate native upstream baseline routes (`/api/sessions/*`, `/v1/skills`, `/v1/toolsets`) from Relay compatibility routes that still require `axiom` or bootstrap (`/api/sessions/search`, `/api/memory`, `/api/config`, legacy skill detail routes, `/api/available-models`, voice aliases). The bootstrap retirement rule is now per route group, not wholesale deletion after one upstream session merge.
|
||||
|
||||
**Verification.** Ran stale-reference searches for `#8556`, `#29302`, and legacy skills routes after edits; remaining references are historical devlog/plans or explicitly marked compatibility/fallback.
|
||||
|
||||
---
|
||||
|
||||
## 2026-06-05 — Prefer upstream `/v1/skills` with legacy fallback
|
||||
|
||||
**Context.** Upstream Hermes Agent now has baseline skill/session API surface area, while Axiom's fork still preserves richer Relay-specific `/api/*` compatibility routes. The Android client should begin consuming upstream-compatible skill listings when present without breaking older fork/bootstrap installs.
|
||||
|
||||
**What changed.** `HermesApiClient.getSkills()` now tries `/v1/skills` first, then falls back to `/api/skills`. Skill parsing accepts upstream OpenAI-style list envelopes (`{"object":"list","data":[...]}`), legacy fork envelopes (`{"skills":[...]}` / `{"items":[...]}`), and direct arrays.
|
||||
|
||||
**Verification.** Added pure Kotlin unit coverage for endpoint order and `/v1/skills` `data` parsing. Verified with `ANDROID_HOME=$HOME/Android/Sdk ./gradlew :app:testGooglePlayDebugUnitTest --tests 'com.hermesandroid.relay.network.HermesApiClientTest'` → BUILD SUCCESSFUL. `git diff --check` passes.
|
||||
|
||||
---
|
||||
|
||||
## 2026-05-26 — Fix voice-mode crash: ExoPlayer audio session id read off-main (barge-in + legacy TTS)
|
||||
|
||||
**Report.** Discord user, sideload latest: voice chat crashes the instant Hermes starts answering — "I hear just 2 letters and it crashes." Stack: `IllegalStateException: Player is accessed on the wrong thread. Current thread: 'DefaultDispatcher-worker-4', Expected thread: 'main'` with the Media3 `player-accessed-on-wrong-thread` doc link and a `Suppressed: ... Dispatchers.IO`.
|
||||
@@ -10,7 +30,26 @@
|
||||
|
||||
**Tests.** Added `VoicePlayerTest` coverage: getter reflects the analytics-listener-cached id, never re-invokes `exoPlayer.audioSessionId` (the off-main call), and defaults to 0 before allocation. Captured the `AnalyticsListener` in the MockK harness. Verified locally: `:app:lintGooglePlayDebug` + `:app:testGooglePlayDebugUnitTest --tests VoicePlayerTest` both green (BUILD SUCCESSFUL).
|
||||
|
||||
**Next.** Cherry-picked onto `android-v0.8.1` hotfix branch (off the `android-v0.8.0` tag) for a focused patch release; merges back to `dev` after.
|
||||
**Next.** Landed on `dev` via PR #60, then shipped as the focused patch release `android-v0.8.1` (cherry-picked off the `android-v0.8.0` tag, PR #62); `main` merged back to `dev`.
|
||||
|
||||
---
|
||||
|
||||
## 2026-05-24 — Background Hermes runs in Realtime Agent voice (ADR 33)
|
||||
|
||||
**Context.** Realtime Agent ran each Hermes turn synchronously *inside* the provider event pump (`_run_brokered_tool` did `return await task`), so a long research/multi-tool/desktop run froze the whole realtime session until it finished. ADR 33 + `docs/plans/2026-05-24-realtime-background-hermes-runs.md` define a three-tier model (foreground / promoted / durable) with the relay as an explicit audio-floor owner. Branch `feature/realtime-background-hermes-runs`.
|
||||
|
||||
**What shipped (phased, per the plan):**
|
||||
|
||||
- **Phase 0 — idle-tolerance probe + verdict.** `scripts/realtime-provider-idle-probe.py` + an "Idle tolerance" section in `docs/realtime-voice-poc.md`. **Ran live against OpenAI** (`VOICE_TOOLS_OPENAI_KEY` in `~/.hermes/.env`): the session survived 10s/20s/30s quiescent windows and returned clean audio on every post-idle turn → verdict **`hold-floor-ok`**. xAI has no creds on the dev box, so its verdict is recorded analytically as `hold-floor-ok` (same `turn_detection:None` multi-turn model; the implementation closes the pending call rather than holding an open response) — confirm on the relay host. Incidental finding logged: OpenAI now wants `session.audio.output.format.rate` at `session.update` (minor `_session_update` follow-up; session still worked).
|
||||
- **Phase 1 — floor owner.** New `plugin/relay/realtime_agent/floor.py`: pure, single-owner audio floor (`provider` / `relay_tts` / `android_filler` mouths; `idle/provider_speaking/hermes_filler/result_pending` labels). Wired behavior-preservingly into the broker (acquire/release on AUDIO_DELTA/AUDIO_DONE/RESPONSE_DONE; relay-TTS render holds the floor; filler gated by `can_speak`). Invariants in `test_realtime_floor.py`.
|
||||
- **Phase 2 — Tier B promotion (was default off).** `_run_brokered_tool` shields the run and waits `promote_after_ms`; if still running it detaches to the background, closes the pending provider call with an interim ack, optionally speaks a handoff, and `_deliver_background_result` speaks the answer once the floor is idle. New events `hermes.run.promoted` / `hermes.run.background_completed` + `tier`/`floor` on progress; 8 new settings. `test_realtime_promotion.py` (promote+pump-responsive, short=no-promote, cancel, detach-resume-replays).
|
||||
- **Phase 3 — default-on + Tier C + Android + docs.** Flipped `promotion_enabled` default **true** (safe: the path closes the pending call rather than holding an open response, so the socket only sees the normal between-turns idle gap). `hermes_run_task(mode="background")` detaches immediately (`tier:"durable"`). Settings exposed on `GET/PATCH /voice/realtime-agent/config`. Android: parse new events → "working on it" chip; Voice Settings → Realtime Agent → Background tasks (promote toggle, spoken-handoff toggle, result-delivery segmented control) → `RelayVoiceClient.updateRealtimeAgentPromotion()`.
|
||||
|
||||
**Why default-on despite the Phase 0 gate.** The implementation closes the pending function call with an interim background ack instead of parking an open provider response, so the worst-case "idle open-response" the gate worried about doesn't occur — the socket sits in the same idle state it does between any two user turns. The probe is retained to confirm per-provider survival; documented in config + ADR.
|
||||
|
||||
**Verified.** Python realtime suite **58 tests green** (`test_realtime_floor`, `test_realtime_promotion`, both provider suites, routes, profile-voice-config). `./gradlew lint` — Kotlin compiles clean; the only 2 lint errors are in the gitignored `local.properties` (absent in CI). Pre-existing unrelated `test_reads_hermes_xai_oauth_credential_pool` failure confirmed on `origin/dev` baseline.
|
||||
|
||||
**Next.** Confirm the xAI idle verdict on the relay host (where xAI creds live); fix the OpenAI `_session_update` rate field; run the lab smoke on a paired device. Open the PR to `dev`.
|
||||
|
||||
---
|
||||
|
||||
|
||||
+2
-2
@@ -101,7 +101,7 @@ Small follow-ons to v0.4 deliberately deferred to keep the v0.4.0 release surfac
|
||||
|
||||
**What the middleware can do (near-term, ships via install.sh).** New aiohttp middleware in `hermes_relay_bootstrap/_command_middleware.py`, installed at the same `_PatchedApplication.__setitem__` hook as the current route injection so it lands before `AppRunner.setup()` freezes the app. Filters by `request.path in ("/v1/runs", "/v1/chat/completions")` — zero-cost fast path for everything else. On chat paths: parses the body, lazy-imports `GATEWAY_KNOWN_COMMANDS` + `resolve_command()` + `gateway_help_lines()` from `hermes_cli.commands`, and splits on command type:
|
||||
- **Stateless commands** (`/help`, `/commands`, and any others the upstream Option B PR ends up supporting without router state) — actually dispatch, emit a synthetic SSE stream matching the runs handler's existing event shape so the Android client at `HermesApiClient.kt:655-715` renders it as a normal assistant turn.
|
||||
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, etc. — most of the registry) — emit a synthetic SSE stream whose content is a short, helpful notice: *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` (post-PR-#8556) or a channel with session state. For commands that work here, type `/help`."* This replaces the LLM hallucination with a deterministic, accurate message that points the user at the real fix.
|
||||
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, etc. — most of the registry) — emit a synthetic SSE stream whose content is a short, helpful notice: *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` or a channel with session state. For commands that work here, type `/help`."* This replaces the LLM hallucination with a deterministic, accurate message that points the user at the real fix.
|
||||
|
||||
**On no match** (unknown command, cli-only command, or plain text): falls through to `handler(request)` unchanged. Fork-detects the same way the existing injection does — if the upstream preprocessor PR lands first, the middleware no-ops.
|
||||
|
||||
@@ -109,7 +109,7 @@ Small follow-ons to v0.4 deliberately deferred to keep the v0.4.0 release surfac
|
||||
|
||||
**Files.** New `hermes_relay_bootstrap/_command_middleware.py` (~150 LOC), one-line append in `_patch.py` inside `_maybe_register_routes`, stdlib `unittest` coverage in `plugin/tests/test_bootstrap_command_middleware.py` mirroring the existing `test_bootstrap_patch.py` harness. Mirrors the upstream Option B PR exactly so the two can be reviewed side-by-side.
|
||||
|
||||
**Phase 2 — stateful dispatch on the session chat stream endpoint (post PR #8556).** Once PR #8556 merges and `/api/sessions/{id}/chat/stream` ships natively in upstream, a separate middleware (or a follow-up upstream PR) can add a preprocessor **scoped to that endpoint only**, leveraging the `session_id` in the URL as the persistence handle. At that point stateful commands become a dict write against session-scoped state — `session.model_override = new_model` — without needing to refactor `GatewayRouter` or plumb api_server into the router. Much smaller than a full router refactor, and it matches upstream's partition: `/v1/*` stays stateless, statefulness lives on `/api/sessions/*`. Blocked on #8556 landing.
|
||||
**Phase 2 — stateful dispatch on the session chat stream endpoint (unblocked by PR #33134 / commit `f7527b0`).** Since `/api/sessions/{id}/chat/stream` now ships natively in upstream, a separate middleware (or a follow-up upstream PR) can add a preprocessor **scoped to that endpoint only**, leveraging the `session_id` in the URL as the persistence handle. At that point stateful commands become a dict write against session-scoped state — `session.model_override = new_model` — without needing to refactor `GatewayRouter` or plumb api_server into the router. Much smaller than a full router refactor, and it matches upstream's partition: `/v1/*` stays stateless, statefulness lives on `/api/sessions/*`.
|
||||
|
||||
## Future — v0.5+
|
||||
|
||||
|
||||
@@ -78,10 +78,10 @@ Things to look into:
|
||||
- **Tool registration discoverability** — `android_*` tools register at gateway import time. There's no canonical "list installed plugin tools" API. Would adding one to upstream make sense, or is `gateway tool list` already enough?
|
||||
- **Versioning + compatibility ranges** — `pip install -e` doesn't enforce version pins between hermes-agent and our plugin. A breaking change in upstream's plugin loader could silently break us. Do we need a `hermes_compat: ">=0.8.0,<1.0.0"` field somewhere?
|
||||
- **`hermes-relay-self-setup` SKILL.md as a precedent** — we just shipped a self-installing skill that an LLM can fetch from a raw GitHub URL and execute. Does this pattern generalize? Could it become a recommended way for any third-party Hermes project to ship setup automation?
|
||||
- **Bootstrap injection** — `hermes_relay_bootstrap/` monkey-patches `aiohttp.web.Application` to inject endpoints into vanilla upstream. This is intentional but feels like a hack. Upstream PR #8556 (`feat/session-api`) will eventually let us delete it — verified 2026-04-15 that its scope covers the full bootstrap surface (sessions, memory, skills, config, available-models). Track that PR's status periodically.
|
||||
- **Gateway slash-command preprocessor — upstream Stage 1 PR.** Sibling follow-up to #8556. Intercepts known gateway commands on `/v1/runs` + `/v1/chat/completions`, dispatches the stateless ones (`/help`, `/commands`) via `gateway_help_lines()`, returns a deterministic "use a channel with session state" notice for the stateful majority. Currently being prepared in `C:/Users/Bailey/Desktop/Open-Projects/hermes-agent-pr-prep/` on branch `feat/api-server-gateway-commands`; awaiting subagent's code + draft PR body before pushing. See `docs/upstream-contributions.md` §5.
|
||||
- **Bootstrap injection shrink path** — `hermes_relay_bootstrap/` monkey-patches `aiohttp.web.Application` to inject endpoints into vanilla/partial upstream. Upstream commit `f7527b0` via PR #33134 now covers baseline sessions/chat/fork/message history, and `/v1/skills` covers list metadata. Do **not** delete the bootstrap wholesale yet: Relay still depends on compatibility routes that upstream lacks or does not match (`/api/sessions/search`, `/api/memory`, `/api/config`, legacy `/api/skills` detail routes, `/api/available-models`, and voice aliases). Shrink per route group only after native parity or client migration.
|
||||
- **Gateway slash-command preprocessor — upstream Stage 1 PR.** Follow-up to the native session-control baseline from PR #33134 / commit `f7527b0`. Intercepts known gateway commands on `/v1/runs` + `/v1/chat/completions`, dispatches the stateless ones (`/help`, `/commands`) via `gateway_help_lines()`, returns a deterministic "use a channel with session state" notice for the stateful majority. Currently being prepared in `C:/Users/Bailey/Desktop/Open-Projects/hermes-agent-pr-prep/` on branch `feat/api-server-gateway-commands`; awaiting subagent's code + draft PR body before pushing. See `docs/upstream-contributions.md` §5.
|
||||
- **Gateway slash-command preprocessor — bootstrap middleware (Stage 1 equivalent).** Sibling shim in `hermes_relay_bootstrap/_command_middleware.py` that mirrors the upstream Stage 1 PR as an aiohttp middleware injected at bootstrap time. Ships the hallucination fix to vanilla-upstream installs before the upstream PR lands. Planned for v0.4.1, after the current bridge feature branch wraps. See `ROADMAP.md` v0.4.1 entry.
|
||||
- **Stage 2 — stateful slash-command dispatch on `/api/sessions/{id}/chat/stream`.** Blocked on PR #8556 merging. Once session primitives ship upstream, add a preprocessor scoped to the session chat stream endpoint only, using `session_id` as the persistence handle. Separate upstream PR + matching bootstrap middleware. See `docs/upstream-contributions.md` §5 ("Stage 2").
|
||||
- **Stage 2 — stateful slash-command dispatch on `/api/sessions/{id}/chat/stream`.** Unblocked by upstream PR #33134 / commit `f7527b0`. Add a preprocessor scoped to the session chat stream endpoint only, using `session_id` as the persistence handle. Separate upstream PR + matching bootstrap middleware. See `docs/upstream-contributions.md` §5 ("Stage 2").
|
||||
|
||||
When the answer becomes clearer, this section becomes either an ADR in `docs/decisions.md` or a Plan under `Plans/`.
|
||||
|
||||
|
||||
@@ -29,6 +29,13 @@ data class VoiceSettings(
|
||||
val autoTts: Boolean = false,
|
||||
val language: String = "",
|
||||
val realtimeTraceDetails: Boolean = false,
|
||||
/**
|
||||
* When true (default), Realtime Agent keeps one provider session/socket open
|
||||
* across turns (persistent conversation). When false, falls back to the
|
||||
* legacy one-session-per-utterance path. See
|
||||
* docs/plans/2026-05-24-realtime-persistent-session.md.
|
||||
*/
|
||||
val realtimePersistentSession: Boolean = true,
|
||||
)
|
||||
|
||||
enum class VoiceEngineMode(val storageValue: String) {
|
||||
@@ -52,6 +59,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
|
||||
private val KEY_AUTO_TTS = booleanPreferencesKey("voice_auto_tts")
|
||||
private val KEY_LANGUAGE = stringPreferencesKey("voice_language")
|
||||
private val KEY_REALTIME_TRACE_DETAILS = booleanPreferencesKey("voice_realtime_trace_details")
|
||||
private val KEY_REALTIME_PERSISTENT_SESSION =
|
||||
booleanPreferencesKey("voice_realtime_persistent_session")
|
||||
|
||||
const val DEFAULT_ENGINE_MODE = "hermes_voice_output"
|
||||
const val DEFAULT_INTERACTION_MODE = "tap"
|
||||
@@ -59,6 +68,7 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
|
||||
const val DEFAULT_AUTO_TTS = false
|
||||
const val DEFAULT_LANGUAGE = ""
|
||||
const val DEFAULT_REALTIME_TRACE_DETAILS = false
|
||||
const val DEFAULT_REALTIME_PERSISTENT_SESSION = true
|
||||
}
|
||||
|
||||
val settings: Flow<VoiceSettings> = dataStore.data
|
||||
@@ -73,6 +83,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
|
||||
language = prefs[KEY_LANGUAGE] ?: DEFAULT_LANGUAGE,
|
||||
realtimeTraceDetails = prefs[KEY_REALTIME_TRACE_DETAILS]
|
||||
?: DEFAULT_REALTIME_TRACE_DETAILS,
|
||||
realtimePersistentSession = prefs[KEY_REALTIME_PERSISTENT_SESSION]
|
||||
?: DEFAULT_REALTIME_PERSISTENT_SESSION,
|
||||
)
|
||||
}
|
||||
.distinctUntilChanged()
|
||||
@@ -100,4 +112,8 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
|
||||
suspend fun setRealtimeTraceDetails(enabled: Boolean) {
|
||||
dataStore.edit { it[KEY_REALTIME_TRACE_DETAILS] = enabled }
|
||||
}
|
||||
|
||||
suspend fun setRealtimePersistentSession(enabled: Boolean) {
|
||||
dataStore.edit { it[KEY_REALTIME_PERSISTENT_SESSION] = enabled }
|
||||
}
|
||||
}
|
||||
|
||||
@@ -99,6 +99,24 @@ data class ServerCapabilities(
|
||||
}
|
||||
}
|
||||
|
||||
internal val HERMES_SKILL_ENDPOINTS = listOf("/v1/skills", "/api/skills")
|
||||
|
||||
internal fun parseSkillListBody(json: Json, body: String): List<SkillInfo>? {
|
||||
try {
|
||||
val parsed = json.decodeFromString<SkillListResponse>(body)
|
||||
val skills = parsed.skills ?: parsed.items ?: parsed.data
|
||||
if (skills != null) return skills
|
||||
} catch (_: Exception) {
|
||||
// Fall through to direct-array compatibility below.
|
||||
}
|
||||
|
||||
try {
|
||||
return json.decodeFromString<List<SkillInfo>>(body)
|
||||
} catch (_: Exception) {
|
||||
return null
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Direct HTTP/SSE client for the Hermes API Server.
|
||||
*
|
||||
@@ -342,27 +360,21 @@ class HermesApiClient(
|
||||
// --- Skills ---
|
||||
|
||||
suspend fun getSkills(): List<SkillInfo> = withContext(Dispatchers.IO) {
|
||||
try {
|
||||
val request = authRequest("$baseUrl/api/skills").get().build()
|
||||
client.newCall(request).execute().use { response ->
|
||||
if (!response.isSuccessful) return@withContext emptyList()
|
||||
val body = response.body?.string() ?: return@withContext emptyList()
|
||||
// Try structured response: { "skills": [...] } or { "items": [...] }
|
||||
try {
|
||||
val parsed = json.decodeFromString<SkillListResponse>(body)
|
||||
val skills = parsed.skills ?: parsed.items
|
||||
for (endpoint in HERMES_SKILL_ENDPOINTS) {
|
||||
try {
|
||||
val request = authRequest("$baseUrl$endpoint").get().build()
|
||||
client.newCall(request).execute().use { response ->
|
||||
if (!response.isSuccessful) return@use
|
||||
val body = response.body?.string() ?: return@use
|
||||
val skills = parseSkillListBody(json, body)
|
||||
if (skills != null) return@withContext skills
|
||||
} catch (_: Exception) { /* fall through */ }
|
||||
// Try direct array: [...]
|
||||
try {
|
||||
return@withContext json.decodeFromString<List<SkillInfo>>(body)
|
||||
} catch (_: Exception) { /* fall through */ }
|
||||
emptyList()
|
||||
}
|
||||
} catch (e: Exception) {
|
||||
Log.w(TAG, "Failed to fetch skills from $endpoint: ${e.message}")
|
||||
}
|
||||
} catch (e: Exception) {
|
||||
Log.w(TAG, "Failed to fetch skills: ${e.message}")
|
||||
emptyList()
|
||||
}
|
||||
|
||||
emptyList()
|
||||
}
|
||||
|
||||
// --- Server personalities ---
|
||||
|
||||
@@ -532,6 +532,64 @@ class RelayVoiceClient(
|
||||
label = "Realtime agent config update",
|
||||
)
|
||||
|
||||
/**
|
||||
* PATCH the ADR 33 background-run promotion settings. Only non-null fields
|
||||
* are sent; the relay echoes back the full [RealtimeVoiceConfig].
|
||||
*/
|
||||
suspend fun updateRealtimeAgentPromotion(
|
||||
promotionEnabled: Boolean? = null,
|
||||
promoteAfterMs: Int? = null,
|
||||
spokenHandoff: Boolean? = null,
|
||||
resultDelivery: String? = null,
|
||||
backgroundDefaultMode: String? = null,
|
||||
progressSpokenAfterMs: Int? = null,
|
||||
progressRepeatMs: Int? = null,
|
||||
maxBackgroundRuns: Int? = null,
|
||||
): Result<RealtimeVoiceConfig> = withContext(Dispatchers.IO) {
|
||||
val httpBase = resolveHttpBase()
|
||||
?: return@withContext Result.failure(IllegalStateException("Relay URL not configured"))
|
||||
val token = resolveBearerToken()
|
||||
if (token.isNullOrBlank()) {
|
||||
return@withContext Result.failure(missingAuthError())
|
||||
}
|
||||
val payload = buildJsonObject {
|
||||
promotionEnabled?.let { put("promotion_enabled", JsonPrimitive(it)) }
|
||||
promoteAfterMs?.let { put("promote_after_ms", JsonPrimitive(it)) }
|
||||
spokenHandoff?.let { put("spoken_handoff", JsonPrimitive(it)) }
|
||||
resultDelivery?.takeIf { it.isNotBlank() }?.let {
|
||||
put("result_delivery", JsonPrimitive(it.trim()))
|
||||
}
|
||||
backgroundDefaultMode?.takeIf { it.isNotBlank() }?.let {
|
||||
put("background_default_mode", JsonPrimitive(it.trim()))
|
||||
}
|
||||
progressSpokenAfterMs?.let { put("progress_spoken_after_ms", JsonPrimitive(it)) }
|
||||
progressRepeatMs?.let { put("progress_repeat_ms", JsonPrimitive(it)) }
|
||||
maxBackgroundRuns?.let { put("max_background_runs", JsonPrimitive(it)) }
|
||||
}
|
||||
val request = Request.Builder()
|
||||
.url(urlWithProfile("$httpBase/voice/realtime-agent/config"))
|
||||
.patch(payload.toString().toRequestBody(JSON_MEDIA_TYPE))
|
||||
.header("Authorization", "Bearer $token")
|
||||
.header("Accept", "application/json")
|
||||
.build()
|
||||
try {
|
||||
okHttpClient.newCall(request).execute().use { response ->
|
||||
if (!response.isSuccessful) {
|
||||
val body = response.body?.string().orEmpty()
|
||||
return@withContext Result.failure(
|
||||
IOException(describeHttpError(response.code, response.message, body))
|
||||
)
|
||||
}
|
||||
val raw = response.body?.string().orEmpty()
|
||||
Result.success(json.decodeFromString(RealtimeVoiceConfig.serializer(), raw))
|
||||
}
|
||||
} catch (e: IOException) {
|
||||
Result.failure(IOException("Realtime promotion update failed: ${e.message ?: "network error"}"))
|
||||
} catch (e: Exception) {
|
||||
Result.failure(e)
|
||||
}
|
||||
}
|
||||
|
||||
suspend fun getVoiceOutputConfig(): Result<VoiceOutputConfig> = withContext(Dispatchers.IO) {
|
||||
val httpBase = resolveHttpBase()
|
||||
?: return@withContext Result.failure(IllegalStateException("Relay URL not configured"))
|
||||
@@ -1133,6 +1191,22 @@ class RelayVoiceClient(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Run a Realtime Agent voice session.
|
||||
*
|
||||
* **One-shot (default):** with [turnInputs] null, this opens a session, sends
|
||||
* the single [inputPcm]/[prompt] turn, and completes + closes the socket on
|
||||
* `voice.response.done` — the historical per-utterance behavior used by the
|
||||
* Voice Lab and the one-shot fallback.
|
||||
*
|
||||
* **Persistent (ADR follow-up, see docs/plans/2026-05-24-realtime-persistent-session.md):**
|
||||
* with [turnInputs] non-null the socket stays open across turns. The first
|
||||
* turn is [inputPcm]/[prompt]; each subsequent [RealtimeTurnInput] read from
|
||||
* the channel is sent on the *same* socket, so the provider keeps a live
|
||||
* conversation. `voice.response.done` invokes [onTurnComplete] instead of
|
||||
* closing; the call returns only when the channel closes (voice-mode exit) or
|
||||
* a fatal socket/provider error occurs.
|
||||
*/
|
||||
suspend fun runRealtimeAgent(
|
||||
prompt: String,
|
||||
inputPcm: ByteArray,
|
||||
@@ -1140,8 +1214,11 @@ class RelayVoiceClient(
|
||||
chatSessionId: String? = null,
|
||||
conversationContext: List<RealtimeConversationContextMessage> = emptyList(),
|
||||
onHandoff: (VoiceHandoffEvent) -> Unit = {},
|
||||
turnInputs: kotlinx.coroutines.channels.ReceiveChannel<RealtimeTurnInput>? = null,
|
||||
onTurnComplete: (RealtimeVoiceSummary) -> Unit = {},
|
||||
onEvent: (RealtimeVoiceEvent, RealtimeAgentSessionControl) -> Unit,
|
||||
): Result<RealtimeVoiceSummary> = withContext(Dispatchers.IO) {
|
||||
val persistent = turnInputs != null
|
||||
val httpBase = resolveHttpBase()
|
||||
?: return@withContext Result.failure(IllegalStateException("Relay URL not configured"))
|
||||
val wsBase = resolveWebSocketBase()
|
||||
@@ -1172,8 +1249,11 @@ class RelayVoiceClient(
|
||||
val lastAudioEventId = AtomicLong(0L)
|
||||
val lastPlayedAudioEventId = AtomicLong(0L)
|
||||
val lastInputChunkId = AtomicLong(0L)
|
||||
val turnStartedAtMs = System.currentTimeMillis()
|
||||
val lastEventAtMs = AtomicLong(turnStartedAtMs)
|
||||
val turnStartedAtMs = AtomicLong(System.currentTimeMillis())
|
||||
val lastEventAtMs = AtomicLong(turnStartedAtMs.get())
|
||||
// True while a turn is awaiting its response. In persistent mode the idle
|
||||
// guard only applies while a turn is active; between-turn idle is normal.
|
||||
val activeTurn = AtomicBoolean(true)
|
||||
val inputChunks = buildList {
|
||||
var offset = 0
|
||||
var chunkId = 1L
|
||||
@@ -1184,8 +1264,33 @@ class RelayVoiceClient(
|
||||
offset = end
|
||||
}
|
||||
}
|
||||
// Highest input chunk id sent on this session. The relay dedups by
|
||||
// input_chunk_seq, so subsequent turns must continue past this.
|
||||
val sessionMaxChunkId = AtomicLong(inputChunks.size.toLong())
|
||||
var audioChunks = 0
|
||||
var audioBytes = 0
|
||||
|
||||
// Persistent-mode: chunk + send one more utterance on the open socket.
|
||||
fun sendTurnPcm(webSocket: WebSocket, pcm: ByteArray, sampleRate: Int) {
|
||||
var offset = 0
|
||||
var sentAny = false
|
||||
while (offset < pcm.size) {
|
||||
val end = (offset + REALTIME_INPUT_CHUNK_BYTES).coerceAtMost(pcm.size)
|
||||
val chunkId = sessionMaxChunkId.incrementAndGet()
|
||||
val encoded = Base64.getEncoder().encodeToString(pcm.copyOfRange(offset, end))
|
||||
webSocket.send(
|
||||
"""{"type":"input_audio.append","chunk_id":$chunkId,"sample_rate":$sampleRate,"audio_base64":"$encoded"}"""
|
||||
)
|
||||
sentAny = true
|
||||
offset = end
|
||||
}
|
||||
if (sentAny) {
|
||||
webSocket.send("""{"type":"input_audio.commit"}""")
|
||||
}
|
||||
turnStartedAtMs.set(System.currentTimeMillis())
|
||||
lastEventAtMs.set(System.currentTimeMillis())
|
||||
activeTurn.set(true)
|
||||
}
|
||||
fun activateSocket(webSocket: WebSocket, generation: Long): Boolean {
|
||||
while (true) {
|
||||
val activeGeneration = activeSocketGeneration.get()
|
||||
@@ -1332,24 +1437,28 @@ class RelayVoiceClient(
|
||||
inputChunkId = event.inputChunkId,
|
||||
)
|
||||
if (event.type == "voice.response.done") {
|
||||
if (completed.compareAndSet(false, true)) {
|
||||
finished.complete(
|
||||
Result.success(
|
||||
RealtimeVoiceSummary(
|
||||
provider = event.provider ?: session.provider,
|
||||
model = event.model ?: session.model,
|
||||
voice = event.voice ?: session.voice,
|
||||
sampleRate = session.sampleRate,
|
||||
audioChunks = audioChunks,
|
||||
audioBytes = audioBytes,
|
||||
firstAudioMs = event.firstAudioMs,
|
||||
responseDoneMs = event.responseDoneMs,
|
||||
eventLogPath = event.eventLogPath ?: session.eventLogPath,
|
||||
)
|
||||
)
|
||||
)
|
||||
val summary = RealtimeVoiceSummary(
|
||||
provider = event.provider ?: session.provider,
|
||||
model = event.model ?: session.model,
|
||||
voice = event.voice ?: session.voice,
|
||||
sampleRate = session.sampleRate,
|
||||
audioChunks = audioChunks,
|
||||
audioBytes = audioBytes,
|
||||
firstAudioMs = event.firstAudioMs,
|
||||
responseDoneMs = event.responseDoneMs,
|
||||
eventLogPath = event.eventLogPath ?: session.eventLogPath,
|
||||
)
|
||||
if (persistent) {
|
||||
// Turn boundary, not session boundary: keep the socket
|
||||
// open for the next utterance.
|
||||
activeTurn.set(false)
|
||||
onTurnComplete(summary)
|
||||
} else {
|
||||
if (completed.compareAndSet(false, true)) {
|
||||
finished.complete(Result.success(summary))
|
||||
}
|
||||
webSocket.close(1000, "done")
|
||||
}
|
||||
webSocket.close(1000, "done")
|
||||
} else if (event.type == "voice.error" || event.type == "voice.session.resume_failed") {
|
||||
completeFailure(event.message ?: "Realtime agent error")
|
||||
webSocket.close(1011, "provider error")
|
||||
@@ -1452,25 +1561,85 @@ class RelayVoiceClient(
|
||||
suspend fun awaitRealtimeAgentCompletion(): Result<RealtimeVoiceSummary> {
|
||||
while (true) {
|
||||
val now = System.currentTimeMillis()
|
||||
val turnElapsedMs = now - turnStartedAtMs
|
||||
val idleElapsedMs = now - lastEventAtMs.get()
|
||||
if (turnElapsedMs >= REALTIME_AGENT_MAX_TURN_MS) {
|
||||
throw IOException("Realtime agent exceeded the turn limit")
|
||||
// In persistent mode the turn/idle guards only apply while a turn
|
||||
// is actually in flight; between-turn idle is expected and must
|
||||
// not trip the stall timeout. The session ends when the turn
|
||||
// channel closes or a fatal error completes `finished`.
|
||||
val guardActive = !persistent || activeTurn.get()
|
||||
if (guardActive) {
|
||||
val turnElapsedMs = now - turnStartedAtMs.get()
|
||||
val idleElapsedMs = now - lastEventAtMs.get()
|
||||
if (turnElapsedMs >= REALTIME_AGENT_MAX_TURN_MS) {
|
||||
throw IOException("Realtime agent exceeded the turn limit")
|
||||
}
|
||||
if (idleElapsedMs >= REALTIME_AGENT_IDLE_TIMEOUT_MS) {
|
||||
throw IOException("Realtime agent stalled waiting for relay events")
|
||||
}
|
||||
val waitMs = minOf(
|
||||
REALTIME_AGENT_WAIT_SLICE_MS,
|
||||
REALTIME_AGENT_MAX_TURN_MS - turnElapsedMs,
|
||||
REALTIME_AGENT_IDLE_TIMEOUT_MS - idleElapsedMs,
|
||||
).coerceAtLeast(1L)
|
||||
withTimeoutOrNull(waitMs) {
|
||||
finished.await()
|
||||
}?.let { return it }
|
||||
} else {
|
||||
withTimeoutOrNull(REALTIME_AGENT_WAIT_SLICE_MS) {
|
||||
finished.await()
|
||||
}?.let { return it }
|
||||
}
|
||||
if (idleElapsedMs >= REALTIME_AGENT_IDLE_TIMEOUT_MS) {
|
||||
throw IOException("Realtime agent stalled waiting for relay events")
|
||||
}
|
||||
val waitMs = minOf(
|
||||
REALTIME_AGENT_WAIT_SLICE_MS,
|
||||
REALTIME_AGENT_MAX_TURN_MS - turnElapsedMs,
|
||||
REALTIME_AGENT_IDLE_TIMEOUT_MS - idleElapsedMs,
|
||||
).coerceAtLeast(1L)
|
||||
withTimeoutOrNull(waitMs) {
|
||||
finished.await()
|
||||
}?.let { return it }
|
||||
}
|
||||
}
|
||||
|
||||
// Persistent mode: drain further utterances and feed each onto the open
|
||||
// socket as a new turn. Closing the channel (voice-mode exit) ends the
|
||||
// session by completing `finished`.
|
||||
val turnReader: Job? = if (turnInputs != null) {
|
||||
launch {
|
||||
try {
|
||||
for (turn in turnInputs) {
|
||||
val ws = currentSocket.get() ?: continue
|
||||
if (turn.prompt.isNotBlank() && turn.inputPcm.isEmpty()) {
|
||||
ws.send(
|
||||
buildRealtimeResponseCreate(
|
||||
text = turn.prompt,
|
||||
toolScaffold = false,
|
||||
renderMode = "verbatim",
|
||||
)
|
||||
)
|
||||
turnStartedAtMs.set(System.currentTimeMillis())
|
||||
lastEventAtMs.set(System.currentTimeMillis())
|
||||
activeTurn.set(true)
|
||||
} else {
|
||||
sendTurnPcm(ws, turn.inputPcm, turn.sampleRate)
|
||||
}
|
||||
}
|
||||
} finally {
|
||||
// Channel closed -> end the persistent session cleanly.
|
||||
if (completed.compareAndSet(false, true)) {
|
||||
finished.complete(
|
||||
Result.success(
|
||||
RealtimeVoiceSummary(
|
||||
provider = session.provider,
|
||||
model = session.model,
|
||||
voice = session.voice,
|
||||
sampleRate = session.sampleRate,
|
||||
audioChunks = audioChunks,
|
||||
audioBytes = audioBytes,
|
||||
firstAudioMs = null,
|
||||
responseDoneMs = null,
|
||||
eventLogPath = session.eventLogPath,
|
||||
)
|
||||
)
|
||||
)
|
||||
}
|
||||
currentSocket.get()?.close(1000, "session ended")
|
||||
}
|
||||
}
|
||||
} else {
|
||||
null
|
||||
}
|
||||
|
||||
val socket = openSocket(resume = false)
|
||||
val routeWatcher = startRouteResumeWatcher(
|
||||
surface = "Realtime agent",
|
||||
@@ -1491,6 +1660,7 @@ class RelayVoiceClient(
|
||||
Result.failure(IOException(e.message ?: "Realtime agent timed out", e))
|
||||
} finally {
|
||||
routeWatcher?.cancel()
|
||||
turnReader?.cancel()
|
||||
}
|
||||
}
|
||||
|
||||
@@ -2080,6 +2250,8 @@ class RelayVoiceClient(
|
||||
eventLogPath = (obj["event_log_path"] as? JsonPrimitive)?.contentOrNull,
|
||||
firstAudioMs = (metrics?.get("first_audio_ms") as? JsonPrimitive)?.doubleOrNull,
|
||||
responseDoneMs = (metrics?.get("response_done_ms") as? JsonPrimitive)?.doubleOrNull,
|
||||
tier = (obj["tier"] as? JsonPrimitive)?.contentOrNull,
|
||||
floor = (obj["floor"] as? JsonPrimitive)?.contentOrNull,
|
||||
raw = raw,
|
||||
)
|
||||
} catch (e: Exception) {
|
||||
@@ -2154,6 +2326,28 @@ data class RealtimeVoiceConfig(
|
||||
val configScope: String? = null,
|
||||
@SerialName("fallback_to_global")
|
||||
val fallbackToGlobal: Boolean = false,
|
||||
/** ADR 33 background-run promotion settings. */
|
||||
val promotion: RealtimeVoicePromotion? = null,
|
||||
)
|
||||
|
||||
/** Wire shape of the `promotion` block on `GET /voice/realtime-agent/config`. */
|
||||
@Serializable
|
||||
data class RealtimeVoicePromotion(
|
||||
val enabled: Boolean = true,
|
||||
@SerialName("promote_after_ms")
|
||||
val promoteAfterMs: Int = 6000,
|
||||
@SerialName("background_default_mode")
|
||||
val backgroundDefaultMode: String = "promote",
|
||||
@SerialName("spoken_handoff")
|
||||
val spokenHandoff: Boolean = true,
|
||||
@SerialName("progress_spoken_after_ms")
|
||||
val progressSpokenAfterMs: Int = 15000,
|
||||
@SerialName("progress_repeat_ms")
|
||||
val progressRepeatMs: Int = 30000,
|
||||
@SerialName("result_delivery")
|
||||
val resultDelivery: String = "speak_when_idle",
|
||||
@SerialName("max_background_runs")
|
||||
val maxBackgroundRuns: Int = 1,
|
||||
)
|
||||
|
||||
@Serializable
|
||||
@@ -2392,6 +2586,9 @@ data class RealtimeVoiceEvent(
|
||||
val eventLogPath: String? = null,
|
||||
val firstAudioMs: Double? = null,
|
||||
val responseDoneMs: Double? = null,
|
||||
// ADR 33: background-run promotion fields.
|
||||
val tier: String? = null,
|
||||
val floor: String? = null,
|
||||
val raw: String,
|
||||
) {
|
||||
val isAudioDelta: Boolean
|
||||
@@ -2467,6 +2664,33 @@ data class RealtimeVoiceSummary(
|
||||
val eventLogPath: String?,
|
||||
)
|
||||
|
||||
/**
|
||||
* One utterance fed into a persistent Realtime Agent session
|
||||
* (see [RelayVoiceClient.runRealtimeAgent] persistent mode). A blank [inputPcm]
|
||||
* with a non-blank [prompt] sends a text turn; otherwise the PCM is chunked and
|
||||
* committed as a spoken turn.
|
||||
*/
|
||||
data class RealtimeTurnInput(
|
||||
val inputPcm: ByteArray,
|
||||
val sampleRate: Int = 16_000,
|
||||
val prompt: String = "",
|
||||
) {
|
||||
override fun equals(other: Any?): Boolean {
|
||||
if (this === other) return true
|
||||
if (other !is RealtimeTurnInput) return false
|
||||
return sampleRate == other.sampleRate &&
|
||||
prompt == other.prompt &&
|
||||
inputPcm.contentEquals(other.inputPcm)
|
||||
}
|
||||
|
||||
override fun hashCode(): Int {
|
||||
var result = inputPcm.contentHashCode()
|
||||
result = 31 * result + sampleRate
|
||||
result = 31 * result + prompt.hashCode()
|
||||
return result
|
||||
}
|
||||
}
|
||||
|
||||
data class VoiceOutputSummary(
|
||||
val provider: String,
|
||||
val model: String,
|
||||
|
||||
@@ -300,5 +300,6 @@ data class SkillInfo(
|
||||
@Serializable
|
||||
data class SkillListResponse(
|
||||
val skills: List<SkillInfo>? = null,
|
||||
val items: List<SkillInfo>? = null
|
||||
val items: List<SkillInfo>? = null,
|
||||
val data: List<SkillInfo>? = null
|
||||
)
|
||||
|
||||
@@ -301,6 +301,28 @@ fun VoiceModeOverlay(
|
||||
)
|
||||
}
|
||||
|
||||
AnimatedVisibility(
|
||||
visible = uiState.backgroundRun != null,
|
||||
enter = fadeIn(tween(140)),
|
||||
exit = fadeOut(tween(180)),
|
||||
) {
|
||||
Surface(
|
||||
shape = RoundedCornerShape(16.dp),
|
||||
color = MaterialTheme.colorScheme.secondaryContainer,
|
||||
modifier = Modifier
|
||||
.fillMaxWidth()
|
||||
.padding(horizontal = 24.dp, vertical = 4.dp),
|
||||
) {
|
||||
Text(
|
||||
text = uiState.backgroundRun?.message
|
||||
?: "Working on it in the background…",
|
||||
style = MaterialTheme.typography.labelMedium,
|
||||
color = MaterialTheme.colorScheme.onSecondaryContainer,
|
||||
modifier = Modifier.padding(horizontal = 16.dp, vertical = 10.dp),
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
Spacer(Modifier.height(8.dp))
|
||||
|
||||
DestructiveCountdownRow(
|
||||
|
||||
@@ -941,6 +941,26 @@ fun VoiceSettingsScreen(
|
||||
},
|
||||
)
|
||||
}
|
||||
Row(
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
verticalAlignment = Alignment.CenterVertically,
|
||||
horizontalArrangement = Arrangement.SpaceBetween,
|
||||
) {
|
||||
Column(modifier = Modifier.weight(1f)) {
|
||||
Text("Persistent session", style = MaterialTheme.typography.bodyLarge)
|
||||
Text(
|
||||
text = "Keep one provider conversation open across turns. Turn off to fall back to a fresh session per utterance.",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
}
|
||||
Switch(
|
||||
checked = voiceSettings.realtimePersistentSession,
|
||||
onCheckedChange = { enabled ->
|
||||
scope.launch { prefsRepo.setRealtimePersistentSession(enabled) }
|
||||
},
|
||||
)
|
||||
}
|
||||
Spacer(Modifier.height(4.dp))
|
||||
ProviderRow(
|
||||
label = "Status",
|
||||
@@ -975,6 +995,109 @@ fun VoiceSettingsScreen(
|
||||
ProviderRow(label = "Auth", value = realtimeAuthLabel(config))
|
||||
}
|
||||
|
||||
realtimeConfig?.promotion?.let { promo ->
|
||||
HorizontalDivider(modifier = Modifier.padding(vertical = 8.dp))
|
||||
Text("Background tasks", style = MaterialTheme.typography.titleSmall)
|
||||
Text(
|
||||
text = "Long Hermes tasks keep running in the background so the conversation stays responsive; the answer is spoken when it's ready.",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
Spacer(Modifier.height(8.dp))
|
||||
Row(
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
verticalAlignment = Alignment.CenterVertically,
|
||||
horizontalArrangement = Arrangement.SpaceBetween,
|
||||
) {
|
||||
Column(modifier = Modifier.weight(1f)) {
|
||||
Text("Promote long tasks", style = MaterialTheme.typography.bodyLarge)
|
||||
Text(
|
||||
text = "Detach a slow run after ${promo.promoteAfterMs} ms instead of waiting silently",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
}
|
||||
Switch(
|
||||
checked = promo.enabled,
|
||||
onCheckedChange = { enabled ->
|
||||
scope.launch {
|
||||
val client = voiceClient ?: return@launch
|
||||
val result = client.updateRealtimeAgentPromotion(
|
||||
promotionEnabled = enabled,
|
||||
)
|
||||
if (result.isSuccess) realtimeConfig = result.getOrNull()
|
||||
}
|
||||
},
|
||||
)
|
||||
}
|
||||
if (promo.enabled) {
|
||||
Spacer(Modifier.height(4.dp))
|
||||
Row(
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
verticalAlignment = Alignment.CenterVertically,
|
||||
horizontalArrangement = Arrangement.SpaceBetween,
|
||||
) {
|
||||
Column(modifier = Modifier.weight(1f)) {
|
||||
Text("Spoken handoff", style = MaterialTheme.typography.bodyLarge)
|
||||
Text(
|
||||
text = "Say a short \"I'm on it\" when a task moves to the background",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
}
|
||||
Switch(
|
||||
checked = promo.spokenHandoff,
|
||||
onCheckedChange = { enabled ->
|
||||
scope.launch {
|
||||
val client = voiceClient ?: return@launch
|
||||
val result = client.updateRealtimeAgentPromotion(
|
||||
spokenHandoff = enabled,
|
||||
)
|
||||
if (result.isSuccess) realtimeConfig = result.getOrNull()
|
||||
}
|
||||
},
|
||||
)
|
||||
}
|
||||
Spacer(Modifier.height(8.dp))
|
||||
Text(
|
||||
"When the answer is ready",
|
||||
style = MaterialTheme.typography.labelMedium,
|
||||
)
|
||||
Spacer(Modifier.height(4.dp))
|
||||
val deliveryOptions = listOf(
|
||||
"speak_when_idle",
|
||||
"notify_then_speak",
|
||||
"visual_only",
|
||||
)
|
||||
val deliveryLabels = listOf("Speak", "Notify", "Show only")
|
||||
SingleChoiceSegmentedButtonRow(modifier = Modifier.fillMaxWidth()) {
|
||||
deliveryOptions.forEachIndexed { index, option ->
|
||||
SegmentedButton(
|
||||
shape = SegmentedButtonDefaults.itemShape(
|
||||
index = index,
|
||||
count = deliveryOptions.size,
|
||||
),
|
||||
onClick = {
|
||||
scope.launch {
|
||||
val client = voiceClient ?: return@launch
|
||||
val result = client.updateRealtimeAgentPromotion(
|
||||
resultDelivery = option,
|
||||
)
|
||||
if (result.isSuccess) realtimeConfig = result.getOrNull()
|
||||
}
|
||||
},
|
||||
selected = option == promo.resultDelivery,
|
||||
) {
|
||||
Text(
|
||||
deliveryLabels[index],
|
||||
style = MaterialTheme.typography.labelSmall,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
HorizontalDivider(modifier = Modifier.padding(vertical = 8.dp))
|
||||
|
||||
Row(
|
||||
|
||||
@@ -17,6 +17,7 @@ import com.hermesandroid.relay.data.BargeInPreferencesRepository
|
||||
import com.hermesandroid.relay.data.BargeInSensitivity
|
||||
import com.hermesandroid.relay.data.ChatMessage
|
||||
import com.hermesandroid.relay.data.MessageRole
|
||||
import com.hermesandroid.relay.data.RealtimeConversationContextMessage
|
||||
import com.hermesandroid.relay.data.ToolCall
|
||||
import com.hermesandroid.relay.data.VoiceEngineMode
|
||||
import com.hermesandroid.relay.data.VoiceIntentTrace
|
||||
@@ -26,6 +27,8 @@ import com.hermesandroid.relay.diagnostics.DiagnosticsLog
|
||||
import com.hermesandroid.relay.network.ChannelMultiplexer
|
||||
import com.hermesandroid.relay.network.RelayVoiceClient
|
||||
import com.hermesandroid.relay.network.RealtimeAgentSessionControl
|
||||
import com.hermesandroid.relay.network.RealtimeTurnInput
|
||||
import com.hermesandroid.relay.network.RealtimeVoiceSummary
|
||||
import com.hermesandroid.relay.network.RealtimeVoiceEvent
|
||||
import com.hermesandroid.relay.network.VoiceHandoffEvent
|
||||
import com.hermesandroid.relay.network.handlers.LocalDispatchResult
|
||||
@@ -131,6 +134,21 @@ data class VoiceUiState(
|
||||
* fresh turn (mic tap), whichever comes first.
|
||||
*/
|
||||
val permissionDeniedCallout: PermissionDeniedCallout? = null,
|
||||
/**
|
||||
* ADR 33: non-null while a Hermes run has been promoted to (or started as) a
|
||||
* background task. Drives a persistent "working on it" chip so the user
|
||||
* knows a long task is still running after the spoken handoff. Cleared when
|
||||
* the background run completes, is cancelled, or errors.
|
||||
*/
|
||||
val backgroundRun: BackgroundRunState? = null,
|
||||
)
|
||||
|
||||
/** ADR 33 background/promoted Hermes run surface for the voice overlay. */
|
||||
data class BackgroundRunState(
|
||||
val runId: String? = null,
|
||||
/** "promoted" (auto-detached long run) or "durable" (explicit mode=background). */
|
||||
val tier: String = "promoted",
|
||||
val message: String = "Working on it in the background…",
|
||||
)
|
||||
|
||||
data class VoiceHandoffStatus(
|
||||
@@ -386,8 +404,28 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
private var voicePreferencesJob: Job? = null
|
||||
private var voiceEngineMode: VoiceEngineMode = VoiceEngineMode.HermesVoiceOutput
|
||||
private var realtimeTraceDetails: Boolean = false
|
||||
private var realtimePersistentSession: Boolean = true
|
||||
private var realtimeAgentControl: RealtimeAgentSessionControl? = null
|
||||
private var realtimeConfirmationControl: RealtimeAgentSessionControl? = null
|
||||
|
||||
// === Persistent realtime-agent session (one socket across turns) ===
|
||||
// The long-lived call to RelayVoiceClient.runRealtimeAgent(persistent) runs
|
||||
// in realtimeSessionJob; further utterances are fed on realtimeTurnChannel.
|
||||
// The per-turn event holders below are hoisted to fields so the single
|
||||
// session-lived event callback can serve every turn; submitRealtimeTurn /
|
||||
// the open path reset them at each turn boundary.
|
||||
private var realtimeSessionJob: Job? = null
|
||||
private var realtimeTurnChannel: kotlinx.coroutines.channels.Channel<RealtimeTurnInput>? = null
|
||||
private var rtUserText: String = ""
|
||||
private var rtAssistantMessageId: String = ""
|
||||
private var rtConversationContext: List<RealtimeConversationContextMessage> = emptyList()
|
||||
private val audioSeen = AtomicBoolean(false)
|
||||
private val audioBytes = AtomicInteger(0)
|
||||
private val bargeInStarted = AtomicBoolean(false)
|
||||
private val lastRealtimeAudioEventId = AtomicLong(0L)
|
||||
private var responseText = StringBuilder()
|
||||
private var inputTranscript = StringBuilder()
|
||||
private val spokenStatusKeys = mutableSetOf<String>()
|
||||
private var voiceRelayPreflight: (suspend () -> Result<Unit>)? = null
|
||||
|
||||
// === PHASE3-voice-intents: voice→bridge intent routing ===
|
||||
@@ -811,8 +849,18 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
"realtimeTraceDetails=${settings.realtimeTraceDetails}",
|
||||
)
|
||||
}
|
||||
// Switching engine away from Realtime Agent (or disabling the
|
||||
// persistent toggle) must drop any open persistent session.
|
||||
if (
|
||||
(voiceEngineMode == VoiceEngineMode.RealtimeAgent &&
|
||||
nextEngineMode != VoiceEngineMode.RealtimeAgent) ||
|
||||
(realtimePersistentSession && !settings.realtimePersistentSession)
|
||||
) {
|
||||
closeRealtimeSession()
|
||||
}
|
||||
voiceEngineMode = nextEngineMode
|
||||
realtimeTraceDetails = settings.realtimeTraceDetails
|
||||
realtimePersistentSession = settings.realtimePersistentSession
|
||||
val mode = when (settings.interactionMode.lowercase()) {
|
||||
"hold" -> InteractionMode.HoldToTalk
|
||||
"continuous" -> InteractionMode.Continuous
|
||||
@@ -1068,6 +1116,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
// Chime BEFORE teardown — AudioTrack release would cut it off otherwise.
|
||||
try { sfxPlayer?.playExit() } catch (_: Exception) { /* ignore */ }
|
||||
cancelRealtimeAgentTurn("exit voice mode")
|
||||
closeRealtimeSession()
|
||||
// B4: tear down the barge-in listener + timers before we kill the
|
||||
// player so AEC doesn't try to track a released audio session.
|
||||
stopBargeInListener()
|
||||
@@ -1740,6 +1789,31 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
responseText = "",
|
||||
)
|
||||
}
|
||||
if (realtimePersistentSession) {
|
||||
// Persistent conversation: one socket across turns. Open lazily on
|
||||
// the first turn (the long-lived call runs in realtimeSessionJob),
|
||||
// then feed further utterances on the channel. Fallback to one-shot
|
||||
// if the toggle is off (Voice Settings).
|
||||
if (realtimeSessionJob?.isActive == true && realtimeTurnChannel != null) {
|
||||
submitRealtimeTurn(chatVm, inputPcm, inputSampleRate)
|
||||
} else {
|
||||
closeRealtimeSession()
|
||||
realtimeTurnChannel = kotlinx.coroutines.channels.Channel(
|
||||
kotlinx.coroutines.channels.Channel.UNLIMITED,
|
||||
)
|
||||
realtimeSessionJob = viewModelScope.launch {
|
||||
runRealtimeAgentTurn(
|
||||
client = client,
|
||||
chatVm = chatVm,
|
||||
userText = "",
|
||||
inputPcm = inputPcm,
|
||||
inputSampleRate = inputSampleRate,
|
||||
persistentOpen = true,
|
||||
)
|
||||
}
|
||||
}
|
||||
return
|
||||
}
|
||||
runRealtimeAgentTurn(
|
||||
client = client,
|
||||
chatVm = chatVm,
|
||||
@@ -2004,6 +2078,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
userText: String,
|
||||
inputPcm: ByteArray,
|
||||
inputSampleRate: Int,
|
||||
persistentOpen: Boolean = false,
|
||||
) {
|
||||
providerRealtimeAgentTurnActive.set(true)
|
||||
streamObserverJob?.cancel()
|
||||
@@ -2031,19 +2106,23 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
)
|
||||
}
|
||||
|
||||
val conversationContext = chatVm.realtimeAgentContextMessages()
|
||||
val assistantMessageId = chatVm.startRealtimeAgentTurn(
|
||||
// Per-turn event state is hoisted to fields so one session-lived callback
|
||||
// serves every turn in persistent mode; reset them for this turn.
|
||||
audioSeen.set(false)
|
||||
audioBytes.set(0)
|
||||
bargeInStarted.set(false)
|
||||
lastRealtimeAudioEventId.set(0L)
|
||||
responseText = StringBuilder()
|
||||
inputTranscript = StringBuilder()
|
||||
spokenStatusKeys.clear()
|
||||
rtUserText = userText
|
||||
rtConversationContext = chatVm.realtimeAgentContextMessages()
|
||||
rtAssistantMessageId = chatVm.startRealtimeAgentTurn(
|
||||
userText = userText,
|
||||
chatSessionId = chatVm.currentSessionId.value,
|
||||
)
|
||||
val conversationContext = rtConversationContext
|
||||
val pcmPlayer = realtimePcmPlayer
|
||||
val audioSeen = AtomicBoolean(false)
|
||||
val audioBytes = AtomicInteger(0)
|
||||
val bargeInStarted = AtomicBoolean(false)
|
||||
val lastRealtimeAudioEventId = AtomicLong(0L)
|
||||
val responseText = StringBuilder()
|
||||
val inputTranscript = StringBuilder()
|
||||
val spokenStatusKeys = mutableSetOf<String>()
|
||||
|
||||
fun emitStatus(
|
||||
key: String,
|
||||
@@ -2105,10 +2184,12 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
chatSessionId = chatVm.currentSessionId.value,
|
||||
conversationContext = conversationContext,
|
||||
onHandoff = ::recordVoiceHandoff,
|
||||
turnInputs = if (persistentOpen) realtimeTurnChannel else null,
|
||||
onTurnComplete = { summary -> onRealtimeTurnComplete(summary) },
|
||||
) { event, control ->
|
||||
realtimeAgentControl = control
|
||||
chatVm.applyRealtimeAgentEvent(
|
||||
assistantMessageId = assistantMessageId,
|
||||
assistantMessageId = rtAssistantMessageId,
|
||||
event = event,
|
||||
showDetailedTrace = realtimeTraceDetails,
|
||||
)
|
||||
@@ -2119,13 +2200,13 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
it.copy(
|
||||
state = VoiceState.Listening,
|
||||
outputAudioActive = false,
|
||||
transcribedText = inputTranscript.toString().ifBlank { userText },
|
||||
transcribedText = inputTranscript.toString().ifBlank { rtUserText },
|
||||
)
|
||||
}
|
||||
}
|
||||
"voice.input_transcript.final" -> {
|
||||
inputTranscript.clear()
|
||||
inputTranscript.append(event.text ?: userText)
|
||||
inputTranscript.append(event.text ?: rtUserText)
|
||||
_uiState.update {
|
||||
it.copy(
|
||||
state = VoiceState.Thinking,
|
||||
@@ -2135,6 +2216,13 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
}
|
||||
}
|
||||
"voice.response.started", "hermes.run.started" -> {
|
||||
if (event.type == "hermes.run.started") {
|
||||
Log.i(
|
||||
TAG,
|
||||
"Realtime Hermes run started run=${event.runId ?: "?"} " +
|
||||
"session=${event.chatSessionId ?: "?"}",
|
||||
)
|
||||
}
|
||||
_uiState.update {
|
||||
it.copy(state = VoiceState.Thinking, outputAudioActive = false)
|
||||
}
|
||||
@@ -2172,6 +2260,12 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
}
|
||||
"hermes.run.progress" -> {
|
||||
val line = realtimeHermesProgressLine(event.message)
|
||||
Log.i(
|
||||
TAG,
|
||||
"Realtime Hermes progress tier=${event.tier ?: "?"} " +
|
||||
"floor=${event.floor ?: "?"} status=${event.statusKey ?: "?"} " +
|
||||
"shouldSpeak=${event.shouldSpeak} run=${event.runId ?: "?"}",
|
||||
)
|
||||
if (event.shouldSpeak) {
|
||||
emitStatus(
|
||||
key = "progress-${event.runId ?: "run"}-${event.statusKey ?: line}",
|
||||
@@ -2188,6 +2282,40 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
)
|
||||
}
|
||||
}
|
||||
"hermes.run.promoted" -> {
|
||||
// The run detached to the background; the provider speaks the
|
||||
// handoff. Surface a persistent chip so the user knows a long
|
||||
// task is still in flight (ADR 33 Tier B/C).
|
||||
val tier = event.tier ?: "promoted"
|
||||
Log.i(
|
||||
TAG,
|
||||
"Realtime run promoted to background tier=$tier " +
|
||||
"run=${event.runId ?: "?"} session=${event.chatSessionId ?: "?"}",
|
||||
)
|
||||
_uiState.update {
|
||||
it.copy(
|
||||
backgroundRun = BackgroundRunState(
|
||||
runId = event.runId,
|
||||
tier = tier,
|
||||
message = if (tier == "durable") {
|
||||
"Started a background task — I'll report back."
|
||||
} else {
|
||||
"This is taking a moment — working on it in the background."
|
||||
},
|
||||
),
|
||||
)
|
||||
}
|
||||
}
|
||||
"hermes.run.background_completed" -> {
|
||||
// Background run finished; the spoken summary follows via the
|
||||
// provider's forced-summary turn. Clear the chip.
|
||||
Log.i(
|
||||
TAG,
|
||||
"Realtime background run completed run=${event.runId ?: "?"} " +
|
||||
"ok=${event.success != false}",
|
||||
)
|
||||
_uiState.update { it.copy(backgroundRun = null) }
|
||||
}
|
||||
"hermes.confirmation.requested" -> {
|
||||
val confirmationId = event.confirmationId
|
||||
if (!confirmationId.isNullOrBlank()) {
|
||||
@@ -2269,6 +2397,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
}
|
||||
}
|
||||
"hermes.run.cancelled" -> {
|
||||
Log.i(TAG, "Realtime Hermes run cancelled run=${event.runId ?: "?"}")
|
||||
realtimeConfirmationControl = null
|
||||
_uiState.update {
|
||||
it.copy(
|
||||
@@ -2276,6 +2405,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
outputAudioActive = false,
|
||||
responseText = "Cancelled.",
|
||||
hermesConfirmation = null,
|
||||
backgroundRun = null,
|
||||
)
|
||||
}
|
||||
}
|
||||
@@ -2294,19 +2424,22 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
}
|
||||
"voice.error" -> {
|
||||
realtimeConfirmationControl = null
|
||||
val rawDetail = event.message ?: "Realtime agent failed"
|
||||
DiagnosticsLog.record(
|
||||
category = DiagnosticCategory.Voice,
|
||||
severity = DiagnosticSeverity.Error,
|
||||
title = "Realtime voice error",
|
||||
detail = event.message ?: "Realtime agent failed",
|
||||
detail = rawDetail,
|
||||
)
|
||||
_uiState.update { it.copy(hermesConfirmation = null) }
|
||||
// Route the raw relay message through the classifier so
|
||||
// provider-auth (and other) failures surface a clear,
|
||||
// actionable message + a Voice-settings snackbar action
|
||||
// instead of a raw "xAI Realtime auth ..." provider string.
|
||||
surfaceError(
|
||||
java.io.IOException(rawDetail),
|
||||
context = "voice_config",
|
||||
)
|
||||
_uiState.update {
|
||||
it.copy(
|
||||
state = VoiceState.Error,
|
||||
error = event.message ?: "Realtime agent failed",
|
||||
hermesConfirmation = null,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -2335,9 +2468,78 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
title = "Realtime voice turn failed",
|
||||
detail = err?.message ?: "Unknown error",
|
||||
)
|
||||
// A persistent session ending in error must drop so the next turn opens
|
||||
// a fresh one rather than submitting into a dead channel.
|
||||
closeRealtimeSession()
|
||||
surfaceError(err, context = "voice_config")
|
||||
}
|
||||
|
||||
/**
|
||||
* Per-turn completion in a persistent realtime session (ADR follow-up). The
|
||||
* socket stays open; this just finalizes the spoken turn and re-arms
|
||||
* continuous listening if enabled.
|
||||
*/
|
||||
private fun onRealtimeTurnComplete(summary: RealtimeVoiceSummary) {
|
||||
DiagnosticsLog.record(
|
||||
category = DiagnosticCategory.Voice,
|
||||
severity = DiagnosticSeverity.Info,
|
||||
title = "Realtime voice turn complete",
|
||||
)
|
||||
Log.i(
|
||||
TAG,
|
||||
"Realtime persistent turn complete provider=${summary.provider} " +
|
||||
"audioChunks=${summary.audioChunks}",
|
||||
)
|
||||
pendingInTtsQueue.set(0)
|
||||
maybeAutoResume()
|
||||
}
|
||||
|
||||
/**
|
||||
* Submit a follow-up utterance onto the already-open persistent session
|
||||
* (turns 2+). Mirrors the per-turn reset the open path does for turn 1.
|
||||
*/
|
||||
private fun submitRealtimeTurn(chatVm: ChatViewModel, inputPcm: ByteArray, inputSampleRate: Int) {
|
||||
val channel = realtimeTurnChannel ?: return
|
||||
drainQueuedLocalTts()
|
||||
try { player?.stop() } catch (_: Exception) { /* ignore */ }
|
||||
firstFrameWatchdogJob?.cancel(); firstFrameWatchdogJob = null
|
||||
clearSpokenChunksState()
|
||||
audioSeen.set(false)
|
||||
audioBytes.set(0)
|
||||
bargeInStarted.set(false)
|
||||
lastRealtimeAudioEventId.set(0L)
|
||||
responseText = StringBuilder()
|
||||
inputTranscript = StringBuilder()
|
||||
spokenStatusKeys.clear()
|
||||
rtUserText = ""
|
||||
rtConversationContext = chatVm.realtimeAgentContextMessages()
|
||||
rtAssistantMessageId = chatVm.startRealtimeAgentTurn(
|
||||
userText = "",
|
||||
chatSessionId = chatVm.currentSessionId.value,
|
||||
)
|
||||
providerRealtimeAgentTurnActive.set(true)
|
||||
_uiState.update {
|
||||
it.copy(
|
||||
state = VoiceState.Thinking,
|
||||
outputAudioActive = false,
|
||||
transcribedText = null,
|
||||
responseText = "",
|
||||
)
|
||||
}
|
||||
val sent = channel.trySend(RealtimeTurnInput(inputPcm, inputSampleRate)).isSuccess
|
||||
Log.i(TAG, "Realtime persistent turn submitted sent=$sent pcmBytes=${inputPcm.size}")
|
||||
}
|
||||
|
||||
/** Tear down the persistent realtime session (voice-mode exit / engine switch). */
|
||||
private fun closeRealtimeSession() {
|
||||
val hadSession = realtimeTurnChannel != null || realtimeSessionJob != null
|
||||
realtimeTurnChannel?.close()
|
||||
realtimeTurnChannel = null
|
||||
realtimeSessionJob?.cancel()
|
||||
realtimeSessionJob = null
|
||||
if (hadSession) Log.i(TAG, "Realtime persistent session closed")
|
||||
}
|
||||
|
||||
private fun realtimeToolStatusLine(toolName: String?): String {
|
||||
val normalized = toolName.orEmpty().lowercase()
|
||||
return when {
|
||||
@@ -3899,6 +4101,7 @@ class VoiceViewModel(application: Application) : AndroidViewModel(application) {
|
||||
|
||||
override fun onCleared() {
|
||||
super.onCleared()
|
||||
closeRealtimeSession()
|
||||
// B4: release the listener + VAD engine native resources before
|
||||
// anything else. stopBargeInListener is defensive / idempotent.
|
||||
stopBargeInListener()
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
package com.hermesandroid.relay.network
|
||||
|
||||
import com.hermesandroid.relay.network.models.SkillListResponse
|
||||
import kotlinx.serialization.json.Json
|
||||
import org.junit.Assert.assertEquals
|
||||
import org.junit.Assert.assertFalse
|
||||
import org.junit.Assert.assertNotEquals
|
||||
@@ -220,6 +222,46 @@ class HermesApiClientTest {
|
||||
assertEquals("http://localhost:8642/v1/models", url)
|
||||
}
|
||||
|
||||
// --- Skills endpoint compatibility ---
|
||||
|
||||
@Test
|
||||
fun skillEndpointOrder_prefersUpstreamV1ThenLegacyApiFallback() {
|
||||
assertEquals(listOf("/v1/skills", "/api/skills"), HERMES_SKILL_ENDPOINTS)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun skillListResponse_parsesUpstreamV1DataEnvelope() {
|
||||
val body = """
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{"name": "android", "description": "Control phone", "category": "android"}
|
||||
]
|
||||
}
|
||||
""".trimIndent()
|
||||
|
||||
val parsed = Json { ignoreUnknownKeys = true }.decodeFromString<SkillListResponse>(body)
|
||||
|
||||
assertEquals(1, parsed.data?.size)
|
||||
assertEquals("android", parsed.data?.first()?.name)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun parseSkillListBody_acceptsUpstreamV1DataEnvelope() {
|
||||
val body = """
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{"name": "android", "description": "Control phone", "category": "android"}
|
||||
]
|
||||
}
|
||||
""".trimIndent()
|
||||
|
||||
val parsed = parseSkillListBody(Json { ignoreUnknownKeys = true }, body)
|
||||
|
||||
assertEquals(listOf("android"), parsed?.map { it.name })
|
||||
}
|
||||
|
||||
@Test
|
||||
fun urlConstruction_deleteSession() {
|
||||
val baseUrl = "http://localhost:8642"
|
||||
|
||||
@@ -259,8 +259,9 @@ GET /api/memory?target=memory // or target=user
|
||||
|
||||
### Skills
|
||||
```
|
||||
GET /api/skills # optional ?category= filter
|
||||
GET /api/skills/{name}
|
||||
GET /v1/skills # native upstream list metadata
|
||||
GET /api/skills # legacy fallback, optional ?category= filter
|
||||
GET /api/skills/{name} # legacy detail route when provided by fork/bootstrap
|
||||
```
|
||||
> `/api/skills/categories` was removed from upstream as dead code (commit 8d023e43) and is not re-injected by the bootstrap.
|
||||
|
||||
@@ -272,7 +273,8 @@ Probe endpoints to detect what's available:
|
||||
GET /health -> basic connectivity
|
||||
GET /api/sessions -> enhanced Hermes session API
|
||||
GET /v1/models -> model listing (OpenAI-compatible)
|
||||
GET /api/skills -> skills support
|
||||
GET /v1/skills -> native skills support
|
||||
GET /api/skills -> legacy skills fallback
|
||||
GET /api/memory -> memory support
|
||||
GET /api/config -> config API
|
||||
|
||||
|
||||
+211
-26
@@ -80,15 +80,19 @@ The app supports two streaming endpoints, selectable in Settings:
|
||||
| **Sessions** (`/api/sessions/{id}/chat/stream`) | Inline text annotations (`` `💻 terminal` ``) — client parses from markdown | Hermes-native SSE (assistant.delta, tool.progress, etc.) or OpenAI-format (delta.content) |
|
||||
| **Runs** (`/v1/runs` + `/v1/runs/{run_id}/events`) | **Structured events** (tool.started, tool.completed) — real-time tool cards | Hermes lifecycle events (message.delta, tool.started, tool.completed, run.completed) |
|
||||
|
||||
**Important upstream note:** The `/api/sessions` CRUD endpoints are moving
|
||||
toward upstream Hermes core through focused PR
|
||||
[#29302](https://github.com/NousResearch/hermes-agent/pull/29302), which covers
|
||||
**Important upstream note:** The `/api/sessions` CRUD/chat endpoints landed in
|
||||
upstream Hermes core via PR
|
||||
[#33134](https://github.com/NousResearch/hermes-agent/pull/33134) / commit
|
||||
[`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde),
|
||||
which salvaged the focused PR
|
||||
[#29302](https://github.com/NousResearch/hermes-agent/pull/29302). It covers
|
||||
session list/create/read/update/delete, messages, fork, chat, and chat stream.
|
||||
Until that reaches a released core build, `hermes_relay_bootstrap/` still ships
|
||||
with the plugin and runs at Python interpreter startup via `.pth`. The bootstrap
|
||||
now composes with partial upstream support: native routes win per method/path,
|
||||
and the relay only injects missing compatibility gaps such as config, skills, or
|
||||
memory when core does not provide them.
|
||||
`hermes_relay_bootstrap/` still ships with the plugin and runs at Python
|
||||
interpreter startup via `.pth` for older cores and for compatibility gaps that
|
||||
upstream still does not provide. The bootstrap composes with partial upstream
|
||||
support: native routes win per method/path, and the relay only injects missing
|
||||
compatibility gaps such as search, config, legacy skills detail routes,
|
||||
available-models, or memory when core does not provide them.
|
||||
|
||||
The app's `probeCapabilities()` returns a per-endpoint snapshot, and `ConnectionViewModel.resolveStreamingEndpoint()` collapses `streamingEndpoint = "auto"` (the default for new installs) to a concrete `"sessions"` or `"runs"` choice based on what the server actually exposes.
|
||||
|
||||
@@ -499,11 +503,14 @@ Adopting from ARC's workflow patterns:
|
||||
### 16. Runtime API Server Patch via .pth Bootstrap (2026-04-12)
|
||||
|
||||
**Context:** The Android app depends on API-server routes for session history,
|
||||
profile/config metadata, skills, and memory-backed UI. Upstream core is now
|
||||
moving in focused pieces rather than one large frontend API patch: PR
|
||||
[#29302](https://github.com/NousResearch/hermes-agent/pull/29302) covers the
|
||||
canonical `/api/sessions/*` surface, while config/skills/memory still remain
|
||||
compatibility routes in this repo until core exposes stable equivalents. Without
|
||||
profile/config metadata, skills, and memory-backed UI. Upstream core has started
|
||||
absorbing the surface in focused pieces rather than one large frontend API patch:
|
||||
PR [#33134](https://github.com/NousResearch/hermes-agent/pull/33134) / commit
|
||||
[`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde)
|
||||
now covers the canonical `/api/sessions/*` surface, and `/v1/skills` covers
|
||||
skill list metadata. Config, legacy skill detail routes, memory, search, and
|
||||
provider-aware model metadata still remain compatibility routes in this repo
|
||||
until core exposes stable equivalents or clients migrate away from them. Without
|
||||
the bootstrap, users on older vanilla upstream builds lose session browsing,
|
||||
metadata-backed settings, and history-on-restart behavior.
|
||||
|
||||
@@ -520,16 +527,17 @@ We considered four options:
|
||||
1. **Zero modifications to hermes-agent's filesystem.** `git pull` / `hermes update` see no local changes, so they always work cleanly. The patch lives entirely in `hermes_relay_bootstrap/` inside our own repo.
|
||||
2. **Single-file containment of all ported logic.** `_handlers.py` is 500 lines of straight-line aiohttp handler code with explicit `adapter` parameters (closures, not bound methods). Easy to audit, easy to delete.
|
||||
3. **Feature detection by method/path, not broad route family.** Native upstream
|
||||
routes win one method/path at a time. This matters because PR #29302 can land
|
||||
`/api/sessions/*` before core has stable config/skills/memory APIs; the
|
||||
bootstrap must not skip those remaining compatibility routes just because a
|
||||
sessions route exists.
|
||||
routes win one method/path at a time. This matters because PR #33134 landed
|
||||
`/api/sessions/*` before core had stable config/legacy skill detail/memory/search
|
||||
APIs; the bootstrap must not skip those remaining compatibility routes just
|
||||
because a sessions or `/v1/skills` route exists.
|
||||
4. **Trust model already established.** The user installed our plugin into their hermes-agent venv. They've already consented to having the plugin import hermes-agent internals (it does this for relay tools, voice endpoints, media registry). Monkey-patching `aiohttp.web.Application` is in the same trust bucket.
|
||||
5. **Surface-by-surface removal.** When PR #29302 or equivalent reaches a
|
||||
released hermes-agent version, the sessions compatibility routes should go
|
||||
quiet automatically. Config, skills, memory, and command preprocessing remain
|
||||
until their native replacements exist. Full bootstrap deletion happens only
|
||||
after every compatibility route group has a stable core equivalent.
|
||||
5. **Surface-by-surface removal.** When a supported hermes-agent baseline includes
|
||||
PR #33134 / `f7527b0`, the sessions compatibility routes should go quiet
|
||||
automatically. Config, legacy skills detail routes, memory, search,
|
||||
available-models, and command preprocessing remain until their native
|
||||
replacements exist or the clients migrate. Full bootstrap deletion happens
|
||||
only after every compatibility route group has a stable core equivalent.
|
||||
6. **`/v1/runs` is genuinely better for chat than `/api/sessions/{id}/chat/stream`.** It's standard upstream, supports `X-Hermes-Session-Id` for continuation, and emits live structured tool events. The fork's chat handler exists because upstream didn't HAVE this clean structured-event runs API at the time the fork was cut — but upstream does now.
|
||||
|
||||
**The Android client adapts via `streamingEndpoint = "auto"`.** New `ServerCapabilities` data class returned by `HermesApiClient.probeCapabilities()` captures per-endpoint presence (`sessionsApi`, `sessionsChatStream`, `runs`, `portable`, `healthy`). `ConnectionViewModel.resolveStreamingEndpoint()` collapses `"auto"` to `"sessions"` (when chat-stream handler is present, i.e. fork or upstream-merged) or `"runs"` (otherwise, i.e. bootstrap-injected vanilla upstream). The setting still supports manual `"sessions"` / `"runs"` overrides for debugging.
|
||||
@@ -548,10 +556,11 @@ We considered four options:
|
||||
- `install.sh` step 2 — copies the `.pth` into the venv site-packages
|
||||
|
||||
**Removal path** is now per surface:
|
||||
1. Sessions: after PR #29302 or equivalent ships in released core, keep the
|
||||
1. Sessions: with PR #33134 / `f7527b0` in the supported core baseline, keep the
|
||||
bootstrap installed but verify it skips native `/api/sessions/*` routes.
|
||||
2. Config/skills/memory: remove those compatibility handlers only after stable
|
||||
core APIs exist and Android probes prefer them.
|
||||
2. Config/legacy skills/memory/search/available-models: remove those
|
||||
compatibility handlers only after stable core APIs exist and Android probes
|
||||
prefer them, or after clients migrate away from the legacy route group.
|
||||
3. Slash middleware: remove after native API-server slash preprocessing exists.
|
||||
4. Full cleanup: delete `hermes_relay_bootstrap/`, delete
|
||||
`hermes_relay_bootstrap.pth`, remove the `.pth` install block, and update
|
||||
@@ -928,7 +937,7 @@ The plan called for adding a `// VOICE HOOK` callback to `ChatViewModel` so `Voi
|
||||
- `plugin/relay/server.py` — route registration alongside `/media/*`
|
||||
- `plugin/tests/test_voice_routes.py` — 14 unit tests, `unittest`-based (pytest conftest issue documented in `CLAUDE.md`)
|
||||
- `app/src/main/kotlin/.../audio/VoiceRecorder.kt` — MediaRecorder amplitude StateFlow
|
||||
- `app/src/main/kotlin/.../audio/VoicePlayer.kt` — MediaPlayer + Visualizer amplitude StateFlow with OEM fallback
|
||||
- `app/src/main/kotlin/.../audio/VoicePlayer.kt` — Media3 ExoPlayer (gapless TTS queue) + Visualizer amplitude StateFlow with OEM fallback; `audioSessionId` served from a thread-safe `@Volatile` cache (read off-main by barge-in)
|
||||
- `app/src/main/kotlin/.../network/RelayVoiceClient.kt` — OkHttp multipart + JSON clients
|
||||
- `app/src/main/kotlin/.../viewmodel/VoiceViewModel.kt` — turn state machine, sentence detection, TTS queue consumer, `ChatViewModel` observation pattern
|
||||
- `app/src/main/kotlin/.../ui/components/VoiceModeOverlay.kt` — full-screen overlay, VoiceState→SphereState mapping, three interaction modes
|
||||
@@ -1525,3 +1534,179 @@ old STT -> Hermes -> TTS pipeline instead of a native realtime agent.
|
||||
- `plugin/relay/realtime_agent/broker.py`
|
||||
- `plugin/relay/realtime_agent/providers/xai.py`
|
||||
- `plugin/relay/realtime_agent/providers/openai.py`
|
||||
|
||||
---
|
||||
|
||||
## ADR 33 - Realtime Agent foregrounds short Hermes turns and promotes long ones to background tasks
|
||||
|
||||
**Status:** Accepted, phased (2026-05-24). Default-on at feature GA, gated by a
|
||||
prerequisite provider-idle-tolerance spike (Phase 0). Supersedes the
|
||||
blocking-broker assumption inside ADR 32's sequence; ADR 32's preamble ->
|
||||
forced-Hermes -> provider-summary shape is preserved for the foreground tier.
|
||||
|
||||
**Context.** Today a Realtime Agent turn runs Hermes *synchronously inside the
|
||||
provider event pump*: `_pump_provider_events` awaits `_handle_provider_tool_call`
|
||||
-> `_run_brokered_tool`, which streams the entire Hermes SSE run to completion
|
||||
before the pump consumes the next provider event (`broker.py`). This is correct
|
||||
and lowest-latency for short Q&A, but it has two costs that grow with task
|
||||
length:
|
||||
|
||||
- The provider realtime socket sits attached-but-idle for the whole run (a live,
|
||||
billed, audio-clocked WebSocket parked for tens of seconds during research,
|
||||
multi-tool, or desktop/build tasks).
|
||||
- The tool surface already advertises a background vocabulary
|
||||
(`hermes_run_task`, `hermes_get_status`, `hermes_cancel`, `hermes_confirm`)
|
||||
and a `hermes_run_status` state machine, but `hermes_get_status` /
|
||||
`hermes_cancel` as *provider* tool calls are unreachable mid-run because the
|
||||
pump is parked. Only the client->relay `response.cancel` path can interrupt.
|
||||
|
||||
The blocking `await` is also an *implicit mutex*: it serializes the three audio
|
||||
sources that can produce `voice.output_audio.delta` (see "Who speaks" below) so
|
||||
they never overlap. Removing it requires replacing that mutex with an explicit
|
||||
floor owner.
|
||||
|
||||
**Who speaks (confirmed against both supported providers).** Up to three mouths
|
||||
exist; only one is active per phase today because of the blocking await:
|
||||
|
||||
1. **Realtime provider** - xAI `grok-voice-latest`, OpenAI `gpt-realtime-2`.
|
||||
Both run with `turn_detection: None` (relay owns turn boundaries; no server
|
||||
VAD) and audio output modality. Speaks the pre-Hermes acknowledgement and the
|
||||
final post-result summary.
|
||||
2. **Relay TTS render** - `xai_tts`/`openai_tts` via `_render_provider_audio`.
|
||||
A separate, non-realtime synthesizer used as the forced-summary fallback and
|
||||
the legacy render path. Emits the *same* `voice.output_audio.delta` wire event
|
||||
as the provider, so Android cannot distinguish them.
|
||||
3. **Android local TTS** - the `should_speak` long-wait filler, driven by
|
||||
`hermes.run.progress`. Client-side, not the provider.
|
||||
|
||||
Because all three converge on one Android `AudioTrack`, **floor arbitration must
|
||||
happen relay-side, before bytes reach the socket.**
|
||||
|
||||
**Decision.** Keep three turn classes; make promotion automatic and default-on,
|
||||
with the relay as the single explicit floor owner.
|
||||
|
||||
- **Tier A - Foreground (short Q&A).** Unchanged from ADR 32: provider preamble
|
||||
-> relay-forced Hermes (still awaited) -> provider summary. Lowest latency,
|
||||
trivial floor. This remains the path for any turn that completes inside the
|
||||
promotion grace window.
|
||||
- **Tier B - Promoted (long task detected late).** Start in Tier A. If the
|
||||
Hermes run has not produced a final result within a grace window
|
||||
(`promote_after_ms`, default ~6000ms, tunable), the relay *detaches* the run
|
||||
from the pump: the run continues as a tracked `asyncio.Task`, the pump resumes
|
||||
consuming provider events, and the provider speaks a short "I've started that -
|
||||
I'll let you know" handoff. Progress continues via `hermes.run.progress`. When
|
||||
the background run completes, the relay injects the result as a tool result and
|
||||
requests a provider summary at the next floor-idle moment.
|
||||
- **Tier C - Explicitly durable.** `hermes_run_task(mode="background")` returns a
|
||||
run handle immediately (no grace window). For tasks the model/profile knows up
|
||||
front are long (research, builds via desktop tools, multi-step). Same
|
||||
completion-injection path as Tier B.
|
||||
|
||||
Promotion is the default behavior, not a flag. Grace-period promotion preserves
|
||||
Tier A latency for the common case and only forks when a run actually proves
|
||||
long, so the user never has to pick a mode.
|
||||
|
||||
**The relay is the single floor owner.** A per-session floor state
|
||||
(`idle | provider_speaking | hermes_filler | result_pending`) gates every audio
|
||||
source:
|
||||
|
||||
- Only one mouth may hold the floor. The provider holds it by default.
|
||||
- A completed background result does **not** barge in. It is queued as
|
||||
`result_pending` and spoken only when the floor returns to `idle` (provider
|
||||
finished, user not mid-utterance). Provider VAD being off means the relay
|
||||
controls `response.create`, so it can withhold the summary until the floor is
|
||||
clear.
|
||||
- Android local filler (`should_speak`) is suppressed whenever the provider holds
|
||||
the floor; it is a Tier B/C long-wait affordance only.
|
||||
- Relay TTS render (mouth 2) may only fire when it owns the floor and the
|
||||
provider has drained, exactly as the forced-summary fallback does today.
|
||||
|
||||
**Settings (ample, per ADR intent).** Surface in Voice Settings -> Realtime
|
||||
Agent, with relay-side `realtime_voice` config as source of truth and per-profile
|
||||
override:
|
||||
|
||||
- `promotion_enabled` (default true) - master switch; false pins Tier A blocking.
|
||||
- `promote_after_ms` (default ~6000) - grace window before Tier B handoff.
|
||||
- `background_default_mode` - whether ambiguous long turns prefer promote vs.
|
||||
stay-foreground.
|
||||
- `spoken_handoff` (default true) - speak the "I've started that" line on
|
||||
promotion vs. silent + visual only.
|
||||
- `progress_spoken_after_ms` / `progress_repeat_ms` - reuse existing
|
||||
`_HERMES_SPOKEN_PROGRESS_*` knobs, now configurable.
|
||||
- `result_delivery` - `speak_when_idle` (default) vs. `notify_then_speak`
|
||||
(chime/visual, speak on user re-engage) vs. `visual_only`.
|
||||
- `max_background_runs` - concurrent background runs per session (default 1 for
|
||||
the MVP; the existing single-`hermes_task` field assumes 1).
|
||||
|
||||
**Protocol additions (relay <-> Android, additive).**
|
||||
|
||||
- `hermes.run.promoted` - run moved to background; carries `run_id`,
|
||||
`promote_after_ms`, `spoken_handoff`.
|
||||
- `hermes.run.background_completed` - background run finished; precedes the
|
||||
provider/relay summary.
|
||||
- Extend `hermes.run.progress` with `tier` and `floor` so the client can render
|
||||
background state distinctly (e.g. a persistent "working on: ..." chip).
|
||||
- `hermes_get_status` / `hermes_cancel` become genuinely reachable as provider
|
||||
tool calls in Tier B/C because the pump is no longer parked; no schema change.
|
||||
|
||||
**Prerequisite (Phase 0 spike) — RESOLVED 2026-05-24.** The spike asked how xAI
|
||||
and OpenAI realtime sessions behave when held open and idle. Verdicts (see
|
||||
`docs/realtime-voice-poc.md`): **OpenAI `hold-floor-ok` (empirical** — survived
|
||||
10/20/30s idle with clean post-idle audio); **xAI `hold-floor-ok`** (shipping
|
||||
Realtime Agent already holds `xai_realtime` sessions open across between-turn
|
||||
idle with `turn_detection: None` + resume TTL; relay-host probe retained as a
|
||||
regression check, not a precondition). The premise was also superseded in
|
||||
implementation: Tier B closes the pending provider call with an interim ack
|
||||
rather than holding an open response, so the socket only sees the normal
|
||||
between-turns idle gap — no provider needs the `must-reopen` fallback today, and
|
||||
default-on is unblocked.
|
||||
|
||||
**Rules.**
|
||||
|
||||
- Hermes remains the only path for tools, memory, current data, research, side
|
||||
effects, durable context, and confirmations (unchanged from ADR 29/32).
|
||||
- Exactly one audio source may hold the floor at a time; the relay enforces this
|
||||
before audio reaches Android. Background results never barge in.
|
||||
- Android must not read raw Hermes output aloud; spoken output is always a
|
||||
provider (or relay-fallback) summary of the compact Hermes result.
|
||||
- Promotion must be cancel-safe: a promoted/background run is cancelable via both
|
||||
the provider `hermes_cancel` tool and the client `response.cancel` path, reusing
|
||||
`_cancel_active_hermes`.
|
||||
- A turn must never strand: if a background run completes while the WS is
|
||||
detached, the result is replayed through the existing event-ring/resume path on
|
||||
reattach.
|
||||
|
||||
**Open questions / risks.**
|
||||
|
||||
- Floor arbitration is the hard part and the existing `native_forced_*` /
|
||||
`_should_forward_provider_response_event` suppression machinery is already the
|
||||
most fragile code in `broker.py`. Concurrency stresses it most; the floor-owner
|
||||
state machine should *replace* ad-hoc suppression, not stack on top of it.
|
||||
- Concurrent background runs (`max_background_runs > 1`) are out of scope for the
|
||||
MVP; the session model assumes a single `hermes_task`.
|
||||
- "Notify then speak" result delivery needs an Android affordance (chime +
|
||||
chip + tap-to-hear) that does not yet exist.
|
||||
|
||||
**Phased rollout.**
|
||||
|
||||
1. **Phase 0** - provider-idle-tolerance spike (above). Gate for default-on.
|
||||
2. **Phase 1** - introduce the relay floor-owner state machine under the current
|
||||
blocking behavior (no functional change), with tests that prove single-floor
|
||||
invariants.
|
||||
3. **Phase 2** - Tier B grace-period promotion behind `promotion_enabled`,
|
||||
default **off**, validated on long-task transcripts.
|
||||
4. **Phase 3** - flip `promotion_enabled` default **on**; add Tier C
|
||||
`mode="background"`; ship the Voice Settings surface.
|
||||
|
||||
**Key Files:**
|
||||
- `docs/decisions.md` (this ADR; supersedes ADR 32's blocking assumption)
|
||||
- `docs/plans/2026-05-24-realtime-background-hermes-runs.md` (companion plan)
|
||||
- `plugin/relay/realtime_agent/broker.py`
|
||||
- `plugin/relay/realtime_agent/hermes_tool_broker.py`
|
||||
- `plugin/relay/realtime_agent/models.py`
|
||||
- `plugin/relay/realtime_agent/providers/xai.py`
|
||||
- `plugin/relay/realtime_agent/providers/openai.py`
|
||||
- `plugin/relay/profile_voice.py`
|
||||
- `docs/relay-protocol.md`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/viewmodel/VoiceViewModel.kt`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/ui/screens/VoiceSettingsScreen.kt`
|
||||
|
||||
@@ -0,0 +1,330 @@
|
||||
# Plan: Realtime Agent background Hermes runs (foreground -> promote -> durable)
|
||||
|
||||
> **Purpose.** Stop the Realtime Agent's provider event pump from blocking on
|
||||
> the full Hermes run. Keep short turns synchronous and low-latency, but let
|
||||
> genuinely long Hermes runs detach to a tracked background task so the provider
|
||||
> session stays responsive and the run can be monitored, summarized on
|
||||
> completion, and cancelled.
|
||||
>
|
||||
> **Decision of record.** ADR 33 (`docs/decisions.md`). This plan is the
|
||||
> phased, file-level execution of that ADR. Read ADR 33 first — it defines the
|
||||
> three tiers, the relay-as-single-floor-owner rule, the settings surface, and
|
||||
> the Phase 0 prerequisite spike.
|
||||
>
|
||||
> **Non-negotiable.** Hermes remains the only authority for tools, memory,
|
||||
> current data, research, side effects, durable context, and confirmations
|
||||
> (ADR 29/32). This plan changes *when and how* a Hermes run is awaited and
|
||||
> spoken — never whether Hermes governs it.
|
||||
|
||||
## Current State (grounded in code)
|
||||
|
||||
The realtime-agent broker already runs a persistent provider event pump; the
|
||||
gap is that tool execution blocks it.
|
||||
|
||||
- `plugin/relay/realtime_agent/broker.py`
|
||||
- `_pump_provider_events` (~`broker.py:1124`) is the long-lived
|
||||
`native_provider_task` — the persistent loop. It survives WS detach/resume.
|
||||
- `FUNCTION_CALL_COMPLETED` -> `_handle_provider_tool_call` (~`broker.py:1628`)
|
||||
-> `_run_brokered_tool` (~`broker.py:1930`, `return await task`) ->
|
||||
`_execute_brokered_tool` (~`broker.py:1956`) `async for`s the entire Hermes
|
||||
SSE stream before the pump consumes the next provider event. **This is the
|
||||
blocking await to remove for long runs.**
|
||||
- `_send_hermes_run_progress` (~`broker.py:2245`) already runs concurrently as
|
||||
`progress_task`, emitting `hermes.run.progress` every
|
||||
`_HERMES_PROGRESS_INTERVAL_SECONDS` and spoken filler after
|
||||
`_HERMES_SPOKEN_PROGRESS_AFTER_SECONDS`. Reuse this for background progress.
|
||||
- `_cancel_active_hermes` (~`broker.py:2295`) already cancels `hermes_task`
|
||||
cross-coroutine. Promotion must keep this working.
|
||||
- `RealtimeAgentSession` (~`broker.py:114`) holds `hermes_task`,
|
||||
`hermes_run_status`, and the `native_forced_*` suppression flags. Floor state
|
||||
will live here.
|
||||
- `_render_provider_audio` (~`broker.py:2509`) is the relay TTS render mouth
|
||||
(forced-summary fallback). It must obey the floor owner.
|
||||
- `_request_playback_drain` / `_send_provider_tool_result` /
|
||||
`_request_provider_response` (~`broker.py:1700`-`1810`) are the existing
|
||||
result-injection primitives the background-completion path will reuse.
|
||||
- `plugin/relay/realtime_agent/hermes_tool_broker.py`
|
||||
- `_TOOL_SURFACE = ("hermes_run_task", "hermes_get_status", "hermes_cancel",
|
||||
"hermes_confirm")` (~`:189`) — background vocabulary already advertised.
|
||||
- `stream_task` yields the run; unchanged by this plan.
|
||||
- `plugin/relay/realtime_agent/models.py` — `SERVER_EVT_*`, `CLIENT_MSG_*`,
|
||||
`HERMES_TOOL_SCHEMAS`. New events/fields added here.
|
||||
- Providers (`providers/xai.py`, `providers/openai.py`) — both run
|
||||
`turn_detection: None` (relay owns `response.create`) and emit audio. This is
|
||||
what makes "withhold the summary until the floor is idle" possible.
|
||||
- Settings: `plugin/relay/profile_voice.py` (`realtime_voice_settings`,
|
||||
`save_profile_voice_section`) + `plugin/relay/config.py`.
|
||||
- Android: `viewmodel/VoiceViewModel.kt`, `network/RelayVoiceClient.kt`,
|
||||
`data/VoicePreferences.kt`, `data/RealtimeConversationContext.kt`,
|
||||
`voice/RealtimeTurnSyncBuilder.kt`, `ui/screens/VoiceSettingsScreen.kt`.
|
||||
- Tests: `plugin/tests/test_realtime_agent_routes.py`,
|
||||
`test_realtime_agent_xai_provider.py`, `test_realtime_agent_openai_provider.py`.
|
||||
|
||||
### Three audio sources ("who speaks")
|
||||
|
||||
All three converge on Android's single `AudioTrack` via `voice.output_audio.delta`:
|
||||
|
||||
1. Realtime provider (xAI/OpenAI) — primary.
|
||||
2. Relay TTS render (`_render_provider_audio`) — fallback; same wire event.
|
||||
3. Android local TTS — `should_speak` filler on `hermes.run.progress`.
|
||||
|
||||
Today the blocking await serializes them. This plan replaces that implicit mutex
|
||||
with an **explicit relay-side floor owner**.
|
||||
|
||||
## Target Architecture
|
||||
|
||||
```text
|
||||
provider event pump (never blocks on a long run)
|
||||
-> short run: await inline (Tier A, unchanged ADR 32 shape)
|
||||
-> long run: detach to tracked task, resume pump, speak handoff (Tier B)
|
||||
-> explicit durable: return handle immediately (Tier C)
|
||||
|
|
||||
v
|
||||
FloorOwner(idle | provider_speaking | hermes_filler | result_pending)
|
||||
|
|
||||
background run completes -> queue result_pending -> speak when floor idle
|
||||
```
|
||||
|
||||
**FloorOwner contract (relay-side, per session):**
|
||||
- Exactly one mouth holds the floor; provider holds it by default.
|
||||
- A completed background result is queued `result_pending`, never barges in. It
|
||||
is spoken (provider summary, or relay-TTS fallback) only on transition to
|
||||
`idle`.
|
||||
- Android `should_speak` filler is suppressed while `provider_speaking`.
|
||||
- `_render_provider_audio` may only fire when it owns the floor and the provider
|
||||
has drained.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Do not change Tier A latency or the ADR 32 preamble -> forced-Hermes ->
|
||||
provider-summary shape for short turns.
|
||||
- Do not give the provider authority over tools/memory/confirmations.
|
||||
- Do not support `max_background_runs > 1` in this plan (single `hermes_task`).
|
||||
- Do not require adb/device testing for verification unless Bailey asks. JVM/py
|
||||
unit tests plus the lab smoke gate cover acceptance.
|
||||
- Do not depend on Hermes upstream changes; the broker already owns the loop.
|
||||
|
||||
## Protocol Additions (additive only)
|
||||
|
||||
Relay -> client (add to `models.py`):
|
||||
- `hermes.run.promoted` — `{run_id, promote_after_ms, spoken_handoff, tier}`.
|
||||
- `hermes.run.background_completed` — `{run_id, ok, tool_count}` (precedes summary).
|
||||
- Extend `hermes.run.progress` with `tier` (`foreground|promoted|durable`) and
|
||||
`floor` (`idle|provider_speaking|hermes_filler|result_pending`).
|
||||
|
||||
No new client->relay messages required — `response.cancel`, `hermes_cancel`, and
|
||||
`hermes_get_status` already exist; the latter two become reachable mid-run once
|
||||
the pump no longer blocks.
|
||||
|
||||
## Settings Surface (per ADR 33)
|
||||
|
||||
Relay-side `realtime_voice` config (source of truth, per-profile override via
|
||||
`profile_voice.py`), mirrored into Android `VoicePreferences`:
|
||||
|
||||
| Key | Default | Meaning |
|
||||
|---|---|---|
|
||||
| `promotion_enabled` | `true` (Phase 3) | Master switch; `false` pins Tier A blocking |
|
||||
| `promote_after_ms` | `6000` | Grace window before Tier B handoff |
|
||||
| `background_default_mode` | `promote` | Ambiguous long turn: `promote` vs `foreground` |
|
||||
| `spoken_handoff` | `true` | Speak "I've started that" on promotion |
|
||||
| `progress_spoken_after_ms` | `15000` | Reuse `_HERMES_SPOKEN_PROGRESS_AFTER_SECONDS` |
|
||||
| `progress_repeat_ms` | `30000` | Reuse `_HERMES_SPOKEN_PROGRESS_REPEAT_SECONDS` |
|
||||
| `result_delivery` | `speak_when_idle` | vs `notify_then_speak` / `visual_only` |
|
||||
| `max_background_runs` | `1` | Fixed at 1 this plan |
|
||||
|
||||
---
|
||||
|
||||
## Phase 0 — Provider idle-tolerance spike (GATES default-on)
|
||||
|
||||
**Goal.** Determine empirically whether xAI and OpenAI realtime sessions tolerate
|
||||
being held open and quiescent (no `response.create`, no input audio) for
|
||||
30–120s, since Tier B holds the floor while a background run completes.
|
||||
|
||||
**Tasks.**
|
||||
1. Add a throwaway script under `plugin/tools/` or a `scripts/` probe (not
|
||||
shipped) that opens each provider realtime socket via the existing adapters,
|
||||
sends `session.update`, then idles 30/60/120s and logs: socket survival, any
|
||||
server-side timeout/close codes, VAD/turn artifacts on the first post-idle
|
||||
`response.create`, and any keep-alive requirement.
|
||||
2. Record findings in `docs/realtime-voice-poc.md` under a new "Idle tolerance"
|
||||
section, per provider.
|
||||
|
||||
**Acceptance.**
|
||||
- A documented per-provider verdict: `hold-floor-ok` | `needs-keepalive` |
|
||||
`must-reopen`. This verdict selects the Tier B fallback strategy per provider.
|
||||
|
||||
**Note.** If a provider is `must-reopen`, Tier B for that provider detaches the
|
||||
run but closes/reopens the provider socket (or keeps a minimal keep-alive)
|
||||
rather than holding it conversational. Capture that in the ADR's Phase 0 line.
|
||||
|
||||
**Status: DONE (2026-05-24).** Verdicts recorded in
|
||||
`docs/realtime-voice-poc.md` → Idle tolerance:
|
||||
- **OpenAI `gpt-realtime-2` — `hold-floor-ok` (empirical).** Probe ran live
|
||||
(10s/20s/30s idle windows survived; clean post-idle audio each time).
|
||||
- **xAI `grok-voice-latest` — `hold-floor-ok`.** No xAI creds on the dev box, but
|
||||
the verdict is not conditional: the shipping Realtime Agent already holds
|
||||
`xai_realtime` sessions open across between-turn idle gaps
|
||||
(`turn_detection: None` + resume TTL) with no idle-close reports. A relay-host
|
||||
probe run is retained as a regression check, not a precondition.
|
||||
|
||||
**Premise superseded.** Phase 2/3 implemented Tier B by *closing the pending
|
||||
provider call with an interim ack* rather than holding an open response, so the
|
||||
"hold the floor conversational while a run completes" worst case this spike
|
||||
guarded against does not occur — the socket only sees the normal between-turns
|
||||
idle gap. Both providers are `hold-floor-ok`, so the default-on gate is satisfied
|
||||
(no provider needs the `must-reopen` fallback today).
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — Introduce FloorOwner under current blocking behavior (no functional change)
|
||||
|
||||
**Goal.** Add the explicit floor state machine and route all three mouths through
|
||||
it, while preserving today's exact audible behavior. This is the de-risking step.
|
||||
|
||||
**Tasks.**
|
||||
1. Add `RealtimeFloor` (new dataclass/enum) to `broker.py` or a new
|
||||
`realtime_agent/floor.py`: state `idle|provider_speaking|hermes_filler|
|
||||
result_pending`, with `acquire(mouth)`, `release(mouth)`, and
|
||||
`can_speak(mouth) -> bool`. Pure, unit-testable, no I/O.
|
||||
2. Hold a `floor: RealtimeFloor` on `RealtimeAgentSession`.
|
||||
3. Gate the three emit paths through `floor.can_speak(...)`:
|
||||
- provider audio in `_pump_provider_events` (`AUDIO_DELTA` -> `provider`),
|
||||
- `_render_provider_audio` (-> `relay_tts`),
|
||||
- `should_speak` progress events (-> `android_filler`; suppress while
|
||||
`provider_speaking`).
|
||||
4. Transition floor on `RESPONSE_STARTED`/`AUDIO_DONE`/`RESPONSE_DONE` and on
|
||||
playback-drain completion.
|
||||
5. Stamp `floor` onto `hermes.run.progress` (additive field).
|
||||
|
||||
**Acceptance.**
|
||||
- All existing `plugin/tests/test_realtime_agent_*` pass unchanged.
|
||||
- New `test_realtime_floor.py` proves single-mouth invariants:
|
||||
background-result-never-barges, filler-suppressed-while-provider-speaks,
|
||||
relay-TTS-only-when-owned.
|
||||
- Manually/log-verified: a short turn sounds identical to pre-change.
|
||||
|
||||
**Test gate.** `python -m unittest plugin.tests.test_realtime_floor` +
|
||||
existing realtime agent tests green.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — Tier B grace-period promotion (default OFF)
|
||||
|
||||
**Goal.** Detach a long Hermes run from the pump after `promote_after_ms`; resume
|
||||
the pump; speak a handoff; deliver the result on completion via the floor owner.
|
||||
|
||||
**Tasks.**
|
||||
1. `models.py`: add `hermes.run.promoted`, `hermes.run.background_completed`,
|
||||
`tier` field; add settings keys to the config schema.
|
||||
2. `profile_voice.py` / `config.py`: read/write the new `realtime_voice` keys
|
||||
with defaults above (`promotion_enabled=false` this phase).
|
||||
3. `broker.py` — split `_run_brokered_tool` for `hermes_run_task`:
|
||||
- Start the run task as today, but `await asyncio.wait({task}, timeout=
|
||||
promote_after_ms)`.
|
||||
- If done in time: Tier A path, unchanged.
|
||||
- If not: emit `hermes.run.promoted`, optionally speak handoff (gated by
|
||||
`spoken_handoff` + floor), set `tier="promoted"`, **return control to the
|
||||
pump** without awaiting the task. Keep `session.hermes_task` set.
|
||||
4. Add a pump-side drain point: between provider events (and on a small timer),
|
||||
check for a finished background task; when finished, emit
|
||||
`hermes.run.background_completed`, then inject via the existing
|
||||
`_send_provider_tool_result` -> `_request_provider_response` path **only when
|
||||
`floor` is `idle`** (else mark `result_pending` and inject on next `idle`).
|
||||
5. Preserve cancellation: `hermes_cancel` (now reachable) and `response.cancel`
|
||||
both route to `_cancel_active_hermes`.
|
||||
6. Resume safety: if WS detaches mid-background-run, the completion event must
|
||||
replay through the existing event-ring/resume path on reattach.
|
||||
|
||||
**Acceptance.**
|
||||
- With `promotion_enabled=true` in test config, a run that exceeds
|
||||
`promote_after_ms` emits `hermes.run.promoted`, the pump processes a subsequent
|
||||
provider event before the run finishes (proves non-blocking), and the result is
|
||||
spoken exactly once after completion.
|
||||
- A run under the window behaves as Tier A (no `promoted` event).
|
||||
- Cancel during a promoted run stops it and emits `hermes.run.cancelled`.
|
||||
- Detach+resume during a promoted run replays `background_completed`.
|
||||
|
||||
**Test gate.** New `test_realtime_promotion.py` (fake provider connection + fake
|
||||
Hermes broker with controllable latency) covering: under-window, over-window,
|
||||
cancel-promoted, detach-resume-completes, result-waits-for-idle-floor.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — Default-on + Tier C durable + Android settings UI
|
||||
|
||||
**Goal.** Flip `promotion_enabled` default to `true` (gated on Phase 0 verdict),
|
||||
add explicit `mode="background"`, and expose settings on Android.
|
||||
|
||||
**Tasks.**
|
||||
1. `models.py` `HERMES_TOOL_SCHEMAS`: document/validate `hermes_run_task.mode`
|
||||
(`run|background`); `background` returns a handle immediately (skip grace
|
||||
window, go straight to Tier B detach).
|
||||
2. Default `promotion_enabled=true`; per-provider Tier B strategy from Phase 0.
|
||||
3. Android `VoicePreferences.kt` + `RelayVoiceClient.kt`: read/write the new
|
||||
settings; surface in `VoiceSettingsScreen.kt` under Realtime Agent (promotion
|
||||
toggle, grace slider, handoff toggle, result-delivery picker).
|
||||
4. Android `VoiceViewModel.kt` + `RealtimeTurnSyncBuilder.kt`: handle
|
||||
`hermes.run.promoted` / `hermes.run.background_completed`; render a persistent
|
||||
"working on: …" chip while `tier=promoted/durable`; implement
|
||||
`notify_then_speak` affordance (chime + tap-to-hear) if that mode is selected.
|
||||
5. Docs: `docs/relay-protocol.md` (new events/fields), `user-docs/features/
|
||||
voice.md` (settings + behavior), `CHANGELOG.md` `[Unreleased]`.
|
||||
|
||||
**Acceptance.**
|
||||
- Default config promotes long runs without user action; short runs unaffected.
|
||||
- `hermes_run_task(mode="background")` returns immediately and completes via the
|
||||
same path.
|
||||
- Android shows promoted state and speaks the result per `result_delivery`.
|
||||
- `./gradlew lint` clean; py tests green; lab smoke
|
||||
(`scripts/realtime-voice-lab-smoke.ps1`) still under threshold for short turns.
|
||||
|
||||
**Test gate.** Full `plugin.tests.test_realtime_agent_*` + new floor/promotion
|
||||
suites; Android unit tests for the new `VoicePreferences`/sync mapping;
|
||||
`./gradlew lint`.
|
||||
|
||||
---
|
||||
|
||||
## Risks & Mitigations
|
||||
|
||||
- **Floor logic stacking on `native_forced_*` suppression.** The forced-summary
|
||||
state machine is the most fragile code in `broker.py`. Mitigation: Phase 1
|
||||
introduces FloorOwner *as the serializer*; migrate suppression decisions to ask
|
||||
the floor rather than adding parallel flags.
|
||||
- **Double-speak (provider + relay TTS).** Mitigation: both gated by
|
||||
`floor.can_speak`; result injection only on `idle`.
|
||||
- **Provider idle close.** Mitigation: Phase 0 verdict picks per-provider Tier B
|
||||
strategy; `must-reopen` providers don't hold the floor.
|
||||
- **Stranded background result on disconnect.** Mitigation: reuse event-ring +
|
||||
resume replay; covered by Phase 2 acceptance.
|
||||
|
||||
## Test Strategy Summary
|
||||
|
||||
- Pure unit: `RealtimeFloor` invariants (Phase 1).
|
||||
- Broker integration with fakes: promotion timing, cancel, resume, idle-gated
|
||||
delivery (Phase 2).
|
||||
- Android: settings round-trip + event handling (Phase 3).
|
||||
- Regression: existing realtime agent suites unchanged throughout; lab smoke for
|
||||
short-turn latency.
|
||||
|
||||
## Key Files
|
||||
|
||||
- `docs/decisions.md` (ADR 33)
|
||||
- `plugin/relay/realtime_agent/broker.py`
|
||||
- `plugin/relay/realtime_agent/floor.py` (new)
|
||||
- `plugin/relay/realtime_agent/hermes_tool_broker.py`
|
||||
- `plugin/relay/realtime_agent/models.py`
|
||||
- `plugin/relay/realtime_agent/providers/xai.py`
|
||||
- `plugin/relay/realtime_agent/providers/openai.py`
|
||||
- `plugin/relay/profile_voice.py`
|
||||
- `plugin/relay/config.py`
|
||||
- `plugin/tests/test_realtime_floor.py` (new)
|
||||
- `plugin/tests/test_realtime_promotion.py` (new)
|
||||
- `docs/relay-protocol.md`
|
||||
- `docs/realtime-voice-poc.md` (Phase 0 findings)
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/viewmodel/VoiceViewModel.kt`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/voice/RealtimeTurnSyncBuilder.kt`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/data/VoicePreferences.kt`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/network/RelayVoiceClient.kt`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/ui/screens/VoiceSettingsScreen.kt`
|
||||
- `user-docs/features/voice.md`
|
||||
- `CHANGELOG.md`
|
||||
@@ -0,0 +1,83 @@
|
||||
# Plan: Persistent realtime-agent session (one socket across turns)
|
||||
|
||||
> **Purpose.** Make Realtime Agent voice a *persistent conversation* instead of a
|
||||
> new session per utterance. Today each turn calls
|
||||
> `RelayVoiceClient.runRealtimeAgent()` which `createRealtimeAgentSession()`s a
|
||||
> fresh relay session + provider socket and closes it on `voice.response.done`.
|
||||
> The provider therefore has no live memory of prior turns (only relay-seeded
|
||||
> context snippets), and every turn pays session-setup latency.
|
||||
>
|
||||
> **Key finding.** This is a **client-only** change. The relay already supports
|
||||
> many turns on one socket — `_handle_provider_native_ws` loops over
|
||||
> `input_audio.append` / `input_audio.commit` / `response.create` and only tears
|
||||
> the session down when the WebSocket disconnects. `voice.response.done` is a
|
||||
> *turn* boundary, not a *session* boundary. The client is what closes the socket
|
||||
> on `voice.response.done` (`RelayVoiceClient.kt:1410`).
|
||||
|
||||
## Design
|
||||
|
||||
Keep one relay session + provider socket open for the lifetime of a Realtime
|
||||
Agent voice-mode session; feed each utterance as a new turn on that socket.
|
||||
|
||||
### RelayVoiceClient
|
||||
- Add an opt-in **persistent mode** to `runRealtimeAgent` via two optional params
|
||||
(one-shot path is byte-for-byte unchanged when they're null):
|
||||
- `turnInputs: ReceiveChannel<RealtimeTurnInput>?` — when non-null, the call is
|
||||
a long-lived session: after the first turn it reads further turns off the
|
||||
channel and sends their `input_audio.append`+`commit` on the open socket.
|
||||
- `onTurnComplete: (RealtimeVoiceSummary) -> Unit` — invoked at each
|
||||
`voice.response.done` instead of completing+closing.
|
||||
- In persistent mode:
|
||||
- `voice.response.done` → call `onTurnComplete`, **do not** close the socket or
|
||||
complete `finished`.
|
||||
- A reader coroutine drains `turnInputs` and sends each turn's PCM chunks.
|
||||
- `finished` completes only when the channel closes (voice-mode exit) or on a
|
||||
fatal socket/provider error.
|
||||
- The per-turn idle/turn-limit guards in `awaitRealtimeAgentCompletion` are
|
||||
scoped to an *active* turn only — between-turn idle is expected and must not
|
||||
trip `REALTIME_AGENT_IDLE_TIMEOUT_MS`.
|
||||
- `RealtimeTurnInput(inputPcm, sampleRate, prompt)` data class.
|
||||
|
||||
### VoiceViewModel
|
||||
- Hold a `realtimeLiveSession` (the persistent call's `Job` + the
|
||||
`SendChannel<RealtimeTurnInput>`), opened lazily on the first Realtime Agent
|
||||
turn and reused for subsequent turns.
|
||||
- Per utterance: instead of a fresh `runRealtimeAgentTurn`, do the per-turn UI
|
||||
setup (state reset, watchdog, `chatVm.startRealtimeAgentTurn`) and
|
||||
`realtimeLiveSession.submit(RealtimeTurnInput(...))`.
|
||||
- Close the session (`channel.close()` + cancel job) on `exitVoiceMode`, engine
|
||||
switch away from Realtime Agent, and `onCleared`.
|
||||
- The continuous event callback (the `when(event.type)` block) is registered once
|
||||
for the session and handles every turn's events.
|
||||
|
||||
### Fallback flag (risk control)
|
||||
- `VoicePreferences.realtimePersistentSession` (default **true**). When false,
|
||||
fall back to the current one-shot `runRealtimeAgentTurn` path. Lets the user
|
||||
flip back on-device without a rebuild if the persistent path misbehaves.
|
||||
|
||||
## Non-Goals
|
||||
- No relay changes (the relay already supports multi-turn sockets).
|
||||
- No change to the one-shot path used by the Voice Lab / stable engine.
|
||||
- Not changing barge-in semantics beyond keeping them working per turn.
|
||||
|
||||
## Risks / must-validate-on-device
|
||||
- Feeding new input while a prior turn's audio is still draining (barge-in vs.
|
||||
new turn). Mitigation: reuse existing barge-in/cancel before submitting a turn.
|
||||
- Per-turn UI state resets must not tear down the shared session.
|
||||
- Between-turn idle must not trip the stall timeout.
|
||||
- Provider/relay session longevity across long pauses (the socket stays open, so
|
||||
no resume-TTL dependency — but confirm providers don't idle-close; see ADR 33
|
||||
Phase 0 `hold-floor-ok`).
|
||||
|
||||
## Test strategy
|
||||
- Unit-test the pure pieces: `RealtimeTurnInput` chunking, the persistent-vs-
|
||||
one-shot branch decision, idle-guard gating by active-turn.
|
||||
- On-device (post-merge, by Bailey): multi-turn conversation with follow-up
|
||||
references ("what did you just say?"), background promotion mid-conversation,
|
||||
barge-in, exit/re-enter voice mode, engine switch.
|
||||
|
||||
## Key Files
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/network/RelayVoiceClient.kt`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/viewmodel/VoiceViewModel.kt`
|
||||
- `app/src/main/kotlin/com/hermesandroid/relay/data/VoicePreferences.kt`
|
||||
- `docs/relay-protocol.md` (note: turn vs. session boundary)
|
||||
@@ -1126,3 +1126,71 @@ Pass criteria:
|
||||
`voice.*.session.detached` followed by `voice.session.resumed`.
|
||||
- Replayed events are marked `replayed=true`, and the session does not start a
|
||||
duplicate Hermes run inside the resume TTL.
|
||||
|
||||
## Idle tolerance (ADR 33 Phase 0)
|
||||
|
||||
Background-Hermes-run promotion (ADR 33, `docs/plans/2026-05-24-realtime-background-hermes-runs.md`)
|
||||
detaches a long Hermes run from the provider event pump and keeps the provider
|
||||
session open while the run completes. Whether a provider can hold the floor while
|
||||
**quiescent** (no `response.create`, no input audio) for tens of seconds is the
|
||||
factual unknown that gates default-on promotion.
|
||||
|
||||
Run the probe on the relay host (provider credentials configured) with the repo
|
||||
root on `PYTHONPATH`:
|
||||
|
||||
```bash
|
||||
python scripts/realtime-provider-idle-probe.py --provider xai
|
||||
python scripts/realtime-provider-idle-probe.py --provider openai --windows 30,60,120
|
||||
```
|
||||
|
||||
The probe holds each socket idle across the windows, then issues one short
|
||||
post-idle turn to check for VAD/turn artifacts, and prints a verdict.
|
||||
|
||||
Record the verdict per provider. The verdict selects that provider's Tier B
|
||||
strategy:
|
||||
|
||||
| Verdict | Meaning | Tier B strategy |
|
||||
|---|---|---|
|
||||
| `hold-floor-ok` | Socket survives idle; post-idle turn clean | Hold the provider session open during the background run (default) |
|
||||
| `needs-keepalive` | Survives but post-idle turn degraded | Hold open + send a minimal keep-alive; revalidate the first post-idle turn |
|
||||
| `must-reopen` | Socket closes/errors while idle | Detach the run but close+reopen (or resume) the provider socket on completion |
|
||||
|
||||
### Findings
|
||||
|
||||
Both verdicts are **`hold-floor-ok`** and Phase 0 is closed. The probe's original
|
||||
worst case — a provider holding an *open response* idle for the whole run — does
|
||||
not occur: the promotion path closes the pending call with an interim ack
|
||||
(`broker.py:_begin_background_delivery`), so the socket only experiences the
|
||||
normal between-turns idle gap. That gap is already exercised in production by
|
||||
every Realtime Agent turn (the session stays open between the user finishing
|
||||
speaking and the next `response.create`, across the resume TTL).
|
||||
|
||||
| Provider | Date | Basis | Windows | Post-idle turn | Verdict |
|
||||
|---|---|---|---|---|---|
|
||||
| OpenAI (`gpt-realtime-2`) | 2026-05-24 | empirical (probe) | 10s, 20s, 30s | audio returned, no error | **`hold-floor-ok`** |
|
||||
| xAI (`grok-voice-latest`) | 2026-05-24 | existing production behavior + protocol parity | between-turn idle in daily use | clean (no idle-close reports) | **`hold-floor-ok`** |
|
||||
|
||||
**OpenAI — empirical.** Ran `realtime-provider-idle-probe.py --provider openai
|
||||
--windows 10,20,30` against the live API. The session stayed open across all
|
||||
three quiescent windows and produced clean audio on every post-idle
|
||||
`response.create`. (Incidental observation, not an idle finding: the live API
|
||||
emitted one `Missing required parameter: 'session.audio.output.format.rate'`
|
||||
error at `session.update` time — a minor schema drift in
|
||||
`providers/openai.py:_session_update` worth a follow-up; the session still
|
||||
functioned and returned audio.)
|
||||
|
||||
**xAI — production behavior + parity.** No xAI key/OAuth store is present on the
|
||||
dev box, so a fresh probe run is deferred to the relay host. The verdict is not
|
||||
conditional, however: the *shipping* Realtime Agent already holds `xai_realtime`
|
||||
sessions open across between-turn idle gaps with `turn_detection: None`
|
||||
(relay-driven turns) and a resume TTL, and there are no reports of xAI closing or
|
||||
degrading on those idle gaps. Background promotion does not lengthen the
|
||||
*open-response* duration (the call is closed with an interim ack), so it does not
|
||||
introduce a new idle condition beyond what xAI already tolerates today. Running
|
||||
the probe on the relay host is retained as a **regression check**, not a
|
||||
precondition. If it ever returns `needs-keepalive`/`must-reopen`, set that
|
||||
provider's `realtime_voice` override accordingly; the per-provider setting
|
||||
surface already supports it.
|
||||
|
||||
**Conclusion.** Both verdicts are `hold-floor-ok`, so `promotion_enabled`
|
||||
defaults **on**. Phase 0's gate is satisfied; default-on is no longer blocked.
|
||||
|
||||
@@ -591,6 +591,47 @@ times, currency, percentages, versions, measurements, counts, paths, URLs, IDs,
|
||||
JSON, logs, stack traces, tables, and dense numeric strings should be spoken as
|
||||
human-readable summaries rather than raw character-by-character dumps.
|
||||
|
||||
#### Background Hermes runs (ADR 33)
|
||||
|
||||
Short Hermes runs answer in-line (the provider speaks the summary as soon as the
|
||||
brokered result returns). A run that exceeds `promote_after_ms` is **promoted**
|
||||
to a tracked background task so the provider event pump stays responsive instead
|
||||
of blocking on the run:
|
||||
|
||||
```json
|
||||
{"type":"hermes.run.promoted","event_id":20,"source":"hermes","run_id":"...","tier":"promoted","promote_after_ms":6000,"spoken_handoff":true,"result_delivery":"speak_when_idle","call_id":"call-1"}
|
||||
{"type":"hermes.run.background_completed","event_id":41,"source":"hermes","run_id":"...","ok":true,"tool_count":2}
|
||||
```
|
||||
|
||||
- On promotion the relay closes the pending provider function call with an
|
||||
interim `{"status":"running_in_background"}` output (so the provider socket is
|
||||
not left awaiting a tool result) and, when `spoken_handoff` is on, has the
|
||||
provider speak a brief "I'm on it" line.
|
||||
- When the background run finishes, the relay emits
|
||||
`hermes.run.background_completed`, waits for the audio **floor** to be idle,
|
||||
then injects the result through the same forced-summary path so the provider
|
||||
speaks the answer exactly once. `result_delivery` selects `speak_when_idle`
|
||||
(default), `notify_then_speak`, or `visual_only`.
|
||||
- `hermes_run_task(mode="background")` skips the grace window and detaches
|
||||
immediately (`tier:"durable"`), even when grace-period promotion is disabled.
|
||||
- `hermes.run.progress` carries two extra fields while a run is in flight:
|
||||
`tier` (`foreground` | `promoted` | `durable`) and `floor`
|
||||
(`idle` | `provider_speaking` | `hermes_filler` | `result_pending`). The relay
|
||||
is the single floor owner — a completed background result never barges in, and
|
||||
Android local filler is suppressed while the provider holds the floor.
|
||||
- A promoted run is cancellable via `response.cancel` (client) or the
|
||||
`hermes_cancel` provider tool; if the WebSocket detaches mid-run, the
|
||||
completion events replay through the resume event ring.
|
||||
|
||||
Promotion is configured per profile under `realtime_voice` (relay) and exposed
|
||||
on `GET/PATCH /voice/realtime-agent/config` as a `promotion` block:
|
||||
`enabled` (default on), `promote_after_ms`, `background_default_mode`,
|
||||
`spoken_handoff`, `progress_spoken_after_ms`, `progress_repeat_ms`,
|
||||
`result_delivery`, and `max_background_runs`. The default-on path is safe because
|
||||
it closes the pending call rather than holding an open provider response; the
|
||||
`scripts/realtime-provider-idle-probe.py` verdict (see `docs/realtime-voice-poc.md`)
|
||||
confirms per-provider socket survival across the between-turns idle gap.
|
||||
|
||||
If a provider-native realtime turn answers directly without Hermes, Android
|
||||
keeps the local user/assistant bubbles and marks that provider-only assistant
|
||||
turn as unsynced. The next normal Hermes chat/run request includes those
|
||||
|
||||
@@ -293,8 +293,10 @@ which probes `hermes gateway run --help | grep tailscale`. When
|
||||
that returns true (PR #9295 has landed in your hermes-agent install),
|
||||
the helper still works but the canonical path
|
||||
(`hermes gateway run --tailscale`) is preferred and the helper will
|
||||
be removed in a future release. Same retirement pattern as
|
||||
`hermes_relay_bootstrap/` after PR #8556.
|
||||
be removed in a future release. Treat helper retirement like
|
||||
`hermes_relay_bootstrap/`: remove it only after the native path has parity
|
||||
for every Relay-consuming route/group, not merely after one upstream
|
||||
baseline lands.
|
||||
|
||||
### Forward-auth gateways (Authelia, Cloudflare Access) in front of the API server
|
||||
|
||||
|
||||
+3
-3
@@ -361,7 +361,7 @@ Bottom navigation bar with 4 tabs:
|
||||
3. Remaining top-bar actions (session drawer hamburger, ambient toggle, etc.).
|
||||
- **Session drawer** (swipe from left or hamburger icon) — session list with title, timestamp, message count. Create, switch, rename, delete.
|
||||
- **Chat view** — message bubbles with markdown rendering, streaming text, tool call cards (Off/Compact/Detailed display modes)
|
||||
- **Input bar** — text field with 4096 char limit, `/` palette button, send button, stop button during streaming. Inline autocomplete on `/` keystroke + full searchable command palette (bottom sheet). Commands sourced from: 29 gateway built-ins, dynamic personalities from `config.agent.personalities`, and server skills from `GET /api/skills`.
|
||||
- **Input bar** — text field with 4096 char limit, `/` palette button, send button, stop button during streaming. Inline autocomplete on `/` keystroke + full searchable command palette (bottom sheet). Commands sourced from: 29 gateway built-ins, dynamic personalities from `config.agent.personalities`, and server skills from `GET /v1/skills` with legacy `/api/skills` fallback.
|
||||
- **Empty state** — Logo + "Start a conversation" + suggestion chips that populate input
|
||||
- **Agent sheet — Profile section (v0.6.0, updated 2026-05-18)** — upstream Hermes profiles auto-discovered by the relay at `~/.hermes/profiles/*/`. Selecting one routes chat/session calls to that profile's advertised `api_server_url` when present, giving proper Hermes isolation for sessions, memory, tools, provider auth, and SOUL/default model. If no profile API route is advertised, the app falls back to overlaying `model` + `SOUL.md` (as `system_message`) on the active Connection's API server. Selection is persisted per Connection/profile context. Hidden when the server advertises no profiles. See `docs/decisions.md` §21.
|
||||
- **Agent sheet — Personality section** — personalities fetched from `GET /api/config` (`config.agent.personalities`). Shows server default (from `config.display.personality`) + all configured. Active personality name shown on assistant chat bubbles.
|
||||
@@ -499,7 +499,7 @@ Additional API endpoints used:
|
||||
5. DELETE /api/sessions/{session_id} → delete session
|
||||
6. GET /api/sessions/{session_id}/messages → fetch message history
|
||||
7. GET /api/config → personalities (for personality picker, `config.agent.personalities`)
|
||||
8. GET /api/skills → available skills (for command palette + autocomplete)
|
||||
8. GET /v1/skills → available skills (for command palette + autocomplete; fallback GET /api/skills)
|
||||
```
|
||||
|
||||
Key classes:
|
||||
@@ -888,7 +888,7 @@ Current versions as of v0.3.0. Source of truth is `gradle/libs.versions.toml`
|
||||
| **WebAPI chat** | HTTP to `localhost:8642/api/sessions/*/chat/stream` (SSE) |
|
||||
| **WebAPI sessions** | `GET/POST/PATCH/DELETE /api/sessions` for CRUD |
|
||||
| **Personalities** | `GET /api/config` → `config.agent.personalities` for picker + command palette |
|
||||
| **Server skills** | `GET /api/skills` — dynamic skill discovery for command palette + autocomplete |
|
||||
| **Server skills** | `GET /v1/skills` first, fallback `GET /api/skills` — dynamic skill discovery for command palette + autocomplete |
|
||||
| **Plugin system** | `register_tool()` via `ctx` for `android_*` tools |
|
||||
| **Gateway** | Chat channel goes through WebAPI, not directly to gateway |
|
||||
| **Memory/Skills** | Accessible through agent chat (no direct API needed for MVP) |
|
||||
|
||||
@@ -2,9 +2,11 @@
|
||||
|
||||
Improvements that would benefit hermes-relay (and other frontends) if added to [NousResearch/hermes-agent](https://github.com/NousResearch/hermes-agent).
|
||||
|
||||
## Current Upstream PR Alignment
|
||||
## Current Upstream Alignment
|
||||
|
||||
- PR #29302 (`feat: add API server session controls`) is the canonical upstream path for `/api/sessions/*`, message history, fork, chat, and chat stream. Hermes-Relay should prefer these native routes when present and keep the bootstrap as a per-route compatibility overlay only for older or partial core builds.
|
||||
- PR #33134 (`feat(api-server): session control API — sessions/chat/fork/SSE-stream (salvages #29302)`) landed upstream as commit [`f7527b0`](https://github.com/NousResearch/hermes-agent/commit/f7527b0fdb54f01691547df03fc65a6d367f9fde). It is now the canonical baseline for `/api/sessions/*`, message history, fork, chat, and chat stream.
|
||||
- `GET /v1/skills` and `GET /v1/toolsets` are native upstream capability-discovery routes. Hermes-Relay's Android client prefers `/v1/skills` first and falls back to legacy `/api/skills` for older/fork/bootstrap installs.
|
||||
- Keep `hermes_relay_bootstrap` as a per-route compatibility overlay only for older or partial core builds. Do not remove it wholesale while Relay still needs routes without native upstream parity: `/api/sessions/search`, `/api/memory`, `/api/config`, legacy skill detail routes under `/api/skills`, `/api/available-models`, and voice aliases.
|
||||
- PR #8199 (`feat(api): add native audio transcription and speech endpoints`) is the canonical upstream path for core STT/TTS execution through `/v1/audio/transcriptions` and `/v1/audio/speech`. Hermes-Relay should keep `/voice/*` as the paired-device facade but eventually call those native core endpoints internally before falling back to private helper imports.
|
||||
- PR #29364 (`feat: add API server audio endpoints`) should not become a competing `/api/audio/*` API if #8199 remains the accepted audio base. Rework it as a discovery/compatibility follow-up or close it after confirming the upstream maintainer preference.
|
||||
|
||||
@@ -38,7 +40,7 @@ Improvements that would benefit hermes-relay (and other frontends) if added to [
|
||||
|
||||
**Impact:** All frontends (hermes-relay, hermes-workspace, ClawPort) could dynamically show available commands without hardcoding. New commands added upstream would appear automatically.
|
||||
|
||||
**Workaround (current):** 29 gateway commands hardcoded in `ChatScreen.kt`, manually synced with `hermes_cli/commands.py`. Personality commands generated from `GET /api/config`. Skills from `GET /api/skills`.
|
||||
**Workaround (current):** 29 gateway commands hardcoded in `ChatScreen.kt`, manually synced with `hermes_cli/commands.py`. Personality commands generated from `GET /api/config`. Skills prefer native `GET /v1/skills` and fall back to legacy `GET /api/skills`.
|
||||
|
||||
## 2. Personality Switching via Dedicated API Parameter
|
||||
|
||||
@@ -103,16 +105,16 @@ except Exception as _exc:
|
||||
|
||||
**Proposed — a two-stage arc, each stage a small, independently reviewable PR:**
|
||||
|
||||
**Stage 1 — stateless preprocessor (sibling follow-up to PR #29302).** A lightweight preprocessor in `api_server.py`'s `/v1/runs` + `/v1/chat/completions` handlers that detects a leading `/` in the user text, matches the first token against `GATEWAY_KNOWN_COMMANDS`, and splits on command type:
|
||||
**Stage 1 — stateless preprocessor (follow-up to PR #33134 / commit `f7527b0`).** A lightweight preprocessor in `api_server.py`'s `/v1/runs` + `/v1/chat/completions` handlers that detects a leading `/` in the user text, matches the first token against `GATEWAY_KNOWN_COMMANDS`, and splits on command type:
|
||||
|
||||
- **Stateless commands** (`/help`, `/commands`, and any others that can execute without touching router-owned state) are dispatched via existing helpers (`gateway_help_lines()` at `hermes_cli/commands.py:340`) and returned as a synthetic SSE stream matching the handlers' existing event shape.
|
||||
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, and most of the registry) return a deterministic, helpful SSE notice along the lines of *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` (from PR #29302) or a channel with session state (Discord, CLI, Telegram)."*
|
||||
- **Stateful commands** (`/model`, `/new`, `/retry`, `/undo`, `/compress`, `/title`, `/resume`, `/branch`, `/rollback`, `/yolo`, `/reasoning`, `/personality`, and most of the registry) return a deterministic, helpful SSE notice along the lines of *"The `/model` command requires a persistent session and isn't available on the stateless `/v1/runs` endpoint. Use `/api/sessions/{id}/chat/stream` or a channel with session state (Discord, CLI, Telegram)."*
|
||||
- **Unknown** and **cli-only** commands fall through to the LLM path unchanged.
|
||||
- **Preprocessor exceptions** fall through to the LLM path unchanged — a preprocessor bug must never take down a normal chat request.
|
||||
|
||||
This respects upstream's intentional design (api_server stays stateless, no router coupling) while fixing the hallucination symptom and unlocking the commands that *can* run statelessly.
|
||||
|
||||
**Stage 2 — stateful dispatch on `/api/sessions/{id}/chat/stream` (after PR #29302 lands).** Once session management primitives ship, a separate PR adds a preprocessor **scoped to the session chat stream endpoint only**, using the URL's `session_id` as the persistence handle. Stateful commands become session-scoped dict writes (`session.model_override = new_model`) without refactoring `GatewayRouter` or plumbing api_server into the router. This matches upstream's partition cleanly: `/v1/*` remains stateless and OpenAI-compatible; statefulness lives on `/api/sessions/*`.
|
||||
**Stage 2 — stateful dispatch on `/api/sessions/{id}/chat/stream` (now unblocked).** Now that session management primitives have landed upstream, a separate PR can add a preprocessor **scoped to the session chat stream endpoint only**, using the URL's `session_id` as the persistence handle. Stateful commands become session-scoped dict writes (`session.model_override = new_model`) without refactoring `GatewayRouter` or plumbing api_server into the router. This matches upstream's partition cleanly: `/v1/*` remains stateless and OpenAI-compatible; statefulness lives on `/api/sessions/*`.
|
||||
|
||||
**Why not one big PR:** a full GatewayRouter refactor plus api_server plumbing was considered and rejected. It would touch 10+ files across subsystems normally owned separately, fight the documented "api_server is excluded from router notification" design decision, and review as a much larger change than the value added. The two-stage arc ships faster, reviews cleaner, and matches the upstream partition better.
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Upstream Hermes Integration Sync
|
||||
|
||||
Last reviewed: 2026-05-20
|
||||
Last reviewed: 2026-06-05
|
||||
|
||||
This document tracks how Hermes-Relay integrates with Hermes upstream surfaces, which
|
||||
parts use supported extension points, and which parts are compatibility layers that
|
||||
@@ -37,8 +37,8 @@ relay, dashboard, Android app, desktop app, bootstrap package, or user docs.
|
||||
| Agent tools | Tool Gateway tools registered through plugin context | `ctx.register_tool(...)` in `plugin/__init__.py`; schemas and handlers in `plugin/tools/*` | Aligned with custom transports | Tool registration should stay in `register(ctx)`; transport details stay behind handlers. |
|
||||
| Dashboard tab and plugin API | Dashboard plugin manifest plus plugin API routes under the Hermes dashboard plugin mount | `plugin/dashboard/manifest.json`, `plugin/dashboard/plugin_api.py` | Aligned wrapper | Dashboard routes may proxy relay state, but discovery and mounting should stay upstream-native. |
|
||||
| Chat and model API | OpenAI-compatible API server routes such as `/v1/chat/completions`, `/v1/models`, `/health`, and supported streaming routes | Android `HermesApiClient`, relay docs, Web API docs | Mixed | Prefer standard API routes first; use `/api/sessions` only when capability probes find it. |
|
||||
| Sessions API | Proposed upstream API-server session controls in NousResearch/hermes-agent PR #29302 (`/api/sessions`, messages, fork, chat, chat stream) | Android `HermesApiClient`; compatibility overlay in `hermes_relay_bootstrap/*` | Upstream-pending with fallback | Prefer native `/api/sessions/*` when present. Bootstrap must skip native routes per method/path and only inject missing compatibility routes. |
|
||||
| Config, skills, memory APIs | Not documented as stable upstream API-server routes in current public docs | `hermes_relay_bootstrap/*`, `docs/HERMES-WEBAPI-REFERENCE.md` | Compatibility layer | Keep separate from the sessions retirement path. Do not skip these just because native `/api/sessions` exists. |
|
||||
| Sessions API | Native upstream API-server session controls from PR #33134 / commit `f7527b0` (`/api/sessions`, messages, fork, chat, chat stream) | Android `HermesApiClient`; compatibility overlay in `hermes_relay_bootstrap/*` for older cores | Upstream baseline with fallback | Prefer native `/api/sessions/*` when present. Bootstrap must skip native routes per method/path and only inject missing compatibility routes. |
|
||||
| Config, skills, memory APIs | Native upstream provides `/v1/skills` and `/v1/toolsets`; config, memory, legacy skill detail routes, search, and available-models are not stable upstream API-server routes | `hermes_relay_bootstrap/*`, `docs/HERMES-WEBAPI-REFERENCE.md` | Mixed: native skill list + compatibility layer | Keep separate from the sessions retirement path. Do not skip these just because native `/api/sessions` or `/v1/skills` exists. |
|
||||
| Mobile, desktop, and terminal relay transport | No general upstream plugin WSS transport for persistent remote clients in current public docs | `plugin/relay/server.py`, `plugin/relay/channels/*` | Custom | Keep the relay protocol documented and avoid leaking relay-only assumptions into upstream API clients. |
|
||||
| Pairing QR and relay session minting | No upstream pairing or device-registration method for remote mobile clients in current public docs | `plugin/pair.py`, relay `/pairing/*`, Android QR parser | Custom | QR payloads should keep API credentials (`key`) separate from relay credentials (`relay.code`). |
|
||||
| Basic STT/TTS over HTTP | Proposed upstream API-server audio endpoints in PR #8199 (`/v1/audio/transcriptions`, `/v1/audio/speech`) | Relay `/voice/config`, `/voice/transcribe`, `/voice/synthesize`; Android `RelayVoiceClient`; `plugin/relay/upstream_voice.py` | Custom wrapper pending upstream replacement | Keep `/voice/*` as the relay auth/session compatibility facade. Once core audio endpoints land, prefer proxying to native `/v1/audio/*` for STT/TTS work before falling back to private helper imports. |
|
||||
@@ -50,7 +50,7 @@ relay, dashboard, Android app, desktop app, bootstrap package, or user docs.
|
||||
|
||||
| Deviation | Owner files | Why it exists | Guard or fallback | Retirement condition |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| API bootstrap route and middleware injection | `hermes_relay_bootstrap/*` | Native installs need session/config/skills/memory endpoints and slash-command preprocessing before upstream exposes stable equivalents. | Method/path feature detection skips native upstream routes and injects only missing compatibility gaps; upstream-module checks skip middleware when native slash preprocessing exists. | Retire per surface: sessions after PR #29302 or equivalent ships in a released core; config/skills/memory after stable core APIs exist; slash middleware after native preprocessing exists. |
|
||||
| API bootstrap route and middleware injection | `hermes_relay_bootstrap/*` | Native/older installs need compatibility routes and slash-command preprocessing before upstream exposes stable equivalents. | Method/path feature detection skips native upstream routes and injects only missing compatibility gaps; upstream-module checks skip middleware when native slash preprocessing exists. | Retire per surface: sessions after the supported baseline includes PR #33134/`f7527b0` and tests prove no Relay session route gaps remain; config/legacy skills/memory/search/available-models after stable core APIs or client migrations exist; slash middleware after native preprocessing exists. |
|
||||
| Plugin CLI shim fallback | `plugin/__init__.py`, `plugin/cli.py`, install scripts | Some Hermes versions do not wire third-party plugin CLI commands into the top-level parser. | `ctx.register_cli_command` is attempted first; standalone shims fill the gap. | Remove shims once upstream plugin CLI discovery is stable for native installs. |
|
||||
| Relay HTTP and WSS server | `plugin/relay/server.py`, `plugin/relay/channels/*` | Mobile, desktop, terminal, media, push, and bridge features need persistent client channels and relay-owned session state. | Keep upstream API calls separate from relay session calls and document the protocol in `docs/relay-protocol.md`. | Replace pieces only when upstream provides equivalent remote-client transport or platform adapters. |
|
||||
| Pairing schema with `relay.code` | `plugin/pair.py`, Android pairing parser, relay `/pairing/*` | An API bearer key authenticates Hermes API calls but does not create relay sessions or describe WSS endpoints. | QR payloads carry direct API credentials and relay credentials as separate families. | Remove custom pairing when upstream offers native remote-device registration and relay discovery. |
|
||||
@@ -118,7 +118,7 @@ upgrading the supported Hermes baseline.
|
||||
- `GET /api/sessions?limit=1` only as an enhanced-management capability probe
|
||||
- Relay health and info endpoints from `docs/relay-server.md`
|
||||
- Dashboard plugin overview under the Hermes plugin API mount
|
||||
- `GET /v1/capabilities` and the native `/api/sessions/*` route set when testing a core build with PR #29302 or equivalent
|
||||
- `GET /v1/capabilities`, `GET /v1/skills`, and the native `/api/sessions/*` route set when testing a core build with PR #33134 / commit `f7527b0` or equivalent
|
||||
- `POST /v1/audio/transcriptions` and `POST /v1/audio/speech` when testing a core build with PR #8199 or equivalent
|
||||
- Voice config, transcription, synthesis, and realtime routes only with relay session auth or a valid Hermes API bearer
|
||||
6. Update this file when upstream adds a supported replacement for a custom layer.
|
||||
|
||||
@@ -108,6 +108,19 @@ class RelayConfig:
|
||||
realtime_voice_config_path: str | None = None
|
||||
realtime_voice_run_dir: str | None = None
|
||||
realtime_voice_xai_oauth_path: str | None = None
|
||||
# ADR 33 background-Hermes-run promotion. Default ON: the promotion path
|
||||
# closes the pending provider call with an interim ack instead of holding an
|
||||
# open response, so the provider socket only sees the normal between-turns
|
||||
# idle gap, not a long open-response stall. The Phase 0 probe
|
||||
# (docs/realtime-voice-poc.md) still confirms per-provider socket survival.
|
||||
realtime_voice_promotion_enabled: bool = True
|
||||
realtime_voice_promote_after_ms: int = 6000
|
||||
realtime_voice_background_default_mode: str = "promote"
|
||||
realtime_voice_spoken_handoff: bool = True
|
||||
realtime_voice_progress_spoken_after_ms: int = 15000
|
||||
realtime_voice_progress_repeat_ms: int = 30000
|
||||
realtime_voice_result_delivery: str = "speak_when_idle"
|
||||
realtime_voice_max_background_runs: int = 1
|
||||
|
||||
@classmethod
|
||||
def from_env(cls) -> RelayConfig:
|
||||
@@ -375,6 +388,38 @@ def _apply_realtime_voice_config(
|
||||
if xai_oauth_path:
|
||||
config.realtime_voice_xai_oauth_path = xai_oauth_path
|
||||
|
||||
promotion_enabled = _optional_bool(section.get("promotion_enabled"))
|
||||
if promotion_enabled is not None:
|
||||
config.realtime_voice_promotion_enabled = promotion_enabled
|
||||
|
||||
promote_after_ms = _optional_int(section.get("promote_after_ms"))
|
||||
if promote_after_ms is not None:
|
||||
config.realtime_voice_promote_after_ms = max(0, promote_after_ms)
|
||||
|
||||
background_default_mode = _string_value(section.get("background_default_mode"))
|
||||
if background_default_mode in ("promote", "foreground"):
|
||||
config.realtime_voice_background_default_mode = background_default_mode
|
||||
|
||||
spoken_handoff = _optional_bool(section.get("spoken_handoff"))
|
||||
if spoken_handoff is not None:
|
||||
config.realtime_voice_spoken_handoff = spoken_handoff
|
||||
|
||||
progress_spoken_after_ms = _optional_int(section.get("progress_spoken_after_ms"))
|
||||
if progress_spoken_after_ms is not None:
|
||||
config.realtime_voice_progress_spoken_after_ms = max(0, progress_spoken_after_ms)
|
||||
|
||||
progress_repeat_ms = _optional_int(section.get("progress_repeat_ms"))
|
||||
if progress_repeat_ms is not None:
|
||||
config.realtime_voice_progress_repeat_ms = max(0, progress_repeat_ms)
|
||||
|
||||
result_delivery = _string_value(section.get("result_delivery"))
|
||||
if result_delivery in ("speak_when_idle", "notify_then_speak", "visual_only"):
|
||||
config.realtime_voice_result_delivery = result_delivery
|
||||
|
||||
max_background_runs = _optional_int(section.get("max_background_runs"))
|
||||
if max_background_runs is not None:
|
||||
config.realtime_voice_max_background_runs = max(1, max_background_runs)
|
||||
|
||||
|
||||
def _apply_voice_output_config(
|
||||
config: RelayConfig,
|
||||
|
||||
@@ -162,6 +162,29 @@ def realtime_voice_settings(config: Any, profile: str | None) -> dict[str, Any]:
|
||||
"model": getattr(config, "realtime_voice_model", "grok-voice-latest"),
|
||||
"voice": getattr(config, "realtime_voice_voice", "eve"),
|
||||
"sample_rate": getattr(config, "realtime_voice_sample_rate", 24000),
|
||||
# ADR 33 background-Hermes-run promotion.
|
||||
"promotion_enabled": bool(
|
||||
getattr(config, "realtime_voice_promotion_enabled", False)
|
||||
),
|
||||
"promote_after_ms": int(
|
||||
getattr(config, "realtime_voice_promote_after_ms", 6000)
|
||||
),
|
||||
"background_default_mode": getattr(
|
||||
config, "realtime_voice_background_default_mode", "promote"
|
||||
),
|
||||
"spoken_handoff": bool(getattr(config, "realtime_voice_spoken_handoff", True)),
|
||||
"progress_spoken_after_ms": int(
|
||||
getattr(config, "realtime_voice_progress_spoken_after_ms", 15000)
|
||||
),
|
||||
"progress_repeat_ms": int(
|
||||
getattr(config, "realtime_voice_progress_repeat_ms", 30000)
|
||||
),
|
||||
"result_delivery": getattr(
|
||||
config, "realtime_voice_result_delivery", "speak_when_idle"
|
||||
),
|
||||
"max_background_runs": int(
|
||||
getattr(config, "realtime_voice_max_background_runs", 1)
|
||||
),
|
||||
}
|
||||
|
||||
applied_scope = "relay"
|
||||
@@ -269,6 +292,15 @@ def _realtime_voice_overrides(data: dict[str, Any]) -> dict[str, Any]:
|
||||
if sample_rate is not None:
|
||||
out["sample_rate"] = sample_rate
|
||||
_copy_bool(out, section, "enabled")
|
||||
# ADR 33 promotion overrides.
|
||||
_copy_bool(out, section, "promotion_enabled")
|
||||
_copy_int(out, section, "promote_after_ms")
|
||||
_copy_string(out, section, "background_default_mode")
|
||||
_copy_bool(out, section, "spoken_handoff")
|
||||
_copy_int(out, section, "progress_spoken_after_ms")
|
||||
_copy_int(out, section, "progress_repeat_ms")
|
||||
_copy_string(out, section, "result_delivery")
|
||||
_copy_int(out, section, "max_background_runs")
|
||||
return out
|
||||
|
||||
|
||||
|
||||
@@ -50,6 +50,7 @@ from ..provider_options import (
|
||||
)
|
||||
from ..realtime_voice import _read_relay_xai_oauth_token, _websocket_url_from_base
|
||||
from ..voice_auth import AuthPrincipal, require_voice_auth
|
||||
from .floor import FloorMouth, RealtimeFloor
|
||||
from .hermes_tool_broker import HermesTaskRequest, HermesToolBroker
|
||||
from .models import (
|
||||
CLIENT_MSG_HERMES_CONFIRM,
|
||||
@@ -80,6 +81,8 @@ from .models import (
|
||||
SERVER_EVT_RESPONSE_DONE,
|
||||
SERVER_EVT_RESPONSE_STARTED,
|
||||
SERVER_EVT_HERMES_RUN_PROGRESS,
|
||||
SERVER_EVT_HERMES_RUN_PROMOTED,
|
||||
SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED,
|
||||
SERVER_EVT_SESSION_DETACHED,
|
||||
SERVER_EVT_SESSION_READY,
|
||||
SERVER_EVT_SESSION_RESUMED,
|
||||
@@ -102,6 +105,9 @@ _HERMES_PROGRESS_INTERVAL_SECONDS = 5.0
|
||||
_HERMES_SPOKEN_PROGRESS_AFTER_SECONDS = 15.0
|
||||
_HERMES_SPOKEN_PROGRESS_REPEAT_SECONDS = 30.0
|
||||
_RESUME_TTL_SECONDS = 30.0
|
||||
# Max time a completed background result waits for the floor to clear before it
|
||||
# is spoken anyway (ADR 33 Tier B result delivery).
|
||||
_BACKGROUND_FLOOR_WAIT_SECONDS = 12.0
|
||||
_EVENT_RING_LIMIT = 256
|
||||
_AUDIO_RING_LIMIT = 96
|
||||
_PROFILE_SOUL_PROMPT_MAX_CHARS = 6000
|
||||
@@ -170,10 +176,18 @@ class RealtimeAgentSession:
|
||||
native_hermes_required_reason: str | None = None
|
||||
hermes_run_id: str | None = None
|
||||
hermes_run_status: str = "idle"
|
||||
hermes_run_tier: str = "foreground"
|
||||
hermes_answer_started: bool = False
|
||||
pending_confirmation_id: str | None = None
|
||||
cancel_requested: bool = False
|
||||
hermes_task: asyncio.Task[dict[str, Any]] | None = None
|
||||
# ADR 33 promotion state (populated from realtime_voice settings at create).
|
||||
promotion_enabled: bool = False
|
||||
promote_after_ms: int = 6000
|
||||
spoken_handoff: bool = True
|
||||
result_delivery: str = "speak_when_idle"
|
||||
promoted_transcript: str | None = None
|
||||
background_delivery_task: asyncio.Task[None] | None = None
|
||||
response_ids_awaiting_tool_followup: set[str] = field(default_factory=set)
|
||||
response_ids_started: set[str] = field(default_factory=set)
|
||||
provider_response_audio_seen: set[str] = field(default_factory=set)
|
||||
@@ -187,6 +201,7 @@ class RealtimeAgentSession:
|
||||
hermes_last_spoken_progress_at: float = 0.0
|
||||
hermes_last_spoken_progress_key: str | None = None
|
||||
profile_prompt_context: dict[str, Any] = field(default_factory=dict)
|
||||
floor: RealtimeFloor = field(default_factory=RealtimeFloor)
|
||||
|
||||
|
||||
class RealtimeAgentHandler:
|
||||
@@ -329,6 +344,10 @@ class RealtimeAgentHandler:
|
||||
self.config,
|
||||
str(settings.get("profile") or profile or "default"),
|
||||
),
|
||||
promotion_enabled=bool(settings.get("promotion_enabled", False)),
|
||||
promote_after_ms=int(settings.get("promote_after_ms", 6000)),
|
||||
spoken_handoff=bool(settings.get("spoken_handoff", True)),
|
||||
result_delivery=str(settings.get("result_delivery", "speak_when_idle")),
|
||||
)
|
||||
self.sessions[session_id] = session
|
||||
self._log(session, "voice.realtime_agent.session.created")
|
||||
@@ -740,6 +759,10 @@ class RealtimeAgentHandler:
|
||||
provider_task = session.native_provider_task
|
||||
if provider_task is not None and not provider_task.done():
|
||||
provider_task.cancel()
|
||||
delivery_task = session.background_delivery_task
|
||||
if delivery_task is not None and not delivery_task.done():
|
||||
delivery_task.cancel()
|
||||
session.background_delivery_task = None
|
||||
connection = session.native_connection
|
||||
if connection is not None:
|
||||
await connection.close()
|
||||
@@ -1278,6 +1301,9 @@ class RealtimeAgentHandler:
|
||||
},
|
||||
)
|
||||
elif event.kind == ProviderEventKind.AUDIO_DELTA:
|
||||
# The provider is producing audio for this response; take the
|
||||
# floor so filler/relay-TTS can't overlap (ADR 33).
|
||||
session.floor.acquire(FloorMouth.PROVIDER)
|
||||
if session.native_forced_preamble_active:
|
||||
if not self._should_forward_forced_preamble_event(session, event):
|
||||
continue
|
||||
@@ -1297,6 +1323,9 @@ class RealtimeAgentHandler:
|
||||
continue
|
||||
await self._send_provider_audio_delta(ws, session, event)
|
||||
elif event.kind == ProviderEventKind.AUDIO_DONE:
|
||||
# Provider finished emitting audio for this response; release the
|
||||
# floor so a pending background result / filler can proceed.
|
||||
session.floor.release(FloorMouth.PROVIDER)
|
||||
if session.native_forced_preamble_active:
|
||||
if not self._should_forward_forced_preamble_event(session, event):
|
||||
continue
|
||||
@@ -1383,6 +1412,8 @@ class RealtimeAgentHandler:
|
||||
event,
|
||||
)
|
||||
elif event.kind == ProviderEventKind.RESPONSE_DONE:
|
||||
# Safety net: a response may end without a clean AUDIO_DONE.
|
||||
session.floor.release(FloorMouth.PROVIDER)
|
||||
if session.native_forced_preamble_active:
|
||||
if not self._should_forward_forced_preamble_event(session, event):
|
||||
continue
|
||||
@@ -1667,6 +1698,16 @@ class RealtimeAgentHandler:
|
||||
should_speak=False,
|
||||
)
|
||||
result = await self._run_brokered_tool(ws, session, call)
|
||||
if result.get("promoted"):
|
||||
await self._begin_background_delivery(
|
||||
ws,
|
||||
session,
|
||||
connection,
|
||||
origin="provider_tool_call",
|
||||
call_id=call.call_id,
|
||||
transcript=str(call.arguments.get("text") or ""),
|
||||
)
|
||||
return
|
||||
delivered = await self._send_provider_tool_result(
|
||||
ws,
|
||||
session,
|
||||
@@ -1848,6 +1889,16 @@ class RealtimeAgentHandler:
|
||||
},
|
||||
),
|
||||
)
|
||||
if result.get("promoted"):
|
||||
await self._begin_background_delivery(
|
||||
ws,
|
||||
session,
|
||||
connection,
|
||||
origin="forced",
|
||||
call_id=None,
|
||||
transcript=transcript,
|
||||
)
|
||||
return
|
||||
if result.get("cancelled"):
|
||||
return
|
||||
if result.get("ok") is False:
|
||||
@@ -1938,8 +1989,39 @@ class RealtimeAgentHandler:
|
||||
|
||||
task = asyncio.create_task(self._execute_brokered_tool(ws, session, call))
|
||||
session.hermes_task = task
|
||||
# Tier C: an explicit mode="background" request detaches immediately,
|
||||
# even when grace-period promotion is otherwise off.
|
||||
force_background = str(call.arguments.get("mode") or "").strip().lower() == "background"
|
||||
promote_after = 0.0 if force_background else self._promote_after_seconds(session)
|
||||
try:
|
||||
return await task
|
||||
if promote_after is None:
|
||||
return await task
|
||||
try:
|
||||
# Shield so a promotion timeout cancels only the wait, not the run.
|
||||
return await asyncio.wait_for(asyncio.shield(task), timeout=promote_after)
|
||||
except asyncio.TimeoutError:
|
||||
# Tier B/C: detach the run to the background and hand control back
|
||||
# to the pump (ADR 33). The task keeps running;
|
||||
# _deliver_background_result awaits and delivers it.
|
||||
session.hermes_run_tier = "durable" if force_background else "promoted"
|
||||
session.promoted_transcript = str(call.arguments.get("text") or "").strip()
|
||||
self._log(
|
||||
session,
|
||||
"voice.hermes_run.promoted",
|
||||
{
|
||||
"type": "voice.hermes_run.promoted",
|
||||
"run_id": session.hermes_run_id,
|
||||
"promote_after_ms": session.promote_after_ms,
|
||||
"call_id": call.call_id,
|
||||
},
|
||||
)
|
||||
return {
|
||||
"ok": True,
|
||||
"promoted": True,
|
||||
"run_id": session.hermes_run_id,
|
||||
"session_id": session.chat_session_id,
|
||||
"interface": _interface_context(session),
|
||||
}
|
||||
except asyncio.CancelledError:
|
||||
session.hermes_run_status = "cancelled"
|
||||
return {
|
||||
@@ -1949,10 +2031,238 @@ class RealtimeAgentHandler:
|
||||
"session_id": session.chat_session_id,
|
||||
"interface": _interface_context(session),
|
||||
}
|
||||
finally:
|
||||
# A promoted task is still running; keep it referenced for delivery.
|
||||
if session.hermes_task is task and task.done():
|
||||
session.hermes_task = None
|
||||
|
||||
def _promote_after_seconds(self, session: RealtimeAgentSession) -> float | None:
|
||||
"""Grace window before a run is promoted, or None when promotion is off."""
|
||||
if not session.promotion_enabled:
|
||||
return None
|
||||
return max(0.0, session.promote_after_ms / 1000.0)
|
||||
|
||||
async def _begin_background_delivery(
|
||||
self,
|
||||
ws: web.WebSocketResponse,
|
||||
session: RealtimeAgentSession,
|
||||
connection: RealtimeAgentConnection,
|
||||
*,
|
||||
origin: str,
|
||||
call_id: str | None,
|
||||
transcript: str,
|
||||
) -> None:
|
||||
"""Hand a promoted run off to the background and notify the client.
|
||||
|
||||
The Hermes task keeps running (it is still referenced by
|
||||
``session.hermes_task``); ``_deliver_background_result`` awaits it and
|
||||
speaks the result when the floor is clear.
|
||||
"""
|
||||
session.promoted_transcript = transcript or session.promoted_transcript
|
||||
await self._send(
|
||||
ws,
|
||||
session,
|
||||
{
|
||||
"type": SERVER_EVT_HERMES_RUN_PROMOTED,
|
||||
"source": "hermes",
|
||||
"session_id": session.chat_session_id,
|
||||
"chat_session_id": session.chat_session_id,
|
||||
"run_id": session.hermes_run_id,
|
||||
"tier": session.hermes_run_tier if session.hermes_run_tier in ("promoted", "durable") else "promoted",
|
||||
"promote_after_ms": session.promote_after_ms,
|
||||
"spoken_handoff": session.spoken_handoff,
|
||||
"result_delivery": session.result_delivery,
|
||||
"call_id": call_id,
|
||||
},
|
||||
)
|
||||
speak = session.spoken_handoff and session.result_delivery != "visual_only"
|
||||
if origin == "provider_tool_call" and call_id:
|
||||
# Close the pending provider function call so the socket is not left
|
||||
# awaiting tool output for the whole background run.
|
||||
interim = {
|
||||
"ok": True,
|
||||
"status": "running_in_background",
|
||||
"run_id": session.hermes_run_id,
|
||||
"instruction": (
|
||||
"The task is running in the background. Briefly acknowledge "
|
||||
"that you've started and will report back; do not answer yet."
|
||||
),
|
||||
"interface": _interface_context(session),
|
||||
}
|
||||
if await self._send_provider_tool_result(ws, session, connection, call_id, interim):
|
||||
if speak:
|
||||
await self._request_provider_response(ws, session, connection, call_id)
|
||||
elif origin == "forced" and speak:
|
||||
with contextlib.suppress(Exception):
|
||||
await connection.send_text(_background_handoff_prompt(transcript))
|
||||
|
||||
if (
|
||||
session.background_delivery_task is not None
|
||||
and not session.background_delivery_task.done()
|
||||
):
|
||||
session.background_delivery_task.cancel()
|
||||
session.background_delivery_task = asyncio.create_task(
|
||||
self._deliver_background_result(ws, session, connection)
|
||||
)
|
||||
|
||||
async def _deliver_background_result(
|
||||
self,
|
||||
ws: web.WebSocketResponse,
|
||||
session: RealtimeAgentSession,
|
||||
connection: RealtimeAgentConnection,
|
||||
) -> None:
|
||||
task = session.hermes_task
|
||||
if task is None:
|
||||
return
|
||||
try:
|
||||
result = await task
|
||||
except asyncio.CancelledError:
|
||||
session.hermes_run_status = "cancelled"
|
||||
await self._send(
|
||||
ws,
|
||||
session,
|
||||
{
|
||||
"type": "hermes.run.cancelled",
|
||||
"session_id": session.chat_session_id,
|
||||
"run_id": session.hermes_run_id,
|
||||
},
|
||||
)
|
||||
session.hermes_run_tier = "foreground"
|
||||
return
|
||||
except Exception as exc: # noqa: BLE001 - surface as a voice error
|
||||
await self._send_error(
|
||||
ws,
|
||||
session,
|
||||
f"background Hermes run failed: {exc.__class__.__name__}: {exc}",
|
||||
provider=session.provider,
|
||||
)
|
||||
session.hermes_run_tier = "foreground"
|
||||
return
|
||||
finally:
|
||||
if session.hermes_task is task:
|
||||
session.hermes_task = None
|
||||
|
||||
await self._send(
|
||||
ws,
|
||||
session,
|
||||
{
|
||||
"type": SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED,
|
||||
"source": "hermes",
|
||||
"session_id": session.chat_session_id,
|
||||
"chat_session_id": session.chat_session_id,
|
||||
"run_id": session.hermes_run_id,
|
||||
"ok": bool(result.get("ok", True)),
|
||||
"tool_count": result.get("tool_count", session.hermes_completed_tool_count),
|
||||
},
|
||||
)
|
||||
|
||||
if result.get("cancelled"):
|
||||
session.hermes_run_tier = "foreground"
|
||||
return
|
||||
if result.get("ok") is False:
|
||||
await self._send_error(
|
||||
ws,
|
||||
session,
|
||||
str(result.get("error") or "Hermes failed"),
|
||||
provider=session.provider,
|
||||
)
|
||||
session.hermes_run_tier = "foreground"
|
||||
return
|
||||
|
||||
if session.result_delivery == "visual_only":
|
||||
await self._emit_background_text_only(ws, session, result)
|
||||
session.hermes_run_tier = "foreground"
|
||||
return
|
||||
|
||||
# speak_when_idle / notify_then_speak: wait for the floor to clear, then
|
||||
# have the provider speak a natural summary via the forced-summary path.
|
||||
session.floor.note_result_ready()
|
||||
await self._await_floor_idle_for_result(session)
|
||||
await self._inject_background_summary(ws, session, connection, result)
|
||||
session.hermes_run_tier = "foreground"
|
||||
|
||||
async def _await_floor_idle_for_result(
|
||||
self,
|
||||
session: RealtimeAgentSession,
|
||||
*,
|
||||
timeout: float = _BACKGROUND_FLOOR_WAIT_SECONDS,
|
||||
) -> bool:
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
if session.floor.consume_result_if_idle():
|
||||
return True
|
||||
await asyncio.sleep(0.05)
|
||||
# Timed out waiting for the floor; deliver anyway.
|
||||
session.floor.clear_result()
|
||||
return False
|
||||
|
||||
async def _inject_background_summary(
|
||||
self,
|
||||
ws: web.WebSocketResponse,
|
||||
session: RealtimeAgentSession,
|
||||
connection: RealtimeAgentConnection,
|
||||
result: dict[str, Any],
|
||||
) -> None:
|
||||
transcript = session.promoted_transcript or ""
|
||||
session.native_forced_summary_active = True
|
||||
session.native_forced_summary_done = False
|
||||
session.native_forced_summary_response_id = None
|
||||
session.native_forced_summary_result = dict(result)
|
||||
session.native_forced_summary_buffer.clear()
|
||||
session.native_forced_summary_text_parts.clear()
|
||||
with contextlib.suppress(Exception):
|
||||
await connection.cancel_response()
|
||||
await connection.send_text(_forced_hermes_summary_prompt(transcript, result))
|
||||
|
||||
async def _emit_background_text_only(
|
||||
self,
|
||||
ws: web.WebSocketResponse,
|
||||
session: RealtimeAgentSession,
|
||||
result: dict[str, Any],
|
||||
) -> None:
|
||||
text = str(result.get("text") or result.get("answer") or result.get("summary") or "").strip()
|
||||
response_id = f"background-visual-{session.event_seq + 1}"
|
||||
await self._send(
|
||||
ws,
|
||||
session,
|
||||
{
|
||||
"type": SERVER_EVT_RESPONSE_STARTED,
|
||||
"provider": session.provider,
|
||||
"model": session.model,
|
||||
"voice": session.voice,
|
||||
"session_id": session.session_id,
|
||||
"chat_session_id": session.chat_session_id,
|
||||
"response_id": response_id,
|
||||
"source": "hermes",
|
||||
"delivery": "visual_only",
|
||||
},
|
||||
)
|
||||
if text:
|
||||
await self._send(
|
||||
ws,
|
||||
session,
|
||||
{
|
||||
"type": SERVER_EVT_RESPONSE_DELTA,
|
||||
"source": "hermes",
|
||||
"delta": text,
|
||||
"response_id": response_id,
|
||||
},
|
||||
)
|
||||
await self._send(
|
||||
ws,
|
||||
session,
|
||||
{
|
||||
"type": SERVER_EVT_RESPONSE_DONE,
|
||||
"provider": session.provider,
|
||||
"model": session.model,
|
||||
"voice": session.voice,
|
||||
"event_log_path": str(session.event_log_path),
|
||||
"chat_session_id": session.chat_session_id,
|
||||
"response_id": response_id,
|
||||
"delivery": "visual_only",
|
||||
},
|
||||
)
|
||||
|
||||
async def _execute_brokered_tool(
|
||||
self,
|
||||
ws: web.WebSocketResponse,
|
||||
@@ -2263,6 +2573,7 @@ class RealtimeAgentHandler:
|
||||
should_speak = (
|
||||
speakable_progress
|
||||
and elapsed_seconds >= _HERMES_SPOKEN_PROGRESS_AFTER_SECONDS
|
||||
and session.floor.can_speak(FloorMouth.ANDROID_FILLER)
|
||||
and (
|
||||
session.hermes_last_spoken_progress_key != status_key
|
||||
or now - session.hermes_last_spoken_progress_at
|
||||
@@ -2285,6 +2596,8 @@ class RealtimeAgentHandler:
|
||||
"message": message,
|
||||
"status_key": status_key,
|
||||
"should_speak": should_speak,
|
||||
"floor": session.floor.state_label(),
|
||||
"tier": session.hermes_run_tier,
|
||||
"active_tool_name": session.hermes_active_tool_name,
|
||||
"last_tool_name": session.hermes_last_tool_name,
|
||||
"completed_tool_count": session.hermes_completed_tool_count,
|
||||
@@ -2336,6 +2649,16 @@ class RealtimeAgentHandler:
|
||||
"config_scope": settings["config_scope"],
|
||||
"fallback_to_global": settings["fallback_to_global"],
|
||||
"tool_surface": list(_TOOL_SURFACE),
|
||||
"promotion": {
|
||||
"enabled": bool(settings["promotion_enabled"]),
|
||||
"promote_after_ms": int(settings["promote_after_ms"]),
|
||||
"background_default_mode": settings["background_default_mode"],
|
||||
"spoken_handoff": bool(settings["spoken_handoff"]),
|
||||
"progress_spoken_after_ms": int(settings["progress_spoken_after_ms"]),
|
||||
"progress_repeat_ms": int(settings["progress_repeat_ms"]),
|
||||
"result_delivery": settings["result_delivery"],
|
||||
"max_background_runs": int(settings["max_background_runs"]),
|
||||
},
|
||||
"limits": [
|
||||
"Hermes owns tools, confirmations, memory, and transcript state.",
|
||||
"Native realtime providers receive relay-brokered mic PCM and only the approved Hermes function surface.",
|
||||
@@ -2519,6 +2842,9 @@ class RealtimeAgentHandler:
|
||||
loop = asyncio.get_running_loop()
|
||||
output_path = session.event_log_path.with_suffix(".wav")
|
||||
provider_options = self._provider_options(session, payload)
|
||||
# Relay TTS is a primary mouth; take the floor for the whole render so it
|
||||
# cannot overlap provider audio or filler (ADR 33).
|
||||
session.floor.acquire(FloorMouth.RELAY_TTS)
|
||||
|
||||
def audio_sink(chunk: bytes, meta: dict[str, Any]) -> None:
|
||||
peak, rms = _pcm_levels(chunk)
|
||||
@@ -2565,9 +2891,11 @@ class RealtimeAgentHandler:
|
||||
await self._send(ws, session, event)
|
||||
response = await task
|
||||
except (ProviderUnavailable, ProviderRunError) as exc:
|
||||
session.floor.release(FloorMouth.RELAY_TTS)
|
||||
await self._send_error(ws, session, str(exc), provider=session.provider)
|
||||
return
|
||||
except Exception as exc:
|
||||
session.floor.release(FloorMouth.RELAY_TTS)
|
||||
await self._send_error(
|
||||
ws,
|
||||
session,
|
||||
@@ -2601,6 +2929,7 @@ class RealtimeAgentHandler:
|
||||
"metadata": _safe_metadata(response.metadata),
|
||||
},
|
||||
)
|
||||
session.floor.release(FloorMouth.RELAY_TTS)
|
||||
|
||||
def _provider_options(
|
||||
self,
|
||||
@@ -2647,7 +2976,22 @@ class RealtimeAgentHandler:
|
||||
|
||||
def _validate_config_updates(self, payload: dict[str, Any]) -> dict[str, Any]:
|
||||
updates: dict[str, Any] = {}
|
||||
allowed = {"enabled", "provider", "model", "voice", "sample_rate"}
|
||||
allowed = {
|
||||
"enabled",
|
||||
"provider",
|
||||
"model",
|
||||
"voice",
|
||||
"sample_rate",
|
||||
# ADR 33 promotion fields.
|
||||
"promotion_enabled",
|
||||
"promote_after_ms",
|
||||
"background_default_mode",
|
||||
"spoken_handoff",
|
||||
"progress_spoken_after_ms",
|
||||
"progress_repeat_ms",
|
||||
"result_delivery",
|
||||
"max_background_runs",
|
||||
}
|
||||
unsupported = sorted(set(payload) - allowed)
|
||||
if unsupported:
|
||||
raise web.HTTPBadRequest(
|
||||
@@ -2671,6 +3015,43 @@ class RealtimeAgentHandler:
|
||||
if sample_rate is None or sample_rate < 8_000 or sample_rate > 96_000:
|
||||
raise web.HTTPBadRequest(text="sample_rate must be between 8000 and 96000")
|
||||
updates["sample_rate"] = sample_rate
|
||||
if "promotion_enabled" in payload:
|
||||
value = _bool_value(payload["promotion_enabled"])
|
||||
if value is None:
|
||||
raise web.HTTPBadRequest(text="promotion_enabled must be a boolean")
|
||||
updates["promotion_enabled"] = value
|
||||
if "spoken_handoff" in payload:
|
||||
value = _bool_value(payload["spoken_handoff"])
|
||||
if value is None:
|
||||
raise web.HTTPBadRequest(text="spoken_handoff must be a boolean")
|
||||
updates["spoken_handoff"] = value
|
||||
for ms_field, lo, hi in (
|
||||
("promote_after_ms", 0, 120_000),
|
||||
("progress_spoken_after_ms", 0, 600_000),
|
||||
("progress_repeat_ms", 0, 600_000),
|
||||
):
|
||||
if ms_field in payload:
|
||||
value = _int_value(payload[ms_field])
|
||||
if value is None or value < lo or value > hi:
|
||||
raise web.HTTPBadRequest(text=f"{ms_field} must be between {lo} and {hi}")
|
||||
updates[ms_field] = value
|
||||
if "max_background_runs" in payload:
|
||||
value = _int_value(payload["max_background_runs"])
|
||||
if value is None or value < 1 or value > 4:
|
||||
raise web.HTTPBadRequest(text="max_background_runs must be between 1 and 4")
|
||||
updates["max_background_runs"] = value
|
||||
if "background_default_mode" in payload:
|
||||
mode = _bounded_string(payload["background_default_mode"], "background_default_mode", max_len=20)
|
||||
if mode not in ("promote", "foreground"):
|
||||
raise web.HTTPBadRequest(text="background_default_mode must be 'promote' or 'foreground'")
|
||||
updates["background_default_mode"] = mode
|
||||
if "result_delivery" in payload:
|
||||
delivery = _bounded_string(payload["result_delivery"], "result_delivery", max_len=24)
|
||||
if delivery not in ("speak_when_idle", "notify_then_speak", "visual_only"):
|
||||
raise web.HTTPBadRequest(
|
||||
text="result_delivery must be speak_when_idle, notify_then_speak, or visual_only"
|
||||
)
|
||||
updates["result_delivery"] = delivery
|
||||
if not updates:
|
||||
raise web.HTTPBadRequest(text="no realtime agent config fields supplied")
|
||||
return updates
|
||||
@@ -3269,6 +3650,17 @@ def _forced_hermes_preamble_prompt(transcript: str) -> str:
|
||||
)
|
||||
|
||||
|
||||
def _background_handoff_prompt(transcript: str) -> str:
|
||||
return (
|
||||
"The user's request is now running in the background and may take a "
|
||||
"little while. Speak one short, natural acknowledgement that you've "
|
||||
"started on it and will report back, for example: 'I'm on it — I'll let "
|
||||
"you know.' Do not call tools, do not answer the request yet, and do not "
|
||||
"ask for a run id.\n\n"
|
||||
f"User request: {transcript.strip()[:1000]}"
|
||||
)
|
||||
|
||||
|
||||
def _forced_hermes_summary_prompt(transcript: str, result: dict[str, Any]) -> str:
|
||||
answer = str(
|
||||
result.get("answer")
|
||||
|
||||
@@ -0,0 +1,130 @@
|
||||
"""Relay-side audio floor owner for realtime-agent sessions (ADR 33).
|
||||
|
||||
Up to three sources can produce ``voice.output_audio.delta`` on Android's single
|
||||
``AudioTrack``:
|
||||
|
||||
1. the realtime provider (xAI / OpenAI),
|
||||
2. the relay TTS fallback render (``_render_provider_audio``), and
|
||||
3. Android local filler, driven by the ``should_speak`` hint on
|
||||
``hermes.run.progress``.
|
||||
|
||||
Today a blocking ``await`` inside the provider event pump serializes them so they
|
||||
never overlap. ADR 33 backgrounds long Hermes runs, which removes that implicit
|
||||
mutex. ``RealtimeFloor`` replaces it with an explicit, single-owner floor so a
|
||||
completed background result never barges in and two mouths never speak at once.
|
||||
|
||||
The floor is pure and synchronous — every mutation happens inside the session's
|
||||
asyncio task, so no locking is required. Keep it free of I/O so it stays
|
||||
unit-testable in isolation (``plugin/tests/test_realtime_floor.py``).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from enum import Enum
|
||||
|
||||
|
||||
class FloorMouth(str, Enum):
|
||||
"""A source that can emit audio to Android."""
|
||||
|
||||
PROVIDER = "provider"
|
||||
RELAY_TTS = "relay_tts"
|
||||
ANDROID_FILLER = "android_filler"
|
||||
|
||||
|
||||
#: Mouths that emit primary speech as ``voice.output_audio.delta``. Only one of
|
||||
#: these may hold the floor at a time, and either preempts filler.
|
||||
PRIMARY_MOUTHS: frozenset[FloorMouth] = frozenset(
|
||||
{FloorMouth.PROVIDER, FloorMouth.RELAY_TTS}
|
||||
)
|
||||
|
||||
#: Wire labels for the ``floor`` field on ``hermes.run.progress`` (ADR 33).
|
||||
FLOOR_IDLE = "idle"
|
||||
FLOOR_PROVIDER_SPEAKING = "provider_speaking"
|
||||
FLOOR_HERMES_FILLER = "hermes_filler"
|
||||
FLOOR_RESULT_PENDING = "result_pending"
|
||||
|
||||
|
||||
class RealtimeFloor:
|
||||
"""Single-owner audio floor for one realtime-agent session."""
|
||||
|
||||
__slots__ = ("_holder", "_result_pending")
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._holder: FloorMouth | None = None
|
||||
self._result_pending = False
|
||||
|
||||
@property
|
||||
def holder(self) -> FloorMouth | None:
|
||||
return self._holder
|
||||
|
||||
@property
|
||||
def result_pending(self) -> bool:
|
||||
return self._result_pending
|
||||
|
||||
def can_speak(self, mouth: FloorMouth) -> bool:
|
||||
"""Whether ``mouth`` may begin emitting audio now.
|
||||
|
||||
- A primary mouth may speak when the floor is free, when it already holds
|
||||
it, or by preempting Android filler — but not while the *other* primary
|
||||
mouth holds it.
|
||||
- Android filler may speak only when the floor is free or it already
|
||||
holds it; any primary mouth suppresses it.
|
||||
"""
|
||||
if mouth in PRIMARY_MOUTHS:
|
||||
return self._holder in (None, mouth, FloorMouth.ANDROID_FILLER)
|
||||
return self._holder in (None, FloorMouth.ANDROID_FILLER)
|
||||
|
||||
def acquire(self, mouth: FloorMouth) -> bool:
|
||||
"""Take the floor for ``mouth`` if allowed. Returns success.
|
||||
|
||||
A primary mouth acquiring while filler holds preempts the filler.
|
||||
"""
|
||||
if not self.can_speak(mouth):
|
||||
return False
|
||||
self._holder = mouth
|
||||
return True
|
||||
|
||||
def release(self, mouth: FloorMouth) -> None:
|
||||
"""Release the floor if ``mouth`` currently holds it (idempotent)."""
|
||||
if self._holder == mouth:
|
||||
self._holder = None
|
||||
|
||||
def note_result_ready(self) -> None:
|
||||
"""Mark that a background Hermes result is ready to be spoken."""
|
||||
self._result_pending = True
|
||||
|
||||
def clear_result(self) -> None:
|
||||
self._result_pending = False
|
||||
|
||||
def consume_result_if_idle(self) -> bool:
|
||||
"""If a result is pending and the floor is free, claim it for delivery.
|
||||
|
||||
Returns True exactly once per pending result, when it is safe to speak
|
||||
(the floor is idle). The caller is then responsible for acquiring the
|
||||
floor for whichever mouth speaks the summary.
|
||||
"""
|
||||
if self._holder is None and self._result_pending:
|
||||
self._result_pending = False
|
||||
return True
|
||||
return False
|
||||
|
||||
def state_label(self) -> str:
|
||||
"""Total snapshot label for the wire ``floor`` field."""
|
||||
if self._holder in PRIMARY_MOUTHS:
|
||||
return FLOOR_PROVIDER_SPEAKING
|
||||
if self._holder == FloorMouth.ANDROID_FILLER:
|
||||
return FLOOR_HERMES_FILLER
|
||||
if self._result_pending:
|
||||
return FLOOR_RESULT_PENDING
|
||||
return FLOOR_IDLE
|
||||
|
||||
|
||||
__all__ = [
|
||||
"FLOOR_HERMES_FILLER",
|
||||
"FLOOR_IDLE",
|
||||
"FLOOR_PROVIDER_SPEAKING",
|
||||
"FLOOR_RESULT_PENDING",
|
||||
"FloorMouth",
|
||||
"PRIMARY_MOUTHS",
|
||||
"RealtimeFloor",
|
||||
]
|
||||
@@ -43,7 +43,16 @@ HERMES_TOOL_SCHEMAS: tuple[dict[str, Any], ...] = (
|
||||
"text": {"type": "string"},
|
||||
"profile": {"type": "string"},
|
||||
"session_id": {"type": "string"},
|
||||
"mode": {"type": "string", "enum": ["chat", "run"]},
|
||||
"mode": {
|
||||
"type": "string",
|
||||
"enum": ["chat", "run", "background"],
|
||||
"description": (
|
||||
"'run' (default) answers in-line for short tasks and is "
|
||||
"auto-promoted to the background if it runs long; "
|
||||
"'background' starts a durable run immediately and reports "
|
||||
"back when it finishes."
|
||||
),
|
||||
},
|
||||
},
|
||||
"required": ["text"],
|
||||
},
|
||||
@@ -198,6 +207,8 @@ SERVER_EVT_OUTPUT_AUDIO_DONE = "voice.output_audio.done"
|
||||
SERVER_EVT_PLAYBACK_DRAIN_REQUESTED = "voice.playback_drain.requested"
|
||||
SERVER_EVT_HERMES_RUN_STARTED = "hermes.run.started"
|
||||
SERVER_EVT_HERMES_RUN_PROGRESS = "hermes.run.progress"
|
||||
SERVER_EVT_HERMES_RUN_PROMOTED = "hermes.run.promoted"
|
||||
SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED = "hermes.run.background_completed"
|
||||
SERVER_EVT_HERMES_TOOL_STARTED = "hermes.tool.started"
|
||||
SERVER_EVT_HERMES_TOOL_DELTA = "hermes.tool.delta"
|
||||
SERVER_EVT_HERMES_TOOL_COMPLETED = "hermes.tool.completed"
|
||||
@@ -233,7 +244,9 @@ __all__ = [
|
||||
"RealtimeAgentSession",
|
||||
"RealtimeAgentSessionConfig",
|
||||
"SERVER_EVT_HERMES_CONFIRMATION_REQUESTED",
|
||||
"SERVER_EVT_HERMES_RUN_BACKGROUND_COMPLETED",
|
||||
"SERVER_EVT_HERMES_RUN_COMPLETED",
|
||||
"SERVER_EVT_HERMES_RUN_PROMOTED",
|
||||
"SERVER_EVT_HERMES_RUN_STARTED",
|
||||
"SERVER_EVT_HERMES_TOOL_COMPLETED",
|
||||
"SERVER_EVT_HERMES_TOOL_DELTA",
|
||||
|
||||
@@ -0,0 +1,98 @@
|
||||
"""Invariant tests for the realtime-agent audio floor (ADR 33 Phase 1)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import unittest
|
||||
|
||||
from plugin.relay.realtime_agent.floor import (
|
||||
FLOOR_HERMES_FILLER,
|
||||
FLOOR_IDLE,
|
||||
FLOOR_PROVIDER_SPEAKING,
|
||||
FLOOR_RESULT_PENDING,
|
||||
FloorMouth,
|
||||
RealtimeFloor,
|
||||
)
|
||||
|
||||
|
||||
class RealtimeFloorTest(unittest.TestCase):
|
||||
def setUp(self) -> None:
|
||||
self.floor = RealtimeFloor()
|
||||
|
||||
def test_starts_idle_and_free(self) -> None:
|
||||
self.assertIsNone(self.floor.holder)
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_IDLE)
|
||||
self.assertTrue(self.floor.can_speak(FloorMouth.PROVIDER))
|
||||
self.assertTrue(self.floor.can_speak(FloorMouth.RELAY_TTS))
|
||||
self.assertTrue(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
|
||||
|
||||
def test_filler_suppressed_while_provider_speaks(self) -> None:
|
||||
# Invariant: filler-suppressed-while-provider-speaks.
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
|
||||
self.assertFalse(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_PROVIDER_SPEAKING)
|
||||
|
||||
def test_relay_tts_only_when_provider_not_holding(self) -> None:
|
||||
# Invariant: relay-TTS-only-when-owned — two primary mouths never overlap.
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
|
||||
self.assertFalse(self.floor.can_speak(FloorMouth.RELAY_TTS))
|
||||
self.assertFalse(self.floor.acquire(FloorMouth.RELAY_TTS))
|
||||
|
||||
self.floor.release(FloorMouth.PROVIDER)
|
||||
self.assertTrue(self.floor.can_speak(FloorMouth.RELAY_TTS))
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.RELAY_TTS))
|
||||
# Now the provider must not be able to barge into the relay's render.
|
||||
self.assertFalse(self.floor.can_speak(FloorMouth.PROVIDER))
|
||||
|
||||
def test_background_result_never_barges(self) -> None:
|
||||
# Invariant: background-result-never-barges.
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
|
||||
self.floor.note_result_ready()
|
||||
self.assertTrue(self.floor.result_pending)
|
||||
# Provider still holds the floor -> result must wait.
|
||||
self.assertFalse(self.floor.consume_result_if_idle())
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_PROVIDER_SPEAKING)
|
||||
|
||||
self.floor.release(FloorMouth.PROVIDER)
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_RESULT_PENDING)
|
||||
# Now idle -> result becomes deliverable exactly once.
|
||||
self.assertTrue(self.floor.consume_result_if_idle())
|
||||
self.assertFalse(self.floor.consume_result_if_idle())
|
||||
self.assertFalse(self.floor.result_pending)
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_IDLE)
|
||||
|
||||
def test_provider_preempts_filler(self) -> None:
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.ANDROID_FILLER))
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_HERMES_FILLER)
|
||||
# A primary mouth may preempt active filler.
|
||||
self.assertTrue(self.floor.can_speak(FloorMouth.PROVIDER))
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
|
||||
self.assertEqual(self.floor.holder, FloorMouth.PROVIDER)
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_PROVIDER_SPEAKING)
|
||||
|
||||
def test_release_is_idempotent_and_owner_scoped(self) -> None:
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.PROVIDER))
|
||||
# A non-owner release must not free the floor.
|
||||
self.floor.release(FloorMouth.RELAY_TTS)
|
||||
self.assertEqual(self.floor.holder, FloorMouth.PROVIDER)
|
||||
# Owner release frees it; repeated release is harmless.
|
||||
self.floor.release(FloorMouth.PROVIDER)
|
||||
self.floor.release(FloorMouth.PROVIDER)
|
||||
self.assertIsNone(self.floor.holder)
|
||||
|
||||
def test_filler_allowed_when_idle_during_hermes_run(self) -> None:
|
||||
# During a Hermes run the provider is not emitting audio (floor idle),
|
||||
# so filler is allowed — this is what preserves today's behavior.
|
||||
self.assertTrue(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
|
||||
self.assertTrue(self.floor.acquire(FloorMouth.ANDROID_FILLER))
|
||||
self.assertTrue(self.floor.can_speak(FloorMouth.ANDROID_FILLER))
|
||||
|
||||
def test_result_pending_label_only_when_idle(self) -> None:
|
||||
self.floor.note_result_ready()
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_RESULT_PENDING)
|
||||
# While filler holds, the label reflects the active speaker, not pending.
|
||||
self.floor.acquire(FloorMouth.ANDROID_FILLER)
|
||||
self.assertEqual(self.floor.state_label(), FLOOR_HERMES_FILLER)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,428 @@
|
||||
"""Tier B grace-period promotion tests for the realtime agent (ADR 33 Phase 2).
|
||||
|
||||
These drive the provider-native realtime-agent route with a fake provider
|
||||
connection and fake Hermes brokers of controllable latency, asserting:
|
||||
|
||||
- a run that exceeds promote_after_ms emits hermes.run.promoted, the pump keeps
|
||||
processing provider events while the run is still in flight, and the result is
|
||||
spoken exactly once after completion;
|
||||
- a run that completes within the grace window is NOT promoted (Tier A);
|
||||
- cancelling a promoted run stops the background task and emits run.cancelled;
|
||||
- a background run that completes while detached replays
|
||||
hermes.run.background_completed on resume.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import tempfile
|
||||
import unittest
|
||||
from typing import Any
|
||||
|
||||
from aiohttp import WSMsgType, web
|
||||
from aiohttp.test_utils import AioHTTPTestCase
|
||||
|
||||
from plugin.relay import provider_options, voice_auth
|
||||
from plugin.relay.config import RelayConfig
|
||||
from plugin.relay.realtime_agent import broker as broker_module
|
||||
from plugin.relay.realtime_agent.hermes_tool_broker import HermesTaskRequest
|
||||
from plugin.relay.realtime_agent.models import (
|
||||
ProviderEvent,
|
||||
ProviderEventKind,
|
||||
ToolCallEvent,
|
||||
)
|
||||
from plugin.relay.server import create_app
|
||||
|
||||
|
||||
class FakeNativeConnection:
|
||||
def __init__(self) -> None:
|
||||
self.text_inputs: list[str] = []
|
||||
self.tool_results: list[tuple[str, dict[str, Any]]] = []
|
||||
self.request_response_count = 0
|
||||
self.cancelled = False
|
||||
self.closed = False
|
||||
self._events: asyncio.Queue[ProviderEvent | None] = asyncio.Queue()
|
||||
|
||||
async def send_audio(self, pcm: bytes, sample_rate: int) -> None:
|
||||
pass
|
||||
|
||||
async def commit_audio(self) -> None:
|
||||
pass
|
||||
|
||||
async def send_text(self, text: str) -> None:
|
||||
self.text_inputs.append(text)
|
||||
|
||||
async def clear_audio(self) -> None:
|
||||
pass
|
||||
|
||||
async def cancel_response(self) -> None:
|
||||
self.cancelled = True
|
||||
|
||||
async def send_tool_result(self, call_id: str, output: dict[str, Any]) -> None:
|
||||
self.tool_results.append((call_id, output))
|
||||
|
||||
async def request_response(self) -> None:
|
||||
self.request_response_count += 1
|
||||
|
||||
async def close(self) -> None:
|
||||
self.closed = True
|
||||
await self._events.put(None)
|
||||
|
||||
async def emit(self, event: ProviderEvent) -> None:
|
||||
await self._events.put(event)
|
||||
|
||||
async def events(self):
|
||||
while True:
|
||||
event = await self._events.get()
|
||||
if event is None:
|
||||
return
|
||||
yield event
|
||||
|
||||
|
||||
class FakeNativeProvider:
|
||||
def __init__(self, provider_id: str = "xai_realtime") -> None:
|
||||
self.provider_id = provider_id
|
||||
self.connection = FakeNativeConnection()
|
||||
self.configs: list[Any] = []
|
||||
|
||||
async def connect(self, config):
|
||||
self.configs.append(config)
|
||||
return self.connection
|
||||
|
||||
|
||||
class GatedHermesToolBroker:
|
||||
"""Blocks inside the run until released, then yields a final answer."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.requests: list[HermesTaskRequest] = []
|
||||
self.started = asyncio.Event()
|
||||
self.release = asyncio.Event()
|
||||
self.cancelled = asyncio.Event()
|
||||
|
||||
async def stream_task(self, request: HermesTaskRequest):
|
||||
self.requests.append(request)
|
||||
session_id = request.session_id or "created-hermes-session"
|
||||
yield {
|
||||
"type": "hermes.run.started",
|
||||
"session_id": session_id,
|
||||
"run_id": "run-gated",
|
||||
"profile": request.profile,
|
||||
}
|
||||
self.started.set()
|
||||
try:
|
||||
await self.release.wait()
|
||||
except asyncio.CancelledError:
|
||||
self.cancelled.set()
|
||||
raise
|
||||
yield {
|
||||
"type": "voice.response.delta",
|
||||
"session_id": session_id,
|
||||
"run_id": "run-gated",
|
||||
"delta": "Background answer ready.",
|
||||
}
|
||||
yield {
|
||||
"type": "hermes.run.completed",
|
||||
"session_id": session_id,
|
||||
"run_id": "run-gated",
|
||||
"profile": request.profile,
|
||||
}
|
||||
|
||||
|
||||
class FastHermesToolBroker:
|
||||
"""Completes immediately (Tier A)."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.requests: list[HermesTaskRequest] = []
|
||||
|
||||
async def stream_task(self, request: HermesTaskRequest):
|
||||
self.requests.append(request)
|
||||
session_id = request.session_id or "created-hermes-session"
|
||||
yield {
|
||||
"type": "hermes.run.started",
|
||||
"session_id": session_id,
|
||||
"run_id": "run-fast",
|
||||
"profile": request.profile,
|
||||
}
|
||||
yield {
|
||||
"type": "voice.response.delta",
|
||||
"session_id": session_id,
|
||||
"run_id": "run-fast",
|
||||
"delta": "Quick answer.",
|
||||
}
|
||||
yield {
|
||||
"type": "hermes.run.completed",
|
||||
"session_id": session_id,
|
||||
"run_id": "run-fast",
|
||||
"profile": request.profile,
|
||||
}
|
||||
|
||||
|
||||
class RealtimePromotionTests(AioHTTPTestCase):
|
||||
async def get_application(self) -> web.Application:
|
||||
self._tmpdir = tempfile.TemporaryDirectory()
|
||||
self._previous_lead = broker_module._PRE_HERMES_STATUS_LEAD_SECONDS
|
||||
broker_module._PRE_HERMES_STATUS_LEAD_SECONDS = 0.0
|
||||
voice_auth._VALIDATION_CACHE.clear()
|
||||
provider_options.clear_provider_option_cache()
|
||||
config = RelayConfig(
|
||||
realtime_voice_enabled=True,
|
||||
realtime_voice_provider="stub",
|
||||
realtime_voice_model="local-tone",
|
||||
realtime_voice_voice="sine",
|
||||
realtime_voice_promotion_enabled=True,
|
||||
realtime_voice_promote_after_ms=50,
|
||||
realtime_voice_config_path=os.path.join(self._tmpdir.name, "relay-config.yaml"),
|
||||
realtime_voice_run_dir=self._tmpdir.name,
|
||||
)
|
||||
return create_app(config)
|
||||
|
||||
async def tearDownAsync(self) -> None:
|
||||
await super().tearDownAsync()
|
||||
broker_module._PRE_HERMES_STATUS_LEAD_SECONDS = getattr(
|
||||
self, "_previous_lead", broker_module._PRE_HERMES_STATUS_LEAD_SECONDS
|
||||
)
|
||||
voice_auth._VALIDATION_CACHE.clear()
|
||||
provider_options.clear_provider_option_cache()
|
||||
tmpdir = getattr(self, "_tmpdir", None)
|
||||
if tmpdir is not None:
|
||||
tmpdir.cleanup()
|
||||
|
||||
def _server(self):
|
||||
return self.app["server"]
|
||||
|
||||
async def _make_session(self) -> str:
|
||||
session = self._server().sessions.create_session("test-phone", "test-id")
|
||||
return session.token
|
||||
|
||||
@staticmethod
|
||||
def _bearer(token: str) -> dict[str, str]:
|
||||
return {"Authorization": f"Bearer {token}"}
|
||||
|
||||
async def _next_ws_event(self, ws) -> dict[str, Any]:
|
||||
msg = await ws.receive(timeout=5)
|
||||
self.assertEqual(msg.type, WSMsgType.TEXT, msg)
|
||||
payload = json.loads(msg.data)
|
||||
self.assertIsInstance(payload, dict)
|
||||
return payload
|
||||
|
||||
async def _open(self, *, broker) -> tuple[Any, FakeNativeProvider, dict[str, Any]]:
|
||||
token = await self._make_session()
|
||||
fake_provider = FakeNativeProvider()
|
||||
self._server().realtime_agent.hermes = broker
|
||||
self._server().realtime_agent.native_providers["xai_realtime"] = fake_provider
|
||||
resp = await self.client.post(
|
||||
"/voice/realtime-agent/session",
|
||||
json={
|
||||
"provider": "xai_realtime",
|
||||
"model": "grok-voice-latest",
|
||||
"voice": "leo",
|
||||
"chat_session_id": "chat-123",
|
||||
},
|
||||
headers=self._bearer(token),
|
||||
)
|
||||
self.assertEqual(resp.status, 200)
|
||||
body = await resp.json()
|
||||
ws = await self.client.ws_connect(body["websocket_path"], headers=self._bearer(token))
|
||||
ready = await self._next_ws_event(ws)
|
||||
self.assertEqual(ready["type"], "voice.session.ready")
|
||||
body["_token"] = token
|
||||
body["_ready"] = ready
|
||||
return ws, fake_provider, body
|
||||
|
||||
async def _emit_tool_call(
|
||||
self,
|
||||
provider: FakeNativeProvider,
|
||||
*,
|
||||
call_id="call-1",
|
||||
mode: str | None = None,
|
||||
) -> None:
|
||||
arguments: dict[str, Any] = {"text": "Research the thing.", "session_id": "chat-123"}
|
||||
if mode is not None:
|
||||
arguments["mode"] = mode
|
||||
await provider.connection.emit(
|
||||
ProviderEvent(
|
||||
ProviderEventKind.FUNCTION_CALL_COMPLETED,
|
||||
response_id="resp-1",
|
||||
payload={
|
||||
"call": ToolCallEvent(
|
||||
call_id=call_id,
|
||||
name="hermes_run_task",
|
||||
arguments=arguments,
|
||||
)
|
||||
},
|
||||
)
|
||||
)
|
||||
|
||||
async def _read_until(self, ws, target: str, *, limit: int = 40) -> list[dict[str, Any]]:
|
||||
events: list[dict[str, Any]] = []
|
||||
for _ in range(limit):
|
||||
event = await self._next_ws_event(ws)
|
||||
events.append(event)
|
||||
if event["type"] == target:
|
||||
return events
|
||||
raise AssertionError(f"did not see {target}; saw {[e['type'] for e in events]}")
|
||||
|
||||
async def test_long_run_promotes_and_pump_stays_responsive(self) -> None:
|
||||
broker = GatedHermesToolBroker()
|
||||
ws, provider, _ = await self._open(broker=broker)
|
||||
try:
|
||||
await self._emit_tool_call(provider)
|
||||
events = await self._read_until(ws, "hermes.run.promoted")
|
||||
promoted = next(e for e in events if e["type"] == "hermes.run.promoted")
|
||||
self.assertEqual(promoted["tier"], "promoted")
|
||||
self.assertEqual(promoted["promote_after_ms"], 50)
|
||||
|
||||
# The pending provider call was closed with an interim background ack.
|
||||
self.assertTrue(provider.connection.tool_results)
|
||||
self.assertEqual(
|
||||
provider.connection.tool_results[0][1].get("status"),
|
||||
"running_in_background",
|
||||
)
|
||||
|
||||
# Prove the pump is NOT blocked: a provider event sent while the run
|
||||
# is still gated is processed and forwarded.
|
||||
self.assertFalse(broker.release.is_set())
|
||||
await provider.connection.emit(
|
||||
ProviderEvent(
|
||||
ProviderEventKind.INPUT_TRANSCRIPT_DELTA,
|
||||
payload={"delta": "still talking"},
|
||||
)
|
||||
)
|
||||
live = await self._read_until(ws, "voice.input_transcript.delta")
|
||||
self.assertTrue(any(e["type"] == "voice.input_transcript.delta" for e in live))
|
||||
|
||||
# Release the run -> it completes in the background and is spoken once.
|
||||
broker.release.set()
|
||||
done = await self._read_until(ws, "hermes.run.background_completed")
|
||||
completed = next(e for e in done if e["type"] == "hermes.run.background_completed")
|
||||
self.assertTrue(completed["ok"])
|
||||
|
||||
# The result is injected for the provider to summarize exactly once.
|
||||
for _ in range(50):
|
||||
summaries = [t for t in provider.connection.text_inputs if "final spoken answer" in t]
|
||||
if summaries:
|
||||
break
|
||||
await asyncio.sleep(0.02)
|
||||
self.assertEqual(len(summaries), 1)
|
||||
finally:
|
||||
broker.release.set()
|
||||
await ws.close()
|
||||
|
||||
async def test_short_run_is_not_promoted(self) -> None:
|
||||
broker = FastHermesToolBroker()
|
||||
ws, provider, body = await self._open(broker=broker)
|
||||
# Widen the grace window so the fast run finishes well within it.
|
||||
self._server().realtime_agent.sessions[body["session_id"]].promote_after_ms = 5000
|
||||
try:
|
||||
await self._emit_tool_call(provider)
|
||||
events = await self._read_until(ws, "voice.playback_drain.requested")
|
||||
types = [e["type"] for e in events]
|
||||
self.assertNotIn("hermes.run.promoted", types)
|
||||
# Real result delivered to the provider (not an interim background ack).
|
||||
self.assertTrue(provider.connection.tool_results)
|
||||
self.assertNotEqual(
|
||||
provider.connection.tool_results[0][1].get("status"),
|
||||
"running_in_background",
|
||||
)
|
||||
finally:
|
||||
await ws.close()
|
||||
|
||||
async def test_explicit_background_mode_promotes_immediately(self) -> None:
|
||||
# Tier C: mode="background" detaches immediately, even with grace-period
|
||||
# promotion turned off and a long grace window.
|
||||
broker = GatedHermesToolBroker()
|
||||
ws, provider, body = await self._open(broker=broker)
|
||||
session = self._server().realtime_agent.sessions[body["session_id"]]
|
||||
session.promotion_enabled = False
|
||||
session.promote_after_ms = 60000
|
||||
try:
|
||||
await self._emit_tool_call(provider, mode="background")
|
||||
events = await self._read_until(ws, "hermes.run.promoted")
|
||||
promoted = next(e for e in events if e["type"] == "hermes.run.promoted")
|
||||
self.assertEqual(promoted["tier"], "durable")
|
||||
|
||||
broker.release.set()
|
||||
done = await self._read_until(ws, "hermes.run.background_completed")
|
||||
self.assertTrue(
|
||||
any(e["type"] == "hermes.run.background_completed" for e in done)
|
||||
)
|
||||
finally:
|
||||
broker.release.set()
|
||||
await ws.close()
|
||||
|
||||
async def test_cancel_during_promoted_run(self) -> None:
|
||||
broker = GatedHermesToolBroker()
|
||||
ws, provider, _ = await self._open(broker=broker)
|
||||
try:
|
||||
await self._emit_tool_call(provider)
|
||||
await self._read_until(ws, "hermes.run.promoted")
|
||||
await broker.started.wait()
|
||||
|
||||
await ws.send_json({"type": "response.cancel"})
|
||||
cancelled = await self._read_until(ws, "hermes.run.cancelled")
|
||||
self.assertTrue(any(e["type"] == "hermes.run.cancelled" for e in cancelled))
|
||||
for _ in range(50):
|
||||
if broker.cancelled.is_set():
|
||||
break
|
||||
await asyncio.sleep(0.02)
|
||||
self.assertTrue(broker.cancelled.is_set())
|
||||
finally:
|
||||
broker.release.set()
|
||||
await ws.close()
|
||||
|
||||
async def test_detach_resume_replays_background_completed(self) -> None:
|
||||
broker = GatedHermesToolBroker()
|
||||
ws, provider, body = await self._open(broker=broker)
|
||||
await self._emit_tool_call(provider)
|
||||
await self._read_until(ws, "hermes.run.promoted")
|
||||
ready = body["_ready"]
|
||||
|
||||
# Detach: close the websocket while the run is still in the background.
|
||||
await ws.close(code=1001, message=b"network changed")
|
||||
session = self._server().realtime_agent.sessions[body["session_id"]]
|
||||
for _ in range(50):
|
||||
if session.detached_at is not None:
|
||||
break
|
||||
await asyncio.sleep(0.01)
|
||||
self.assertIsNotNone(session.detached_at)
|
||||
|
||||
# Complete the run while detached -> events recorded to the ring.
|
||||
broker.release.set()
|
||||
for _ in range(100):
|
||||
if any(
|
||||
e.get("type") == "hermes.run.background_completed"
|
||||
for e in session.event_ring
|
||||
):
|
||||
break
|
||||
await asyncio.sleep(0.02)
|
||||
|
||||
ws2 = await self.client.ws_connect(
|
||||
body["websocket_path"], headers=self._bearer(body["_token"])
|
||||
)
|
||||
try:
|
||||
await ws2.send_json(
|
||||
{
|
||||
"type": "session.resume",
|
||||
"resume_token": body["resume_token"],
|
||||
"last_event_id": ready["event_id"],
|
||||
"last_audio_event_id": 0,
|
||||
"last_played_audio_event_id": 0,
|
||||
"last_input_chunk_id": 0,
|
||||
}
|
||||
)
|
||||
events = await self._read_until(ws2, "voice.replay.done", limit=60)
|
||||
types = [e["type"] for e in events]
|
||||
self.assertIn("voice.session.resumed", types)
|
||||
self.assertIn("hermes.run.background_completed", types)
|
||||
replayed = next(
|
||||
e for e in events if e["type"] == "hermes.run.background_completed"
|
||||
)
|
||||
self.assertTrue(replayed.get("replayed"))
|
||||
finally:
|
||||
await ws2.close()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,151 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Phase 0 idle-tolerance probe for ADR 33 / the background-Hermes-runs plan.
|
||||
|
||||
Opens a provider-native realtime-agent socket (xAI or OpenAI) via the existing
|
||||
relay adapters, configures it, then holds it QUIESCENT — no ``response.create``,
|
||||
no input audio — for a series of idle windows. It logs whether the socket
|
||||
survives, any server-side close codes/timeouts, and any turn/VAD artifacts on the
|
||||
first post-idle ``response.create``.
|
||||
|
||||
This answers the only factual unknown that gates ADR 33 default-on: can a
|
||||
provider session hold the floor while a background Hermes run completes, or must
|
||||
that provider's Tier B path keep-alive / reopen the socket instead?
|
||||
|
||||
This is a DEV PROBE — it is not imported by the relay and ships nothing. Run it
|
||||
on the relay host where provider credentials are configured (env or xAI OAuth
|
||||
store), with the repo root on PYTHONPATH:
|
||||
|
||||
python scripts/realtime-provider-idle-probe.py --provider xai
|
||||
python scripts/realtime-provider-idle-probe.py --provider openai --windows 30,60,120
|
||||
|
||||
Record the per-provider verdict in docs/realtime-voice-poc.md ("Idle tolerance").
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import contextlib
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
# Allow running as a bare script from the repo root.
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from plugin.relay.realtime_agent.models import ( # noqa: E402
|
||||
HERMES_TOOL_SCHEMAS,
|
||||
ProviderEventKind,
|
||||
RealtimeAgentSessionConfig,
|
||||
)
|
||||
from plugin.relay.realtime_agent.providers.openai import ( # noqa: E402
|
||||
OpenAIRealtimeAgentProvider,
|
||||
)
|
||||
from plugin.relay.realtime_agent.providers.xai import ( # noqa: E402
|
||||
XAIRealtimeAgentProvider,
|
||||
)
|
||||
|
||||
_PROVIDERS = {
|
||||
"xai": (XAIRealtimeAgentProvider, "grok-voice-latest", "ember", 24000),
|
||||
"openai": (OpenAIRealtimeAgentProvider, "gpt-realtime-2", "alloy", 24000),
|
||||
}
|
||||
|
||||
|
||||
def _session_config(provider_id: str) -> RealtimeAgentSessionConfig:
|
||||
_, model, voice, rate = _PROVIDERS[provider_id]
|
||||
return RealtimeAgentSessionConfig(
|
||||
provider=provider_id,
|
||||
model=model,
|
||||
voice=voice,
|
||||
sample_rate=rate,
|
||||
profile=None,
|
||||
hermes_session_id=None,
|
||||
instructions="You are an idle-tolerance probe. Do not speak unless asked.",
|
||||
provider_options={},
|
||||
tools=HERMES_TOOL_SCHEMAS,
|
||||
)
|
||||
|
||||
|
||||
async def _drain_events(connection, stop: asyncio.Event, sink: list[str]) -> None:
|
||||
"""Record provider events (esp. ERROR/close) while we idle."""
|
||||
try:
|
||||
async for event in connection.events():
|
||||
sink.append(f"{time.monotonic():.1f}s {event.kind.value}")
|
||||
if event.kind == ProviderEventKind.ERROR:
|
||||
sink.append(f" ERROR payload={event.payload!r}")
|
||||
if stop.is_set():
|
||||
return
|
||||
except Exception as exc: # noqa: BLE001 - probe wants the raw failure
|
||||
sink.append(f"{time.monotonic():.1f}s events() raised {exc!r}")
|
||||
finally:
|
||||
stop.set()
|
||||
|
||||
|
||||
async def probe(provider_id: str, windows: list[int]) -> int:
|
||||
cls, *_ = _PROVIDERS[provider_id]
|
||||
provider = cls()
|
||||
print(f"[probe] connecting {provider_id} ...")
|
||||
connection = await provider.connect(_session_config(provider_id))
|
||||
print("[probe] connected + configured")
|
||||
|
||||
stop = asyncio.Event()
|
||||
sink: list[str] = []
|
||||
reader = asyncio.create_task(_drain_events(connection, stop, sink))
|
||||
|
||||
verdict = "hold-floor-ok"
|
||||
try:
|
||||
for window in windows:
|
||||
print(f"[probe] idling {window}s (no response.create, no audio) ...")
|
||||
try:
|
||||
await asyncio.wait_for(stop.wait(), timeout=window)
|
||||
# stop fired => the socket closed / errored while idle.
|
||||
print(f"[probe] socket DID NOT survive {window}s idle window")
|
||||
verdict = "must-reopen"
|
||||
break
|
||||
except asyncio.TimeoutError:
|
||||
print(f"[probe] survived {window}s idle")
|
||||
|
||||
# Probe a post-idle turn for VAD/turn artifacts.
|
||||
print("[probe] requesting a short post-idle response ...")
|
||||
before = len(sink)
|
||||
await connection.send_text("Say the single word: ready.")
|
||||
await asyncio.sleep(8)
|
||||
produced = sink[before:]
|
||||
got_audio = any("audio_delta" in line for line in produced)
|
||||
got_error = any("error" in line for line in produced)
|
||||
print(f"[probe] post-idle audio={got_audio} error={got_error}")
|
||||
if got_error or not got_audio:
|
||||
verdict = "needs-keepalive"
|
||||
with contextlib.suppress(Exception):
|
||||
await connection.cancel_response()
|
||||
finally:
|
||||
stop.set()
|
||||
with contextlib.suppress(Exception):
|
||||
await connection.close()
|
||||
reader.cancel()
|
||||
with contextlib.suppress(asyncio.CancelledError):
|
||||
await reader
|
||||
|
||||
print("\n[probe] event trace:")
|
||||
for line in sink:
|
||||
print(" " + line)
|
||||
print(f"\n[probe] VERDICT ({provider_id}): {verdict}")
|
||||
print("[probe] record this in docs/realtime-voice-poc.md -> Idle tolerance")
|
||||
return 0
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--provider", choices=sorted(_PROVIDERS), default="xai")
|
||||
parser.add_argument(
|
||||
"--windows",
|
||||
default="30,60,120",
|
||||
help="comma-separated idle windows in seconds",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
windows = [int(w) for w in str(args.windows).split(",") if w.strip()]
|
||||
return asyncio.run(probe(args.provider, windows))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -12,7 +12,7 @@ The plugin is a thin observer — it never modifies state, never writes to your
|
||||
|
||||
**On your server:**
|
||||
|
||||
- hermes-agent with the Dashboard Plugin System (upstream commit `01214a7f` on `axiom`, or any later `main` once [PR #8556](https://github.com/NousResearch/hermes-agent/pull/8556) and its dashboard followups merge). `hermes dashboard start` must already work for you.
|
||||
- hermes-agent with the Dashboard Plugin System (upstream commit `01214a7f` on `axiom`, or any later `main` that includes the dashboard plugin follow-ups). `hermes dashboard start` must already work for you.
|
||||
- The canonical Hermes-Relay install — if you ran the one-liner on the [Quick Start](/guide/getting-started), you're done. The installer symlinks `~/.hermes/plugins/hermes-relay` → the plugin subtree and the dashboard scanner picks up `plugin/dashboard/manifest.json` automatically.
|
||||
- A gateway restart after install: `systemctl --user restart hermes-gateway`.
|
||||
|
||||
|
||||
@@ -227,6 +227,29 @@ voice-output stream cannot compete with the realtime provider. The final spoken
|
||||
answer is generated by the realtime provider after Hermes returns a compact
|
||||
result, so tool output is summarized naturally instead of read aloud.
|
||||
|
||||
#### Background tasks
|
||||
|
||||
Some requests take a while — research, multi-step work, a long command. Instead
|
||||
of freezing the conversation until they finish, Realtime Agent **promotes** a
|
||||
slow run to the background: the agent says a short "I'm on it" and you can keep
|
||||
talking, ask something else, or just wait. When the task finishes, the agent
|
||||
speaks the answer. Asking for something explicitly long starts a background task
|
||||
right away.
|
||||
|
||||
You control this under **Voice Settings → Realtime Agent → Background tasks**:
|
||||
|
||||
- **Promote long tasks** — turn the behavior on or off. With it off, the agent
|
||||
waits silently until the task finishes (the old behavior).
|
||||
- **Spoken handoff** — whether the agent says a short acknowledgement when a task
|
||||
moves to the background, or just shows it on screen.
|
||||
- **When the answer is ready** — *Speak* it as soon as you're not mid-sentence,
|
||||
*Notify* and speak when you re-engage, or *Show only* (no spoken answer).
|
||||
|
||||
While a background task is running, a small "working on it" chip stays visible in
|
||||
the voice screen. You can cancel a background task at any time the same way you
|
||||
cancel any turn. Hermes still owns the task end-to-end — promotion only changes
|
||||
*when* the answer is spoken, never who runs the tools.
|
||||
|
||||
Provider-native Android paths stream mic PCM to a relay-owned realtime provider
|
||||
WebSocket session. Android commits the captured utterance, the active provider
|
||||
owns input transcription and speech generation, and Hermes still owns profile
|
||||
|
||||
@@ -55,7 +55,7 @@ All commands are fetched dynamically from the server where possible:
|
||||
- **Configuration**: `/model`, `/personality`, `/reasoning`, `/yolo`, `/verbose`, `/voice`
|
||||
- **Info**: `/help`, `/status`, `/usage`, `/insights`, `/commands`
|
||||
- **Personalities**: generated from server config (`config.agent.personalities`) — `/personality victor`, `/personality creative`, etc.
|
||||
- **Skills**: dynamically fetched from `GET /api/skills` — 90+ server skills grouped by category (creative, devops, research, etc.)
|
||||
- **Skills**: dynamically fetched from `GET /v1/skills` with fallback to legacy `GET /api/skills` — server skills grouped by category (creative, devops, research, etc.)
|
||||
|
||||
## Tool Execution
|
||||
|
||||
|
||||
@@ -18,9 +18,9 @@ Hermes-Relay voice endpoints can reuse this same Bearer token. Relay validates i
|
||||
|
||||
## How endpoints get served
|
||||
|
||||
Installing the plugin via `install.sh` is enough to make all of the endpoints below work — including the management ones (`/api/sessions/*`, `/api/memory`, `/api/skills`, `/api/config`, `/api/available-models`). The plugin wires the gateway up at install time so these are served on the same `:8642` host as the standard `/v1/*` endpoints, with the same `Authorization: Bearer …` auth.
|
||||
Current upstream hermes-agent includes the baseline session API (`/api/sessions/*`, chat, stream, fork, messages) and native capability discovery (`/v1/skills`, `/v1/toolsets`). Installing the Relay plugin via `install.sh` keeps older or partial core builds compatible by injecting only the missing management endpoints Relay still consumes, such as `/api/sessions/search`, `/api/memory`, `/api/config`, legacy `/api/skills` detail routes, `/api/available-models`, and voice aliases. These are served on the same `:8642` host as the standard `/v1/*` endpoints, with the same optional `Authorization: Bearer ***` auth.
|
||||
|
||||
**Chat streaming uses standard `/v1/runs`** by default — it emits structured `tool.started`/`tool.completed` SSE events for live tool progress cards in the Android app. The app's `Settings → Chat → Streaming endpoint = "Auto"` (default) probes per-endpoint capability and picks the best chat path automatically; you can manually force `Sessions` or `Runs` mode for debugging.
|
||||
**Chat streaming uses standard `/v1/runs`** by default — it emits structured `tool.started`/`tool.completed` SSE events for live tool progress cards in the Android app. The app's `Settings → Chat → Streaming endpoint = "Auto"` (default) probes per-endpoint capability and picks the best chat path automatically; you can manually force `Sessions` or `Runs` mode for debugging. Skill discovery prefers `GET /v1/skills` and falls back to legacy `GET /api/skills`.
|
||||
|
||||
## Endpoints
|
||||
|
||||
|
||||
Reference in New Issue
Block a user