Compare commits

...
Author SHA1 Message Date
Bailey Dixon 3a99842011 Release v1.2.0 (android + plugin) (#95)
Merge dev -> main for android-v1.2.0 and plugin-v1.2.0.
2026-06-21 17:31:30 -04:00
Bailey Dixon d261a1c374 feat: support static pet packs 2026-06-21 17:18:24 -04:00
Bailey DixonandClaude Opus 4.8 cf30b0dbc2 ci(android): auto-publish Play Store listing on main pushes
The Play Store Listing workflow now publishes the listing (screenshots,
graphics, and text) automatically when its path-scoped assets change on main,
in addition to manual workflow_dispatch. PRs and dev pushes still validate
only, and it skips gracefully (a notice, not a failure) when the
PLAY_SERVICE_ACCOUNT_JSON secret is absent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 16:41:25 -04:00
Bailey DixonandClaude Opus 4.8 8ea813d8d8 docs(android): document the deterministic screenshot harness
Add a "Deterministic rendering" section to docs/screenshot-automation.md (run
command, how to add a view, real-screen vs curated-frame for config/data
screens, the JDK-21 and no-plugin gotchas, and the Play-listing publish flow),
plus a CLAUDE.md Key Files pointer so the harness is discoverable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 16:41:05 -04:00
Bailey DixonandClaude Opus 4.8 a806726cb2 docs(android): add App Themes gallery to the user docs
New Themes feature page showing the eight-theme gallery (the same chat reskinned
by every theme), wired into the docs sidebar and the features index.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 16:40:45 -04:00
Bailey DixonandClaude Opus 4.8 45519e9fc8 chore(android): refresh 1.2.0 store screenshots (deterministic 1:1 renders)
Regenerate all eight phone screenshots host-side at exact 2:1; replace the
command-palette and settings scenes with App Themes and Appearance (the latter
the real AppearanceSettingsScreen, rendered 1:1). Re-export the Play graphics
and README grid; screenshots.py validate is clean (no 2:1 crop warnings).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 16:40:24 -04:00
Bailey DixonandClaude Opus 4.8 7746d7de98 test(android): add Roborazzi host-side screenshot harness
Deterministic, device-free store/docs screenshot renderer: renders real
screens/components with mock data at exactly 1080x2160 (no Play 2:1 clipping).
Drops the AGP-9-incompatible Roborazzi Gradle plugin (keeps the runtime) and
runs unit tests on JDK 21 for the markdown code-highlighter.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 16:40:01 -04:00
Bailey DixonandClaude Opus 4.8 3bec0d22b8 release(plugin): plugin-v1.2.0
Bump plugin/dashboard metadata to 1.2.0 (in sync). Release notes cover the
relay enhancement layer + agent-context injection (sensitive-media block,
/context/injected audit, dashboard toggles, default-on), provider-aware
enhanced voice (Gemini + xAI), isolated TUI-tuned tmux, and voice cleanup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 23:24:48 -04:00
Bailey DixonandClaude Opus 4.8 222fab4fb9 release(android): android-v1.2.0
Promote [Unreleased] -> [1.2.0]; backfill the agent-pet system, in-app
crash reporting, per-profile icons, clean mode, permissions screen, the
"Standard"->"Vanilla Hermes" rename, and PDF/image crash fixes that
shipped to dev without changelog bullets. appVersionCode 13 -> 14.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 23:24:33 -04:00
Bailey DixonandClaude Opus 4.8 b27f3a0a7d docs(android): pet kit — frames must visibly animate, not just register
A generation-method change over-corrected: the registration/safe-box prompt
produced 16 near-identical frames (measured interframe diff ~0.02/255), so pets
rendered static even though frameCount is 16 and the renderer cycles all of
them. Clarify across the prompt template, gotchas, and pet-spec that the cells
are an animation, NOT copies — lock only the identity/anchor (position+scale),
but the moving parts (eyes, mouth, hands, hair, accent) must visibly progress
through the full motion arc across all 16 frames; over-locking is its own
distinct failure. Also carries the chroma-key + safe-box authoring guidance.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 23:05:51 -04:00
Bailey DixonandClaude Opus 4.8 3cbf0333ae fix(android): drop the customized ring when a profile icon image is shown
The 2dp "customized" ring is meant to mark the letter avatar; on an actual
profile photo it just looks like a bad outline. Suppress it whenever
LocalAgentIconPath is set (sheet header + Settings); the ring still shows for
the letter fallback.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 22:53:54 -04:00
Bailey DixonandClaude Opus 4.8 954d2522ed feat(android): use the per-profile icon for header/navbar avatars too
The profile icon only reached the per-message label; the circular header avatars
(agent sheet, chat top bar, Settings) still showed the generated letter. Add a
shared AgentAvatarFace that renders the LocalAgentIconPath image when set, else
the name's initial, and use it in all three. The chat header keeps its letter
cross-fade for the no-icon case (image short-circuits before AnimatedContent).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 22:41:00 -04:00
Bailey DixonandClaude Opus 4.8 fa973dd2df docs(todo): mark per-profile icon + static-image avatar shipped; ignore build logs
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 22:31:09 -04:00
Bailey DixonandClaude Opus 4.8 d827e460e0 feat(android): static-image avatars + per-profile agent icon (client-side)
Two custom-identity features (they share ConnectionViewModel, so one commit):

Static-image avatar: "Add a pet" now accepts a single image (PNG/JPG/GIF/WebP),
detected by magic bytes, and auto-wraps it as a one-frame static pet (idle.png +
a synthesized minimal pet.json) — a custom avatar with no manifest authoring.
importZip -> importUri; importPetFromZip -> importPet.

Per-profile agent icon: a client-side twin of ProfileDisplayAliasStore. New
ProfileIconStore (own DataStore, keyed per (connection, profile), never sent to
Hermes) holds a path to an image copied into files/profile-icons/ (not a SAF
URI, so it survives without persistable permission). Wired through
ProfileController next to profileDisplayAlias, exposed on ConnectionViewModel,
provided at the app root as LocalAgentIconPath, and rendered as a small circular
Coil image beside the agent name in MessageBubble. Picker (AgentIconRow) sits
under the local-name row in ConnectionInfoSheet. Scope: small name-adjacent icon
only; the big avatar stays global. Tests for both; PetImporter image-wrap +
ProfileIconStore scoping/clear.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 22:29:51 -04:00
Bailey DixonandClaude Opus 4.8 3cd8791ce4 feat(android): in-app pet state preview in Appearance
Testing a pet meant inducing each state by driving the agent (run a tool for
working, fail a turn for error, start voice for speaking). Add a preview under
the speed/stabilize controls (pet selected only): a ~140dp canvas rendering the
active pet, a FilterChip row for the seven sustained states, and Greet/Done
buttons that replay the one-shots. Pure UI on the existing AgentAvatar seam —
no new ViewModel/pref/renderer; it calls activeAvatar.Render(AvatarRenderState(
state=...)) with a user-picked state, so it also reflects the live speed and
stabilize settings. Working = Thinking + toolCallBurst; Greet remounts via key;
Done drives a momentary Speaking->Idle transition.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 21:50:31 -04:00
Bailey DixonandClaude Opus 4.8 d1bf6245fd feat(android): auto-stabilize pet frames (re-center on content)
AI-generated sprite sheets keep a character's appearance consistent but not its
position/scale across cells, so the pet floats/jumps as it plays (audited: 34px
vertical drift over 16 cells, 8/16 frames touching the cell edge). Add decode-
time stabilization: scan each frame's opaque pixels (alpha bbox) and shift the
draw so the content's center sits at the cell center. Works for sheets (per
cell) and sequences (per bitmap); one-time scan on IO with a reused buffer.

Exposed as a global LocalPetStabilize (pet_stabilize pref) with a "Stabilize
frames" Switch in Appearance, default on. Keys the decode produceState so
toggling re-decodes. Fixes an installed pet at render time with no re-import.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 21:04:29 -04:00
Bailey DixonandClaude Opus 4.8 f083ceacf0 docs(android): stress frame registration in the pet prompt kit
On-device audit of a 4x4 pet showed the character's vertical center drifting
34px across the 16 cells with 8/16 frames touching the cell edge — the image
model kept appearance consistent but not position/scale, so the pet floats and
the next frame's edge bleeds in. The renderer slices/centers exact cells
faithfully, so this is an authoring (registration) gap, not an engine bug. Add
registration instructions to the prompt template (lock head/shoulders, same
position + scale, only small secondary motion) and the consistency caveat
(registration degrades with cell count; drop to 3x3/2x2 if a 4x4 drifts).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:48:41 -04:00
Bailey Dixon a43d395108 Merge: transport-tier stepper + dashboard default-on into dev 2026-06-20 20:40:23 -04:00
Bailey DixonandClaude Opus 4.8 47d4d4f532 feat(android): transport-tier stepper in session details + dashboard default-on display
Android: SessionPathDetails (agent sheet → Connection) gains a vertical basic→best transport ladder (Completions → Runs → Sessions → Gateway) via a new TransportTierStepper, using the same resolveChatTransportStatus as the status badge — active tier filled+highlighted, server-unsupported tiers muted, with the resolver's reason beneath. Dashboard: the Agent-context toggles now read as ON when the env is unset (matching the new config default) via a strict-bool coercion, and the label says 'On by default for relay installs'; dist rebuilt.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:40:00 -04:00
Bailey DixonandClaude Opus 4.8 242665348d docs(android): pet cell-resolution guidance (256px cells, size for biggest surface)
Pixelation is a resolution axis (cell px), separate from smoothness (frame
count): one frame set is contain-fit into every surface, so author for the
largest (the full-screen chat background) and small placements (voice overlay)
downscale and stay sharp. Bump the kit default to 256px cells (a 1024x1024
sheet for 4x4), note 512px is fine for a sprite sheet (one bitmap), and that
the old "<=256px" note was for frame-sequences. Updates custom-avatars.md,
pet-prompt-kit.txt, pet-spec.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:20:41 -04:00
Bailey DixonandClaude Opus 4.8 5544c23f05 feat(android): pet playback-speed control in Appearance
A pet that feels too fast/slow needed re-authoring + re-importing to tune. Add a
global playback-speed multiplier (pet_speed pref, 0.5x-1.5x, default 1.0) as a
Slider in Appearance, shown when a pet is selected. It's provided at the app
root via a new LocalPetPlaybackSpeed composition local and read live in
PetAvatar.Render (rememberUpdatedState), so dragging it re-times the pet
instantly with no restart. Applies to every clip including one-shots and
composes with intensity (baseFps * speed * intensityFactor, clamped 1-60). The
sphere avatar ignores it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:20:08 -04:00
Bailey DixonandClaude Opus 4.8 3d5a94d818 docs: correct default-on for relay agent-context injection (CHANGELOG/DEVLOG)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:18:56 -04:00
Bailey DixonandClaude Opus 4.8 aeaf7282f3 feat(relay): enable agent-context injection by default for relay installs
The relay plugin install is itself the opt-in, and the wrap is fail-open, auditable (chat 'Relay context (server-side)'), and reversible from the dashboard toggle — so default the master + media-sensitivity gates ON. Vanilla upstream (no plugin) is unaffected; set RELAY_AGENT_CONTEXT_ENABLED=0 to opt out. Tests updated for the new default + explicit-off coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:10:20 -04:00
Bailey DixonandClaude Opus 4.8 fc8aaff749 docs(android): default pet kit to 4x4 (16-frame) sheets for smooth motion
A 2x2 (4-frame) sheet reads steppy at any fps. The renderer already slices any
N×M grid (decodeClip derives cols/rows from sheet size / cell size; drawPetFrame
indexes col=i%cols, row=i/cols), so "support 4x4" is an authoring default, not a
renderer change. Default the kit to a 4x4 grid (16 frames): prompt template,
manifest example, and pet-prompt-kit.txt now use frameCount 16 with fps matched
to the count (idle ~8 -> ~2s loop); 2x2/4 stays documented as the
easier-consistency fallback. pet-spec notes any rectangular grid works. Adds a
PetLoaderTest case for a 16-frame sheet.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:06:27 -04:00
Bailey DixonandClaude Opus 4.8 27e62ff768 fix(android): smooth pet frame loop (remove double-wait frame skip)
PetAvatar.Render's frame loop awaited withFrameNanos (one vsync) AND
delay(1000/fps) each iteration, so every frame waited ~16ms longer than its
duration; the surplus accumulated until the loop skipped a frame to catch up —
a periodic hitch, worst at low fps. Drop the delay: withFrameNanos already
paces the loop at vsync, and the accumulator advances the sprite only when a
frame's worth of real time has elapsed, so playback is smooth and intensity's
variable rate no longer causes skips.

Also document that smoothness comes from frame count (8-16), not fps, and to
match fps to count (calm states 3-4); lowered the example/kit idle+listening
fps to 4.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 19:48:00 -04:00
Bailey DixonandClaude Opus 4.8 c4b1a02ba5 docs(todo): relay enhancement-layer follow-ups + retirement notes
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 15:22:22 -04:00
Bailey Dixon f3fc8e557c Merge: relay enhancement layer + agent-context injection into dev
# Conflicts:
#	DEVLOG.md
2026-06-20 15:09:25 -04:00
Bailey DixonandClaude Opus 4.8 b581756ffd docs(relay): enhancement-layer design + structured-media plan + DEVLOG/CHANGELOG
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 15:06:26 -04:00
Bailey Dixon 3ae2495188 feat: audit relay context and chat transport 2026-06-20 15:01:15 -04:00
Bailey DixonandClaude Opus 4.8 092d6a0c8f feat(android): in-app add/remove/refresh for custom pet avatars
Appearance could select avatars but not add or remove a pet — the only path
was adb push into app-scoped external storage, which scoped storage stalls on
(confirmed hanging on a Samsung device). And the avatar list loaded once at
startup, so even a pushed pet never appeared without a restart; users saw only
the Sphere.

- PetImporter (new): "Add a pet" launches a SAF .zip picker and unpacks into
  pets/. Hardened with a zip-slip guard, per-file/total/count ceilings, and
  post-extract validation through the same PetSpec.toAvatar the loader uses.
- PetLoader.deletePet: remove a pack by resolved manifest id, behind a confirm
  dialog; falls back to the Sphere if the deleted pet was selected.
- Live refresh: an avatarsRefreshTick keys the avatar produceState in RelayApp,
  so import/delete and opening Appearance re-scan pets/ without an app restart
  (resolves the process-scoped-load TODO). Results surface as snackbars.
- AppearanceSettingsScreen: "Add a pet" + "Rescan" buttons and an
  "Installed pets" management list with per-pet remove.
- Tests: PetImporterTest (root/nested import, no-manifest, missing-idle,
  zip-slip refused) and PetLoaderTest delete cases.

Built and installed to the sideload debug build; new unit tests pass (the 12
build failures are the pre-existing DataStore/FileStorage JVM cases).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 14:56:49 -04:00
Bailey DixonandClaude Opus 4.8 3c0f6f6cca docs(android): AI pet authoring kit + JSON schema for custom avatars
Pets are pure data, so the only barrier to making one is sourcing the art.
Document an AI-generation workflow plus a machine-readable contract:

- A reference-image-first, character-agnostic prompt template
  ({character}/{style}/{accent}) and a per-state motion table mapping image
  generation onto the agent-state vocabulary, a full 9-state manifest, and a
  one-download pet-prompt-kit.txt.
- A draft-07 JSON Schema (user-docs/public/pet.schema.json) mirroring the
  loader structural rules (required idle, frames-XOR-sheet, positive sheet
  dims) so editors and AI agents can validate a pet.json before installing it.
- A vendor-neutral "let an AI agent build the pack" callout (Codex/Claude Code
  as examples) stating the acceptance criteria and image-gen prerequisite.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 14:56:25 -04:00
Bailey Dixon 41b341ed09 feat(plugin): add relay agent context injection 2026-06-20 14:50:49 -04:00
Bailey DixonandClaude Opus 4.8 b5fd63bb93 fix(android): stop PDF viewer crash when document closes mid-measure
PdfDoc.pageCount was a lazy getter delegating to PdfRenderer.pageCount, so a LazyColumn measure pass racing DisposableEffect's onDispose { doc.close() } could call getPageCount() on an already-closed renderer -> IllegalStateException 'Document already closed' (caught in the wild by the crash reporter). Capture pageCount once at open time (a PDF's count is immutable) so it never reads the renderer after close, and skip page render when the doc is already closed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 13:34:30 -04:00
Bailey DixonandClaude Opus 4.8 5077ddd244 fix(android): render server-local chat images whose path has a space
An agent referenced /mnt/.../Coralee Adshade/undressher.jpg three ways and none rendered — all because of the space in the path:

- MEDIA:/path bare marker used /\S+, which stops at the space, so the marker never matched and showed as raw text. Now /.+? (allows spaces; OkHttp re-encodes for /media/by-path).

- ![](<path with spaces>): the markdown angle-bracket URL form wasn't accepted — the regex kept the leading '<' and stopped at the space, failing the startsWith("/") server-local check. Regex now accepts <...> and normalizeImageSrc strips the brackets.

- ![](/path%20encoded): the percent-encoded space wasn't decoded, so the relay looked up a literal '%20' directory and 404'd. normalizeImageSrc now percent-decodes absolute paths (protecting a literal '+').

Verified on-device: the previously-raw MEDIA: line now renders the image. File and relay were fine; this was entirely client-side path handling.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 13:11:49 -04:00
Bailey DixonandClaude Opus 4.8 52990aaf37 feat(android): keep crash report until acknowledged, not just first view
CrashReportGate consumed (read+deleted) the report on first read, so it vanished after one glance even if the user never acted on it. Switch to peek-on-read + clear-on-acknowledge: the report now survives relaunches until the user Dismisses or Reports it (Copy keeps it available), so a crash you saw but didn't report isn't lost.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 13:01:33 -04:00
Bailey DixonandClaude Opus 4.8 133a785839 fix(android): stop crash on server-local chat images (kotlin.Result in suspend)
RelayServerImage crashed on app open with 'kotlin.Result cannot be cast to byte[]': the resolver returned Result<ByteArray> from a suspend fun, and runCatching { fetch() } nested Result-in-Result, which Kotlin's value-class Result collapses incorrectly at runtime. Replace the suspend fetch path's kotlin.Result with a purpose-built ServerImageResult sealed type (Success/Failure) so the resolver boundary never returns kotlin.Result from a suspend function.

Caught in the wild by the new in-app crash reporter (Galaxy S25 Ultra, SDK 36).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 12:55:56 -04:00
Bailey DixonandClaude Opus 4.8 b1a0a7b21d fix(android): use GitHub's stable title+body params for crash-report prefill
Issue-form field-id prefill (template=bug_report.yml&<id>=...) is a GitHub public-preview feature and silently did not apply — only the title carried. Switch to the stable classic ?title=&body=&labels=bug route (blank_issues_enabled is true), with a markdown body that mirrors the form's sections (Affected area / What happened / Environment / Crash) plus a sanitization reminder, so the auto-captured report reliably prefills.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 12:23:50 -04:00
Bailey DixonandClaude Opus 4.8 a455e4688f feat(android): in-app crash reporting + QR camera hardening for foldables
Add privacy-respecting crash capture (no Firebase): an uncaught handler persists a structured report then re-raises so the system dialog and Play Android vitals still collect it. On next launch a show-once dialog offers Copy + a pre-filled GitHub bug_report.yml issue with device/version/trace.

Harden QrPairingScanner camera init — try/catch around ProcessCameraProvider.get() (main thread) and InputImage.fromMediaImage() (analyzer thread), with a graceful CameraUnavailableCard -> manual pairing fallback instead of a force-close. Addresses a Galaxy Z Fold7 'keeps crashing during setup' report.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 12:04:34 -04:00
Bailey DixonandClaude Opus 4.8 60093e383d feat(android): pet intensity modulation — clip speeds up under load
Complete the pet reactivity story (voice · tools · activity). The activity
ramp (intensity, ~0.7 while streaming) was already fed to every avatar but
pets ignored it; now an opt-in pet quickens its clip as the agent works.

- Live playback-rate modulation in PetAvatar.Render, opt-in via
  reactive.intensity: the base/working loop's fps scales by
  1 + intensity*PET_INTENSITY_RATE (0.6 -> ~1.4x typical, 1.6x peak, capped at
  PET_MAX_FPS). Read live via rememberUpdatedState so speed tracks the agent
  mid-clip without restarting the long-lived frame loop (re-keying on a
  continuously-animated float would thrash). One-shots excluded (!playOnce) so
  greet/done keep their authored rate.
- Flipped PET_RENDERER_CAPABILITIES.intensity to true; the loader's existing
  reactive.intensity && capability formula now lets a declared intensity:true
  through, so the pet honestly advertises Activity. No loader change.
- Tests: declared intensity is honored (Voice · Activity); split the prior
  clamp test so tools-without-a-working-clip still stays off the badge.
- docs/pet-spec.md: intensity row rewritten from Reserved to the speedup
  behavior; removed from Forthcoming (only attention remains there).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 10:16:38 -04:00
Bailey DixonandClaude Opus 4.8 d5a1ef54f0 feat(android): pet one-shot reaction layer (greet + celebrate)
The event tier of pet behavior: a reaction clip that plays ONCE over the base
loop, then returns — the touch that turns a status display into a character
(cf. the Peon Pet's celebrate-on-finish).

- Pet-local triggers, no host plumbing: reactions ride the activity-state
  transitions the avatar already sees. PetOneShot.Greet fires on first
  composition (the pet appears); PetOneShot.Done fires when a productive turn
  ends (Streaming/Speaking -> Idle; Thinking/Error -> Idle don't celebrate).
  Both opt-in (only if the pet ships the clip) and require >= 2 frames.
- Play-once-then-revert in PetAvatar.Render: a `playOnce` frame mode runs the
  clip 0->end (no modulo wrap), parks on the last frame, clears the active
  reaction, and recomposition hands back to the base loop. A reaction overlays
  everything (incl. working). Suppressed under reduced motion; an
  ONE_SHOT_MAX_MS (4s) backstop guarantees it never lingers on decode failure.
- PetLoader resolves friendly aliases (greet/wake, done/celebrate) from explicit
  `states` keys only (no fallback). One-shots are reactions, not a reactivity
  signal, so they don't touch the picker badge.
- Test: a pack with greet/done keys loads and the badge stays Voice (no
  accidental Tools/Activity coupling). Render-time playback is on-device/
  Compose-test territory (flagged in TODO).
- docs/pet-spec.md: new "One-shot reactions" section (Greet/Done table, opt-in,
  play-once, reduced-motion), an Expressive authoring tier. `attention`-on-
  notification stays Forthcoming (needs a host event the avatar lacks).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 10:03:03 -04:00
Bailey DixonandClaude Opus 4.8 217daeddf1 feat(android): pet working/tool-use overlay reacting to tool calls
Give pets a distinct "agent is running a tool" behavior, separate from
thinking — the strongest cross-system convention (MS Agent Think vs Process;
pi-animations Thinking·Working·Tool) is that acting should look different
from thinking.

- Pet-local overlay derived from the already-plumbed toolCallBurst, NOT a 7th
  SphereState — zero blast radius on the Sphere or call sites. PetAvatar.Render
  swaps to an optional workingClip when toolCallBurst >= 0.5 during a
  thinking/writing turn, releasing ~600ms after the last tool as the burst
  decays. Error keeps its own clip; burst is ~0 outside tool activity.
- Opt-in + clip-driven: workingClip resolves only from an explicit `working`
  key (no fallback). Shipping one IS the tool-reactivity capability — it drives
  both the swap and the Tools badge (reactivity.tools = workingClip != null &&
  PET_RENDERER_CAPABILITIES.tools), so the declared reactive.tools flag is no
  longer needed and can't over-promise. Flipped PET_RENDERER_CAPABILITIES.tools
  to true.
- Tests: a working clip lights the Tools badge; a working clip with missing
  files does not; declared-but-no-clip still clamps to Voice.
- docs/pet-spec.md: `working` moved from Forthcoming into the implemented model
  (state-table row, "working overlay" subsection, Rich tier = 7 clips,
  reactivity table tools row). Forthcoming trimmed to one-shots + intensity.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 09:52:43 -04:00
Bailey DixonandClaude Opus 4.8 f6b0afec9f feat(android): honest pet reactivity badge + behavior-model spec
The pet picker badge read reactivity straight from pet.json, so a manifest
could advertise tools/intensity the renderer never delivered. Clamp the
effective reactivity to what the renderer actually honors, and document a
real agent-state -> behavior model so pets can show thinking/writing/etc.

- PetAvatar.PET_RENDERER_CAPABILITIES: single source of truth for the live
  signals Render consumes today (voice only). PetLoader.toAvatar clamps a
  pet's reactivity to declared-AND-supported, so the badge can't over-promise.
- Friendly `writing` clip alias for the Streaming (output) state; tidied the
  Speaking/Error fallback chains. Backward compatible.
- docs/pet-spec.md: new "Agent states & pet behavior" section — state meanings,
  friendly clip-key vocabulary + fallback chains, a Minimal->Rich authoring
  ladder, and a "Forthcoming behavior" tier (working/tool clip, one-shot
  reactions, intensity modulation) grounded in prior art (MS Agent .acs set,
  pi-animations, Peon Pet). Reactivity table notes the clamp.
- PetLoaderTest: declared tools/intensity are dropped from the badge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 09:45:48 -04:00
Bailey DixonandClaude Opus 4.8 11b0bb391d fix(chat): paint a reopened session's real model from the session.resume result
On reopen/prewarm the gateway already returns the session's model in the
session.resume RPC result's `info`, but the client read only `session_id` and
discarded it — so the header/picker showed the global DEFAULT until the first
turn's async session.info arrived (~15-30s later), though the send itself
correctly used the session's stored model. Read info.model/provider/effort/yolo/
fast/usage from the resume result (resumeForPrewarm + ensureSession) into the same
_server* flows the session.info event feeds, via a shared applySessionInfo helper,
so a reopened session shows its actual model immediately.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 23:01:52 -04:00
Bailey DixonandClaude Opus 4.8 60eb993b15 fix(android): make side-loaded avatars/skins reachable + unify storage
The only documented way to install a pet avatar or sphere skin was an
`adb push` to external app-scoped storage, but both loaders read from
internal `filesDir` (`/data/data/<pkg>/files/`), which is not
`adb push`-able on a non-rooted device. The documented side-load path
could never work on either flavor.

- UserContentDir (new): shared resolver preferring external app-scoped
  storage (getExternalFilesDir, the /sdcard/Android/data/<pkg>/files/
  path adb push reaches, no runtime permission on API 19+) with internal
  filesDir fallback. Single source of truth for where pets AND sphere
  skins live — fixes the bug once for both.
- PetLoader / SphereSkinLoader: resolve through UserContentDir; add pure
  load(dir: File) overloads so the validation/skip-invalid logic is
  unit-testable without an Android Context.
- PetLoaderTest (17) + SphereSkinLoaderTest (6): parse, id/label
  fallbacks, schema + missing-idle + missing-file rejection, the
  safeChild path-traversal guard, fps clamping, one-bad-pack isolation,
  sort order, empty/absent dirs.
- AppearanceSettingsScreen: "Add your own pet" pointer so the feature is
  discoverable with no pets installed (mirrors the sphere-skin pointer).
- docs/pet-spec.md + docs/sphere-spec.md: correct the storage prose, both
  flavor paths, cross-link the two specs, fix an "Agent sphere" naming
  drift, add undecodable-image + per-frame-memory authoring caveats.
- user-docs/features/custom-avatars.md (new) + nav: user-facing page on
  the avatar→skin model, reactivity badges, adding skins/pets, reduced
  motion, troubleshooting.

Follow-ups (TODO.md): per-frame memory cap/downsample, decoded-clip cache.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 22:50:30 -04:00
Bailey DixonandClaude Opus 4.8 6abb28e7ce fix(chat): model picker "Server default" caption shows the real default, not the override
The in-chat picker's "Server default" row captioned itself from fallbackModelDetail
(gatewayCurrentModel ?? profile ?? serverModelName). selectModel() force-sets
gatewayCurrentModel to the active override, so once you picked a model the row read
"Current: <your override>" — presenting the override AS the server default. Caption
it from serverModelName (/api/config, never touched by overrides) instead — the same
source the agent drawer already uses correctly. The selected-row highlight was already
right; only the caption was wrong.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 20:55:30 -04:00
Bailey DixonandClaude Opus 4.8 bb1beed488 fix(media): surface server-image fetch failure reason; gate media badge on pairing
Server-local agent images (markdown ![](/abs/path) via /media/by-path) rendered a
generic "this image is on the server" placeholder on ANY failure, hiding why. The
resolver now returns Result<ByteArray>, the failed phase carries the reason, and
the inline notice shows it (sandbox 403 / not-found 404 / unauthorized / decode /
unsupported path) for debugging. Also gate mediaUrlConfigured() (the media-
capability badge + SSE media hint) on a current paired token, not just a relay
URL, so the badge agrees with what the fetch can actually do.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 20:49:25 -04:00
Bailey DixonandClaude Opus 4.8 8537f75ab1 feat(chat): clean-mode text persists and slides up; persistent new-chat hint
AgentTextFlow no longer fades lines away — they slide in and PERSIST, scrolling
up within a bounded ~1/3-screen viewport with a soft top-edge fade so the avatar
above stays unobstructed (a calmer, minimal accumulate-and-scroll feel rather
than ephemeral disappearing text). The clean-mode discoverability hint is now a
persistent pill shown ONLY on the empty/new-chat view, replacing the timed popup
that re-fired too often.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 18:59:55 -04:00
Bailey DixonandClaude Opus 4.8 68a6ff6c00 fix(chat): bind the picked model on a new gateway chat
createNewChat pre-created an api_server session for both transports; on the
gateway that handed the next turn a concrete id, forcing ensureSession down the
session.resume branch (the api_ id resumes against the shared launch state.db on
the default profile), which bypasses the model/provider/effort/fast binding that
only runs on session.create. New chats therefore ran the DEFAULT model while the
picker still showed the last pick. On the gateway transport, drop the gateway
session + null the id so the next send hits session.create and binds the
carried-over model. SSE keeps pre-creating (it needs a concrete id). Also fixes
the same latent effort/fast gap on new gateway chats.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 18:59:54 -04:00
Bailey DixonandClaude Opus 4.8 c8e8d67560 feat(android): allow image/attachment viewers to rotate to landscape
The full-screen ChatImageViewer and AttachmentViewer call AllowDeviceRotation()
(SENSOR) while open, overriding the app-wide portrait lock so wide images and
video can be viewed in landscape; portrait is restored on dismiss.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 18:27:11 -04:00
Bailey DixonandClaude Opus 4.8 349bee04ae feat(android): lock app to portrait orientation
Single-activity app, so screenOrientation=portrait on MainActivity locks the
whole app. tools:ignore for the deliberate LockedOrientationActivity lint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 18:21:24 -04:00
Bailey Dixon 3f51c23969 Merge: chat clean-mode + swappable avatar/pets into dev 2026-06-19 18:06:29 -04:00
Bailey DixonandClaude Opus 4.8 024515678e feat(chat): clean text-flow mode + swappable avatar with pet plugin system
Clean mode: long-press the chat background enters a full-screen ambient mode
(evolved from ambientMode) — a centered agent avatar with the assistant reply
flowing in as themed monospace text that materializes, dwells, and fades (bounded
6-line buffer), a thin composer, explicit exit, and full reduced-motion/TalkBack
fallbacks to static readable text.

AgentAvatar seam: a swappable AgentAvatar { Render(AvatarRenderState, modifier) }
with SphereAvatar as the default (the morphing sphere + its skin system nested
unchanged). Every sphere call site (chat, clean mode, voice overlay, onboarding,
splash) routes through LocalAgentAvatar; the Appearance picker is now "Agent avatar".

Pets: users can drop animated avatars in files/pets/<id>/pet.json (frame-sequence
or sprite-sheet, no new deps - off-thread BitmapFactory + rate-capped Canvas loop),
selected via an agent_avatar pref (mirrors sphere_skin) and persisted/switched in
Appearance. Fresh install with no pets behaves exactly as today. See docs/pet-spec.md.

Spec: docs/plans/2026-06-18-chat-clean-mode-and-pets.md
Follow-ups in TODO.md: process-scoped pack load, clip re-decode flash, tools/intensity pet reactivity.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 18:03:56 -04:00
Bailey Dixon 378a50eaf0 Merge: voice overhaul (overlay fixes, per-profile voice, settings IA, waveform output sync) into dev 2026-06-19 17:00:49 -04:00
Bailey DixonandClaude Opus 4.8 43135fe4b9 feat(voice): overlay fixes, per-profile voice, settings IA, and waveform output sync
Overlay: kill click-through (focus-mode pointer-consuming scrim + gesturesEnabled=
!voiceMode on the drawer), de-wrap the topbar (trimmed collapsed header + FlowRow
pills), and add a gear link to Voice Settings that exits voice mode before navigating.

Per-profile voice: VoicePreferencesRepository is now scope-aware — engine mode,
audio route, and the enhanced overrides namespace per (connection, profile) and
layer over global defaults; ergonomic prefs stay global. VoiceViewModel re-seeds
on profile change. The relay path already carried per-profile voice end-to-end.

Settings IA: single Voice scope banner, a "Voice for this profile" section, merged
Enhanced + Voice Output into one Text-to-Speech card (Advanced expander), dead
controls behind a "Coming soon" expander, SectionCards extracted, and the
relay-config fetch lifted into VoiceSettingsViewModel. Standard reads "Global voice".

Waveform: the output/Speaking waveform now unfolds only on the first real
playback-amplitude frame (VoicePlayer attaches the Visualizer on audio-session-id
to fix a deep-buffer cold-start race) instead of leading audio off the state flip.

Spec: docs/plans/2026-06-18-voice-overhaul.md
Follow-ups in TODO.md: connectionId namespacing wiring; realtime-PCM waveform gating.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 22:50:38 -04:00
Bailey DixonandClaude Opus 4.8 265ebf7df8 Merge: reconcile optimistic message ids to server ids into dev
Reloader follow-up (off the same worktree): loadMessageHistory now reconciles
live client-UUID message ids to their server ids (position+role+content,
consume-once) before the delta-merge, and the merge adopts the server id in place
— so gateway assistant rows and user rows carry tokens/badges/attachments by id,
no drop-and-reinsert. Content-fallback drops to a pure safety net. 5 new
ChatHandlerTest cases; build + lint green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 22:05:07 -04:00
Bailey DixonandClaude Opus 4.8 2b74f4552c docs: route follow-ups to TODO.md; codify in CLAUDE.md + AGENTS.md
DEVLOG records what happened; TODO.md is the single home for follow-ups /
deferred work / known gaps. Adds attachment (B3/A6/C5/thumbnails/D5), voice,
and chat follow-ups to TODO.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 22:00:51 -04:00
Bailey DixonandClaude Opus 4.8 911926cddb docs(plans): voice overhaul + clean-mode/pets roadmap specs
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 21:57:20 -04:00
Bailey Dixon 80c7337563 Merge: attachment experience (in-app previews, sensitive-media blur, richer capture) into dev 2026-06-18 21:56:33 -04:00
Bailey DixonandClaude Opus 4.8 ad6b7468cd refactor(chat): reconcile optimistic message ids to server ids before the delta-merge
The delta-merge + priorById carry are keyed by message id, but live
(optimistic) ids only sometimes match the reloaded server transcript: SSE
assistant rows are swapped to the server id mid-turn (replaceMessageId), but
gateway assistant rows keep a local UUID (the gateway exposes no per-message
server id during the turn) and USER rows of every transport keep a local
UUID. So the id-keyed carry silently missed those rows — a gateway turn's
tokens/badges survived only if a content match happened to cover them, and
user rows were drop-and-reinserted with attachments rescued only by the
content fallback.

Reconcile live ids to server ids inside loadMessageHistory before building
the carry map: match each still-unreconciled, non-clientOnly live row to an
unclaimed server row by (role, marker-stripped content), consume-once in
document order, and adopt the server id (prior.copy now sets id = messageId).
SSE assistant rows already carry a server id and are skipped (no double-swap);
clientOnly orphans have no server row and are never mapped; a row that matches
no slot is left alone (graceful fallback on truncation/compaction/divergence).
The content-keyed outbound-attachment fallback stays as the safety net, but is
now fed only by rows that did NOT reconcile, so a reconciled row and the queue
can't double-supply the same attachment. Net: gateway assistant AND user rows
now carry tokens/badges/attachments BY ID, in place, and every subsequent
reload matches by id.

run.started (SSE/runs) was considered for an earlier user-id swap but omitted:
the gateway (primary transport) exposes no such id, the first-reload
reconciliation already covers SSE/runs user rows, and a new callback through
three SSE methods + GatewayTurnCallbacks + the ViewModel would be redundant
surface.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 21:56:17 -04:00
Bailey DixonandClaude Opus 4.8 aa1b239e64 feat(attachments): in-app previews, sensitive-media blur, and richer capture
Inbound: new AttachmentViewer renders image/video/audio/pdf/text in-app
(Media3 + PdfRenderer) with a shared Share/Save/Open-externally toolbar; tapping
an attachment now previews in-app instead of firing ACTION_VIEW. Off-thread card
thumbnails, inline-image save menus, and configurable sensitive-media blur
(OFF/FLAGGED/ALL_IMAGES) applied in card, inline image, and viewer.

Sensitivity is model-emitted metadata only (no classifier): the relay carries a
`sensitive` bit via register_media -> X-Media-Sensitive header ->
FetchedMedia.sensitive -> Attachment.sensitive; the standard path uses a markdown
spoiler/sentinel convention. Adds D6 content re-sniff via _IMAGE_MAGIC.

Outbound: permissionless Photo Picker + camera capture + clipboard paste behind a
Photos/Files/Camera/Paste menu, unified through ingestAttachmentFromUri.

Design spec: docs/plans/2026-06-18-attachment-experience.md
Deferred: download progress/cancel (B3), multi-image gallery (A6), agent-side
sensitivity config gate (C5).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 21:50:55 -04:00
Bailey DixonandClaude Opus 4.8 b76565f6f6 Merge: chat history reloader hardening into dev
Brings in the reloader-gaps worktree (off dev): preserve user-sent attachments
across reload, replace the id-prefix orphan whitelist with a clientOnly flag, and
delta-merge the history reload instead of wholesale-replacing the transcript.
14 new ChatHandlerTest cases; build + lint green. User-message-id reconciliation
(deeper run.started fix) follows as a separate change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 21:35:27 -04:00
Bailey DixonandClaude Opus 4.8 50e638f282 refactor(chat): delta-merge the history reload instead of wholesale replace
loadMessageHistory rebuilt every ChatMessage from server data on each
post-turn reload and reassigned the whole list, carrying client-only state
forward only through a hand-picked field list — the root of the
drop-on-reload class and needless row churn.

Make the per-row reconcile a delta-merge keyed by id: a server message that
matches a local row now copies that row and refreshes only the
server-authoritative fields (content, tool calls, cards, reasoning, role,
timestamp), so EVERY client-only field survives automatically instead of a
curated subset — and an unchanged row produces an equal object, so Compose
doesn't re-render it. A server message with no local row is inserted; a
client-only orphan is kept; a row that was server-backed but is no longer in
the transcript is dropped (genuine server-side delete/fork/truncate). Server
reasoning stays authoritative when present, but live-streamed thinking is no
longer blanked when the transcript omits it. Ordering, media-marker
re-dispatch, card extraction, and the MAX_MESSAGES cap are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 21:16:53 -04:00
Bailey DixonandClaude Opus 4.8 70e94a1aa8 refactor(chat): mark client-only bubbles with a flag instead of id-prefix sniffing
Client-only bubbles (no server-side row) survived the post-turn reload
only if their id matched a known prefix (voice-intent-/steer-/ask-/
system-notice-) or they carried an "Error" badge. Any new client-only
bubble type silently dropped, and the badge check could mis-handle a turn
that errored after persisting.

Add ChatMessage.clientOnly (default false) and set it at every creator:
addSystemNotice, appendAskCardMessage, appendLocalVoiceIntentTrace (both
bubbles), appendLocalVoiceIntentResult, the steer echo, and — where
provenance is only known after the fact — markError (gateway terminal
error on a non-persisted turn) and attachRealtimeTurnTrace (a trace is
attached only for provider-only, non-Hermes-backed realtime turns).

loadMessageHistory now preserves any prior message with clientOnly == true
whose id is absent from the reloaded transcript, replacing the id-prefix
whitelist and the Error-badge sniff. A turn that errored after persisting
keeps its Error badge but IS in the transcript, so it reconciles normally;
only clientOnly + absent-from-transcript marks a preservable orphan.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 21:05:19 -04:00
Bailey DixonandClaude Opus 4.8 373939ce95 fix(chat): preserve user-sent attachments across the history reload
loadMessageHistory rebuilt each ChatMessage with no attachments, so a
user-sent image/file (outbound Attachment, state LOADED, relayToken null)
vanished from its bubble after the post-turn reload. Inbound media
(MEDIA: markers) is re-fetched via the marker re-dispatch, but outbound
attachments are neither in server content nor re-dispatched, so they were
dropped.

Carry outbound-only attachments (relayToken == null) forward across the
reload. priorById matches by id, but user-message ids are never reconciled
to the server id (only the assistant placeholder is swapped via
replaceMessageId), so an id-only carry never fires for user bubbles. Add a
content-keyed, consume-once fallback so outbound attachments survive even
when the reloaded user row carries a fresh server id. Inbound
(relayToken != null) attachments are intentionally excluded to avoid
double-adding what the marker re-dispatch re-fetches.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 20:53:22 -04:00
Bailey DixonandClaude Opus 4.8 f76203c227 docs(diagram): add diagrams/README — file roles + keep-in-sync note
Records the three representations of the architecture model (path-architecture.html,
CombineModel.vue, this SVG) that must be updated together, the canonical gating
sources, and how to regenerate the SVG/PNG.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 20:40:50 -04:00
Bailey DixonandClaude Opus 4.8 a918bdb5fe docs(diagram): add "how Hermes-Relay connects" architecture diagram
Hand-authored SVG (+ PNG raster + editable .excalidraw source) showing the
two-axis model at a glance: Vanilla Hermes (Chat/Manage/Voice, no plugin) as the
always-on backbone, the optional Relay plugin fanning out to the app + CLI
(Terminal/Bridge/relay voice/desktop tools), and the sideload gate sitting on
Device Control.

- Embed the SVG at the top of the user-docs Architecture page (served from public/).
- Add the PNG to the README "What it is" section.

Generated with the excalidraw-diagram skill's design methodology; published as a
dependency-free SVG (the skill's CDN-based render pipeline can't egress in this
sandbox, so the .excalidraw is included as the editable source).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 20:36:54 -04:00
Bailey DixonandClaude Opus 4.8 3128d8cf66 fix(chat): preserve client-only message details across the post-turn reload
loadMessageHistory wholesale-replaces the transcript from server data, which
rebuilds content/tool-calls/reasoning but carries NONE of the per-message state
the server does not persist: token usage + cost, provenance badges, tapped-card
confirmations, and the voice/realtime sync traces. Each had to be patched
individually (badges were; tokens were not), so a normal reply lost its
input/output token subtext the moment the turn finished -- the error bubble kept
it only because errored turns skip the reload.

Replace the badge-only carry map with an id-keyed priorById and carry ALL
client-only fields forward for any message id that still matches -- preserve by
default, instead of a per-field whitelist the next new field always forgets.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 20:19:19 -04:00
Bailey DixonandClaude Opus 4.8 9475f8bec4 feat(chat): session-scoped model display + "show system messages" debug toggle
Model display: the chat header subtitle and the agent detail sheet now resolve
the model from the session scope (selectedModelOverride -> gateway session.info
-> profile -> server default), matching the input chip and footer, so a
mid-session switch shows everywhere. The agent-sheet header took the global model
name but the session provider (showed "gpt-5.5 . xAI Grok"); it now takes a
sessionModelName so model+provider come from one scope, and adds a quiet
"Server default: ..." caption only when the session runs a different model than
the host default -- the always-visible global-vs-session split.

Debug toggle: a default-off "Show system messages" switch in Chat Settings
(DataStore-backed, mirrors parseToolAnnotations) drives ChatHandler.showSystemMarkers
to reveal the otherwise-hidden upstream "[System: ...]" steering markers.

Updates CHANGELOG (Unreleased) and DEVLOG.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 20:10:10 -04:00
Bailey DixonandClaude Opus 4.8 bb3d89d5c4 fix(chat): land model switch on the live session + stop swallowing gateway errors
Model switch: selectModel called fire-and-forget prewarm() then setModel(), so
config.set{key:"model"} ran with no live session and upstream applied it as a
GLOBAL write instead of switching the session. Added suspending prewarmAwait()
that selectModel awaits before setModel, so the switch lands session-scoped (the
same _apply_model_switch path the CLI/TUI /model uses) -- or defers to the next
session.create override when there is genuinely no session, never writing global
config.

Errors: dispatchOn (the main-thread turn-callback wrapper) omitted onStatusUpdate,
so the server's terminal-error lifecycle line hit a default no-op -- the turn was
never badged Error and onComplete's post-turn history reload wiped the client-only
error bubble. Wired onStatusUpdate through dispatchOn (also restores live gateway
status lines) and hardened loadMessageHistory to re-inject local Error-badged
messages the server transcript lacks, so no reload path can swallow a failure.

Also hides upstream role:system "[System: ...]" steering markers from the
transcript by default (desktop/TUI parity), behind a ChatHandler.showSystemMarkers
flag.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 20:09:44 -04:00
Bailey DixonandClaude Opus 4.8 4f291e1625 refactor(naming): rename user-facing "Standard" -> "Vanilla Hermes"
"Standard" was overloaded — it read as both an app feature tier and the
unmodified-upstream server state, which was confusing. Rename all
user-visible strings, docs, onboarding copy, and the matching test
assertions to "Vanilla Hermes" so the no-plugin path reads unambiguously.
Code identifiers, enum constants, and the persisted "standard" route value
are unchanged — that is an internal name only.

Also lands this session's architecture work:
- docs/path-architecture.html — connection-path + chat-transport
  resolution flowchart, plus the build-flavor (googlePlay/sideload)
  capability axis.
- user-docs CombineModel "how the pieces combine" three-tier model and
  the release-tracks/index wording that makes the plugin-vs-flavor
  prerequisites explicit.
- Aligns docs/security.md, upstream-surface-matrix.md, and spec.md on the
  device-control 403 codes (device_control_sideload_only / sideload_only).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 20:05:32 -04:00
Bailey Dixon ddb691a3f3 Merge: connection-UX + cold-start perf + profile-swap audit fixes into dev 2026-06-18 16:58:52 -04:00
Bailey DixonandClaude Opus 4.8 800cc0b6ec feat(voice): note that standard voice uses the host's global TTS, not the profile
Standard (no-plugin dashboard) voice rides upstream POST /api/audio/speak, which
is text-only global TTS — TTSSpeakRequest has no profile field and
text_to_speech_tool has no profile scope (web_server.py) — so switching the chat
profile does not change the spoken voice on standard-only installs. The relay
voice path IS profile-aware and is left untouched.

- On a profile change, when the EFFECTIVE voice route is Standard and the
  profile is non-default, record a quiet Voice diagnostics line explaining the
  limitation and pointing to the Relay plugin for profile-aware voice.
- AutoVoiceAudioClient gains effectiveRoute, resolving Auto against live
  readiness (relay-first) so the notice never claims the relay path has this
  limitation.
- StandardHermesVoiceClient passes profile= on /api/audio/speak defensively
  (upstream ignores extra fields today; forward-compatible if upstream adds
  profile-aware TTS).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 16:53:17 -04:00
Bailey DixonandClaude Opus 4.8 9200b25224 feat(chat): agent-sheet toggles say "confirms on your next message" when ready
The YOLO/Fast controls showed an indefinite "Checking…" spinner whenever their
value was null. After a new chat or profile switch the value is intentionally
unconfirmed and only re-settles from session.info on the user's next message —
so an endless spinner reads as broken. When the gateway is Ready (socket up) but
the value is still null, the placeholder now reads "Confirms on your next
message" instead; the spinner is reserved for the genuine still-probing state.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 16:53:04 -04:00
Bailey DixonandClaude Opus 4.8 0800ddeb4b fix(chat): keep yolo/fast/effort/personality per-session across new chats + profile switches
Setting reasoning effort, fast, or YOLO BEFORE a new chat's first message ran a
sessionless gateway config.set, which upstream applies as GLOBAL writes — and
YOLO via os.environ["HERMES_YOLO_MODE"], leaking approval-bypass into every
other session. Profile switches also leaked stale state: a stale personality
overlay was injected onto the new profile's first SSE turn, and reasoning effort
was re-fetched sessionless (reading the launch/global profile's value, not the
newly-selected one).

Verified against upstream tui_gateway/server.py: session.create consumes
model/provider (model_override), reasoning_effort (create_reasoning_override),
and fast (priority service tier) as PER-SESSION overrides, but does NOT accept
yolo.

- GatewaySessionModel now carries nullable reasoningEffort + fast (model also
  nullable) and binds them on session.create with upstream's param names; null
  fields leave the profile/server default intact.
- selectReasoningEffort/setFast/setYolo skip the sessionless config.set on a
  brand-new chat (no live session); effort/fast ride session.create, YOLO is
  stashed and applied session-scoped from the turn's onSessionId.
- _selectedReasoningEffort is now nullable (null = unknown) so a profile/
  connection switch shows the chip as unconfirmed until session.info, never a
  stale value that could ride session.create.
- Profile/connection switches reset personality to default + effort to unknown
  (alongside yolo/fast) so neither a stale overlay nor chip carries over; the
  optimistic getReasoningSettings() fetch in activateGatewayProfile is dropped.
- GatewayChatClientTest gains reasoning_effort/fast session.create binding cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 16:52:53 -04:00
Bailey DixonandClaude Opus 4.8 2fed7bc479 fix(android): profile drawer — whole-pill badges + expandable descriptions
Badges (Active / "N skills" / SOUL / model) no longer split internally
(maxLines=1, softWrap=false, Clip) — so no vertical "S O U L" or "141\nskills"
under width pressure — and wrap as WHOLE pills in their own FlowRow on a
dedicated line below the description. Long profile descriptions truncate to 2
lines with a gated "More"/"Show less" affordance (only shown on real overflow),
so a long description can't crunch the badges. One new optional ProfileRadioRow
param (secondaryExpandable, default false); only the profile-list caller opts in.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 16:16:40 -04:00
Bailey DixonandClaude Opus 4.8 bb0c9f76eb feat(android): cold-start perf + clearer, honest connection UX
Cold-start keystore contention (~2.9s -> ~0.95s to Paired, 3 keyset builds -> 1):
- SecureStoreCache (sync ConcurrentHashMap.computeIfAbsent) builds each prefs
  file's Tink keyset once process-wide; buildRawTokenStore shared factory.
- Defer the throwaway legacy-sentinel AuthManager's keyset build (eagerHydrate);
  re-gate the pre-StrongBox migration on file name + a marker (read legacy once).
- Unify the dashboard cookie store onto the connection's token keyset
  (tokenStoreKey provider) with a one-shot, marker-gated cookie migration.

Honest loading, never stale, never hidden:
- LoadedFadeIn / RelaySkeletonLine; fade-ins on header subtitle, agent sheet,
  context meter, session drawer, Manage.
- Standard upstream controls (Model, YOLO, Fast, reasoning effort) never hidden:
  live when ready, "checking..." while loading, disabled-with-reason when the
  transport can't use them (GatewayToggleControl); bounded picker loading rows.

Connection clarity:
- Session-path summary in the agent sheet (friendly transport + route + honest
  capability chips), absorbing the old "Show routes" expander.
- Redesigned the Connections detail screen (removed API/Voice/Relay redundancy,
  lighter hierarchy).
- Injected-context "media capability" is transport-aware (no false "not set" on
  the gateway, where the relay renders server-local images client-side).

UI polish:
- Connection toast -> live stepper + finger-tracking dismiss + error link.
- Chat header: approvals -> amber icon, Share -> overflow, endpoint chip dropped
  (footer strip now tappable -> Connections), no "none" personality.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 15:37:11 -04:00
Bailey Dixon e27f5e8b9a Merge pull request #93 from Codename-11/Codename-11/fix-ui-ux-issues
fix(chat): apply model pick on new chats, render relay images, smooth profile switch
2026-06-18 14:16:18 -04:00
Bailey DixonandClaude Opus 4.8 f3aba63977 fix(chat): apply model pick on new chats, render relay images, smooth profile switch
UI/UX fixes from a profile-switching + chat-composer audit, verified against
upstream tui_gateway/server.py.

* Model picker now applies on a fresh chat. The gateway model is a per-session
  override; we set it via config.set on live sessions only, so a brand-new
  chat's config.set carried no session_id (upstream no-ops it) and
  session.create omitted the model -> the agent ran on the account's global
  default. Added a live GatewayChatClient.sessionModelProvider (mirrors
  sessionProfileProvider) that binds model/provider onto session.create, which
  upstream honors as the session's model_override -- matching the desktop
  client. Mid-session switches still use config.set; a profile switch retires
  an explicit pick (the profile owns its model) and seeds the picker label
  up-front so it doesn't lag the round-trip. SSE paths already carried the
  model. GatewayChatClientTest gains 3 model-binding cases.

* Server-local images render through the relay. Markdown ![](/path) images only
  understood http(s) -> a server path fell to an "image is on the server"
  notice that never consulted the relay (only the MEDIA: marker path did).
  Added a RelayServerImageResolver CompositionLocal (provided by ChatScreen from
  ChatViewModel.resolveServerImage) that fetches an absolute path via the relay
  /media/by-path route, decodes, caches (bounded LRU), and renders inline with
  tap-to-zoom. On SSE the agent is also told it can surface images/files by path
  when a relay route is configured (shown in the "What the agent sees" sheet).
  Standard no-plugin connections are unchanged.

* Smoother profile switch. switchProfileContext no longer clears the message
  list before the async history fetch; the previous transcript is held and
  swapped atomically, so the LazyColumn's animateItem() cross-fades old->new
  instead of blanking to an empty/Loading state.

Verified: :app:compileSideloadDebugKotlin + unit-test compile + the
GatewayChatClientTest suite are green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 14:15:25 -04:00
Bailey Dixon acfe55f958 Merge pull request #92 from Codename-11/fix/dashboard-ci-and-agent-parity
ci(dashboard): restore requests dep + close agent-framework parity gaps
2026-06-18 11:13:26 -04:00
Bailey DixonandClaude Opus 4.8 8f45c7a4cd ci(dashboard): also install relay reqs (aiohttp) for plugin.relay import
First pass added `requests` and got the suite collecting (18 tests ran), but
one test imports `plugin.relay.tailscale`, which loads plugin/relay/server.py
-> `import aiohttp`. Restore `-r relay_server/requirements.txt` (aiohttp +
pyyaml) alongside fastapi/httpx/requests so the full import chain resolves.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 11:08:58 -04:00
Bailey DixonandClaude Opus 4.8 b996b379a7 docs(agents): close framework-parity gaps
Make the agent guidance work across frameworks that don't read AGENTS.md
natively, and make AGENTS.md self-sufficient beyond Android.

- AGENTS.md: add the Plugin (Python 3.11 aiohttp) and Desktop CLI (Node >=21,
  zero-dep) stack rules to the non-negotiables (was Android-only).
- GEMINI.md: thin pointer to AGENTS.md for Gemini CLI.
- .github/copilot-instructions.md: thin pointer + quick non-negotiables for
  GitHub Copilot.

Both shims point at AGENTS.md as the single source of truth (which links on to
CLAUDE.md for depth) so rules are single-sourced and can't drift.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 11:02:54 -04:00
Bailey DixonandClaude Opus 4.8 b0db3638f0 ci(dashboard): restore requests dep dropped by streamline
The repo-automation streamline trimmed the dashboard API test deps to
`fastapi httpx`, but `python -m unittest plugin.dashboard.test_plugin_api`
imports the `plugin` package, whose __init__ eagerly loads android_tool and
desktop_tool — both of which `import requests`. Without it the test module
fails to import (ModuleNotFoundError: requests), failing CI on dev.

Restore just `requests` (the only third-party need in that chain beyond the
already-present fastapi/httpx); no need to bring back relay_server/requirements
or pytest.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 11:02:44 -04:00
Bailey Dixon bcfd509d35 Merge branch 'Codename-11/repo-automation' into dev 2026-06-18 10:45:19 -04:00
Bailey Dixon d92d18a607 chore: streamline repo automation 2026-06-18 10:45:11 -04:00
Bailey Dixon 4dfa1ecd18 Merge pull request #91 from Codename-11/fix/terminal-tui-and-chrome
feat(terminal): scrollable compact key bar, TUI-correct input, isolated tmux
2026-06-18 10:18:01 -04:00
Bailey DixonandClaude Opus 4.8 64e5a5f75e feat(terminal): scrollable compact key bar, TUI-correct input, isolated tmux
Refines the terminal screen against Orca's mobile terminal and hardens the
relay PTY for correct TUI behavior.

Android:
- Extra-keys bar: horizontalScroll with fixed-min-width keys (labels no
  longer clip), compacted to ~32dp keys / 12sp to match Orca's sizing.
- Mode-aware special keys: window.termSendKey reads xterm's DECCKM and
  encodes arrows/Home/End as SS3 vs CSI; PASTE routes through term.paste()
  for bracketed paste so multi-line paste no longer auto-runs.
- Compact header: custom ~52dp row replaces the 64dp TopAppBar; status shown
  once inline (dot + word, ellipsized) and tappable for the info sheet. Tab
  strip hidden for single-tab sessions (new-tab "+" moves to the header).
- Removed a redundant navigationBarsPadding gap below the keys; added an 8px
  bottom gap in the terminal so the last row clears the key bar.

Relay:
- Terminal sessions spawn on a dedicated -L hermes-relay tmux socket with a
  generated config: escape-time 0, tmux-256color + truecolor, mouse,
  focus-events, set-clipboard, aggressive-resize, status off. Isolated from
  the user's own tmux; persistence unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 09:57:26 -04:00
Bailey Dixon 839c279da7 Merge pull request #90 from Codename-11/docs/native-encryption-devlog
docs(devlog): backfill native secure routes entry (#88)
2026-06-17 22:53:40 -04:00
Bailey DixonandClaude Opus 4.8 fcc6e7601d docs(devlog): backfill native secure routes entry (PR #88)
PR #88 (feature/native-encryption) landed without a DEVLOG entry; record the split connection model (Features vs Route) + plugin secure-proxy route.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 22:53:11 -04:00
Bailey Dixon bebaeb7418 Merge pull request #89 from Codename-11/docs/native-encryption-changelog
docs(changelog): native secure routes entry (backfill for #88)
2026-06-17 22:50:31 -04:00
Bailey DixonandClaude Opus 4.8 9d23eb32ad docs(changelog): add native secure routes (connections features vs routes) entry
Backfills the [Unreleased] CHANGELOG bullet for PR #88 (feature/native-encryption), which landed without one: connections now split Features from Route, plus a plugin Secure proxy route.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 22:49:41 -04:00
Bailey Dixon 99239c8417 Merge pull request #87 from Codename-11/Codename-11/app-theming-enhancements
feat(theme): theme-aware brand tokens, app themes, and hot-swappable sphere
2026-06-17 22:17:02 -04:00
Bailey Dixon 208a1a6ebc Merge pull request #88 from Codename-11/feature/native-encryption
feat(android): native secure routes — split connection features from routes
2026-06-17 22:12:41 -04:00
Bailey DixonandClaude Opus 4.8 a11e7f9420 fix(theme): qualify LocalContext reference in RelayApp
RelayApp never imports LocalContext (every other use is fully qualified
as androidx.compose.ui.platform.LocalContext); the sphere-skin wiring
used the short form, breaking compilation. Match the file convention.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 22:03:46 -04:00
Bailey DixonandClaude Opus 4.8 1224414a80 feat(theme): theme-aware brand tokens, app themes, and hot-swappable sphere
Fix the formerly hardcoded-dark chat/Manage surfaces and add real app
themes plus a pluggable agent sphere.

Brand tokens: convert the dark-only RelayRefresh object into a
snapshot-backed facade over an active BrandPalette, so the ~150 existing
RelayRefresh.X call sites repaint with the theme without edits. The
Material ColorScheme is now derived from the palette (toColorScheme),
and a new LocalBrand CompositionLocal backs new code. Flourishes and
markdown syntax highlighting across 13 files now follow the active
palette (LocalBrand.current.isDark) rather than the system setting.

App themes: ship 8 looks via an AppThemes registry — Hermes Relay
(light+dark) plus ports of the Nous Hermes dashboard baselines (Teal,
Nous Blue, Midnight, Ember, Mono, Cyberpunk, Rose). Hybrid model: the
brand honors Light/Dark/Auto; character themes are fixed-mode. New
appTheme pref + swatch gallery in Appearance.

Hot-swappable sphere: a SphereSkin layer over the untouched core
algorithm (parity mirror preserved). Built-in Adaptive (follows theme),
Classic, Aurora, Solar, Mono skins plus user-authored JSON skins
(SphereSpec/SphereSkinLoader, data-only + validated). Reactivity
(voice/tools/intensity) is declared per skin, gated in the renderer, and
shown as capability badges. Auto-follow-theme with per-skin override.
Format documented in docs/sphere-spec.md.

Reviewed, not compiled (no SDK in worktree). gradlew lint + on-device
verify pending via Android Studio.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 21:52:24 -04:00
Bailey DixonandClaude Opus 4.8 b7d6a2fb31 fix(docs-site): keep hero sphere canvas backing store synced to its css box
On mobile the hero phone-preview morphing sphere rendered at ~1/3 size and
hugged the top-left of the frame. The canvas backing store (sized once in
resize() from a clientWidth snapshot, with the dpr transform) drifted from
drawSphere()'s live per-frame clientWidth reads, so the grid was drawn into a
coordinate space that no longer matched the store — and canvas drawing starts
at (0,0), hence the top-left pin. Mobile triggered it via late-resolving 88cqw
container-query width (resize() bailed on cw<=0, leaving the 300x150 default
store with no dpr transform that the truthy-width guard never retried) and via
the 88cqw->80cqw boot->chat width tween that never resized screenEl.

Add syncCanvasSize(): measure the real box with getBoundingClientRect(),
reallocate the backing store only on an actual pixel-size change (re-applying
the dpr transform), and return the css-px dims to draw against. drawSphere()
now calls it every frame and draws against that single measurement, so the
store and draw math can no longer diverge and a not-ready layout self-heals on
the next frame. Point the ResizeObserver at the canvas (not screenEl) so the
boot->chat width tween is tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 21:50:43 -04:00
Bailey Dixon f38870a4c5 Merge pull request #86 from Codename-11/feature/chat-ux-transparency
feat(chat): injected-context audit sheet + spoken-turn badges; UX polish
2026-06-17 21:47:52 -04:00
Bailey Dixon dbfde1ffc4 Merge remote-tracking branch 'origin/dev' into feature/chat-ux-transparency
# Conflicts:
#	DEVLOG.md
2026-06-17 21:20:04 -04:00
Bailey Dixon 2cea7d1618 merge: native encryption route model 2026-06-17 21:18:58 -04:00
Bailey Dixon 54337826bd feat(android): split connection features from routes 2026-06-17 21:18:29 -04:00
Bailey Dixon 7daa301075 feat(android): merge permissions review screen 2026-06-17 21:13:25 -04:00
Bailey Dixon b5bdf0a81f feat(android): add permissions review screen 2026-06-17 20:54:57 -04:00
Bailey DixonandClaude Opus 4.8 fa18e1c88e docs: changelog + devlog for chat transparency batch
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 20:53:03 -04:00
Bailey DixonandClaude Opus 4.8 c71751cd06 feat(chat): spoken-turn badges + injected-context audit sheet
Voice-mode replies get a Voice chip and realtime replies keep Realtime Agent; both share a speaker glyph (MessagePathBadge gained an optional leading icon). composeInjectedContext() single-sources the per-turn system_message build for both startStream (sent) and previewInjectedContext() (shown); tapping the ContextMeterBar opens InjectedContextSheet, with the gateway persona labeled server-side. loadMessageHistory now preserves provenance badges by id across the post-turn reload, also fixing the pre-existing Stopped/Error loss.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 20:53:03 -04:00
Bailey DixonandClaude Opus 4.8 578f67a352 fix(ui): clearer version-skew error + opaque connection toast
RelayErrorClassifier maps a 400 whose body names an unsupported field to a non-retryable Relay-update-needed message, distinct from a bad value such as an unsupported codec. ConnectionStatusToast composites its container over the theme surface so the floating overlay is opaque; the in-flow banner stays translucent by design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 20:53:03 -04:00
Bailey DixonandClaude Opus 4.8 9739b88af3 fix(docs-site): keep hero sphere canvas backing store synced to its css box
On mobile the hero phone-preview morphing sphere rendered at ~1/3 size and
hugged the top-left of the frame. The canvas backing store (sized once in
resize() from a clientWidth snapshot, with the dpr transform) drifted from
drawSphere()'s live per-frame clientWidth reads, so the grid was drawn into a
coordinate space that no longer matched the store — and canvas drawing starts
at (0,0), hence the top-left pin. Mobile triggered it via late-resolving 88cqw
container-query width (resize() bailed on cw<=0, leaving the 300x150 default
store with no dpr transform that the truthy-width guard never retried) and via
the 88cqw->80cqw boot->chat width tween that never resized screenEl.

Add syncCanvasSize(): measure the real box with getBoundingClientRect(),
reallocate the backing store only on an actual pixel-size change (re-applying
the dpr transform), and return the css-px dims to draw against. drawSphere()
now calls it every frame and draws against that single measurement, so the
store and draw math can no longer diverge and a not-ready layout self-heals on
the next frame. Point the ResizeObserver at the canvas (not screenEl) so the
boot->chat width tween is tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 20:44:29 -04:00
Bailey Dixon e04e55c35c Merge pull request #84 from Codename-11/Codename-11/connectionviewmodel-decomposition
refactor(viewmodel): decompose ConnectionViewModel into transport/pairing/profile collaborators (ADR 34 follow-up)
2026-06-17 19:23:21 -04:00
Bailey Dixon 2b7b698285 Merge remote-tracking branch 'origin/dev' into Codename-11/connectionviewmodel-decomposition
# Conflicts:
#	DEVLOG.md
2026-06-17 19:22:18 -04:00
Bailey Dixon 7ce8e6270a Merge pull request #85 from Codename-11/docs/changelog-voice-enhancements
docs(changelog): voice-mode enhancements [Unreleased] entry
2026-06-17 19:19:16 -04:00
Bailey DixonandClaude Opus 4.8 c43b6a6014 docs(changelog): add Unreleased entry for voice-mode enhancements
Public-facing Keep-a-Changelog entry (Added/Changed/Fixed) for the voice work
merged in #83: enhanced voice control (Gemini & xAI), render-path visibility,
spoken-output formatting, and the realtime/synthesis fixes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 19:18:31 -04:00
Bailey Dixon 52e4159c4c Merge pull request #83 from Codename-11/Codename-11/voice-mode-enhancements
feat(voice): relay fixes + provider-aware enhanced voice (Gemini + xAI) + diagnostics
2026-06-17 19:01:18 -04:00
Bailey DixonandClaude Opus 4.8 ab646eba27 Merge origin/dev into voice-mode-enhancements
Resolved conflicts from dev's ADR 34 network package fence:
- StandardHermesVoiceClient.kt: took dev's refactored network/upstream version
  (the interface/adapter/AutoVoiceAudioClient now live in network/shared +
  network/relay), then re-applied the standard-voice polish (25MB transcribe
  guard, 413/400 copy, MAX_TRANSCRIBE_BYTES).
- network/relay/RelayVoiceAudioClientAdapter.kt: re-applied the
  enhancedOverridesProvider param (RelayApp's auto-merged call requires it).
- DEVLOG.md: kept dev's entries + prepended the voice-mode-enhancements entry.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 19:00:18 -04:00
Bailey DixonandClaude Opus 4.8 2b6fb2c3c8 docs(devlog): record ConnectionViewModel decomposition (3 collaborators, Relay deferred)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:55:21 -04:00
Bailey DixonandClaude Opus 4.8 a3fb37fc94 docs: voice enhanced-voice surface, route ownership, render-path troubleshooting
- upstream-surface-matrix.md: "Voice Surfaces (standard vs. relay)" with an
  explicit route-ownership table (every /voice/* route is relay-owned; only
  dashboard /api/audio/* is upstream; no upstream streaming/WS audio route) and
  an enhanced-voice matrix across both relay paths.
- spec.md Phase V: /voice/synthesize overrides, tts.enhanced block, and the
  voice_output auto_speech_tags control.
- user-docs/features/voice.md: "Enhanced Voice (Gemini & xAI)" section, the
  streaming speech-tags toggle, the settings Render-path row + Diagnostics
  breadcrumb in troubleshooting, and corrected the stale ~/voice-memos note.
- DEVLOG: session entry covering the fixes, enhanced voice, and docs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:50:08 -04:00
Bailey DixonandClaude Opus 4.8 334a4f0ee4 feat(app): voice spoken-output hint, enhanced-voice UI, render-path visibility
- VoiceViewModel: enrich STABLE_VOICE_INTERFACE_CONTEXT so the model formats
  replies for speech (rides the non-persisted system_message slot, no history
  pollution). ChatViewModel forces voice turns onto SSE since the gateway
  prompt.submit has no system-message slot.
- Enhanced-voice UI: EnhancedVoiceOverrides + EnhancedVoiceCapabilities;
  RelayVoiceClient.synthesize sends the generic override fields; provider-aware
  "Enhanced Voice (<provider>)" Voice Settings card (curated dropdown for
  Gemini, free-text for xAI; persona for Gemini, language for xAI).
- Streaming: VoiceOutputConfig.auto_speech_tags + updateVoiceOutputConfig
  param + an "Expressive speech tags" switch in the Hermes Chat + Voice Output
  card (xai_tts), persisted with the existing Save buttons.
- Render-path visibility: a per-session DiagnosticsLog entry naming the active
  path (streaming /voice/output vs basic /voice/synthesize), plus a persistent
  "Render path" row in the settings card derived from voiceOutputConfig.
- Standard voice polish: pre-flight 25MB transcribe guard + friendly 413/400
  copy; harden the dashboard audio HEAD probe to also try /api/audio/speak.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:49:53 -04:00
Bailey DixonandClaude Opus 4.8 b3562dd6cb feat(relay): voice fixes + provider-aware enhanced voice (Gemini + xAI)
Correctness fixes:
- broker.py: non-native realtime-agent loop dropped playback.drained, tearing
  down sessions every turn on non-native providers. Extract a shared
  _handle_common_client_message dispatcher used by both loops so they can't
  drift; add input_audio.clear to the non-native path. Preserves the native
  loop's per-message provider_task.done() break.
- voice.py: synthesize now owns a temp output_path and deletes it after
  streaming (no more ~/voice-memos leak).
- realtime_voice.py: bind the lab WS session to its creating principal
  (_auth_matches_session), mirroring voice_output.py.

Enhanced voice (per-request, no fork; upstream imports isolated in
upstream_voice.py):
- /voice/synthesize accepts voice/model/audio_tags/persona_prompt/language,
  mapped onto Gemini (_generate_gemini_tts) or xAI (_generate_xai_tts).
- /voice/config advertises a provider-aware tts.enhanced capability block.
- /voice/output streaming renderer honors xAI auto_speech_tags as a per-profile
  voice_output: setting (threaded through config/env/YAML/settings/session/
  provider_options/config_payload/PATCH, mirroring text_normalization); the
  relay applies upstream_voice.apply_xai_speech_tags() per chunk. No Gemini
  streaming provider in voice_lab, so Gemini enhanced voice is synthesize-only.

Tests: non-native playback.drained regression (red-on-bug), Gemini + xAI
synthesize overrides, enhanced-block + extract pure-function coverage,
auto_speech_tags PATCH round-trip, apply_xai_speech_tags call-through/fail-soft.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:49:38 -04:00
Bailey DixonandClaude Opus 4.8 0d969468a8 refactor(viewmodel): extract ProfileController from ConnectionViewModel
Move the agent-profiles cluster into
viewmodel/connection/ProfileController.kt:

- the merged agentProfiles list (relay auth.ok union dashboard
  /api/profiles) + refreshDashboardProfiles + profile-scoped
  session/message fetch
- the per-connection selected-profile state machine
  (selectProfile / resolvePendingProfileFrom / pending-name resolution)
- the three persistence stores (selection / session / displayAlias,
  exposed as public vals so the ViewModel's connection-lifecycle
  orchestrators keep their clear/persist call sites byte-identical)
- profileDisplayAlias + activeSessionTransport + per-profile
  last-session restore

ConnectionViewModel keeps its public getters/functions and delegates.
Because the profile state machine is co-driven by ViewModel-level
lifecycle observers (connection switch, active-connection change,
agent-profile arrival, gateway-availability settle), those observers
stay in the ViewModel and call profileController.* lifecycle hooks in
their original order — the orchestration stays put; only the state +
logic moved, so the state machine is now unit-testable in isolation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:44:09 -04:00
Bailey DixonandClaude Opus 4.8 ceb707581e refactor(viewmodel): extract UpstreamTransportController from ConnectionViewModel
Move the upstream dashboard/gateway transport cluster into
viewmodel/connection/UpstreamTransportController.kt:

- per-connection encrypted DashboardCookieStore cache + accessors
- a single consolidated DashboardApiClient factory (was 4+ build sites)
- the cached GatewayChatClient (lazy build, mid-turn LAN/Tailscale
  retarget) + gateway availability tier + sticky-Unsupported verdict
- the per-endpoint capability snapshot + chatMode, and the
  streamingEndpoint-preference resolution that reads them

ConnectionViewModel keeps its public getters/functions and delegates;
rebuildApiClient pushes the probed capability snapshot via
setCapabilitiesAndMode. The @Synchronized gateway-cache lock moves with
the state (now the controller instance), preserving mutual exclusion.

Deliberately NOT moved: the HermesApiClient SSE/runs client
(_apiClient/_chatApiClient), API-server reachability/health, and
rebuildApiClient/rebuildChatApiClient — those are written inline by
several ViewModel-level orchestrators (saveStandardApiConnection,
saveApiAndProbeVoice, testApiConnection, updateApiServerUrl, revalidate)
interleaved with diagnostics + callbacks; lifting them would need a wide
mutable surface that relocates the coupling rather than removing it (per
the decomposition plan's stop-if-too-entangled rule).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:30:08 -04:00
Bailey DixonandClaude Opus 4.8 5e5f76076f refactor(viewmodel): extract PairingController from ConnectionViewModel
Move the paired-devices list (GET /sessions) + management
(load/revoke/extend/revokeChannelGrant) and the insecure-ack DataStore
flags into viewmodel/connection/PairingController.kt. ConnectionViewModel
keeps its public getters/functions and delegates unchanged — a pure
mechanical lift, behavior preserved verbatim.

First step of the ConnectionViewModel decomposition (ADR 34 follow-up).
The pairing orchestrator (applyPairingPayload) stays in the ViewModel:
it is glue across the upstream/relay/connection-store collaborators, not
a cohesive unit that moves cleanly behind this seam.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:16:48 -04:00
Bailey Dixon e00d439b61 Merge pull request #82 from Codename-11/docs/cvm-decomposition-plan
docs(plans): ConnectionViewModel decomposition plan
2026-06-17 17:52:38 -04:00
Bailey DixonandClaude Opus 4.8 19a7c84c18 docs(plans): ConnectionViewModel decomposition plan (ADR 34 follow-up)
Worktree-runnable plan to break the 5.5k-line ConnectionViewModel god object
into focused viewmodel/connection/ controllers (Upstream/Relay transport,
Pairing, Profiles) behind its frozen public surface, plus an optional
ChatTransportProvider seam. Behavior-preserving extraction only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 17:52:14 -04:00
Bailey Dixon 7739a372a9 Merge pull request #81 from Codename-11/feature/upstream-relay-isolation
refactor(network): fence vanilla-upstream from Relay surfaces (ADR 34)
2026-06-17 17:42:16 -04:00
Bailey DixonandClaude Opus 4.8 a014ce8707 docs(devlog): record upstream/relay isolation work (ADR 34)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 17:34:34 -04:00
Bailey DixonandClaude Opus 4.8 afc9b3fd1a test(ci): vanilla-upstream route-surface contract (ADR 34)
Prove the Android standard-path route surface exists on unmodified
NousResearch/hermes-agent — the invariant CLAUDE.md asserts but that was never
tested (the staging server runs a fork with relay routes compiled in).

- scripts/check-upstream-route-contract.py source-parses upstream's declared
  routes (aiohttp add_* + FastAPI decorators): no server boot, no pip install,
  no model keys. Two tiers: REQUIRED standard-path routes fail the build if
  missing; mode-dependent routes (auth-gate, /api/pty, /v1/models) only warn.
  Refuses to pass against our fork via a fork-marker guard.
- .github/workflows/ci-contract.yml checks out vanilla upstream with NO relay
  bootstrap, asserts the checkout is vanilla, runs the contract. Weekly
  schedule tracks upstream main as a drift siren; PR/push use a pinned ref.

Verified locally against the upstream clone: 12/12 REQUIRED routes present;
auth-gate routes correctly advisory (absent in the loopback-token build).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 17:27:13 -04:00
Bailey DixonandClaude Opus 4.8 0448fd36df test(network): enforce upstream/relay/shared fence with Konsist
Add a JUnit-level architecture test (Konsist) asserting the ADR 34 package
fence on production code: network.upstream must not import network.relay and
vice-versa, and network.shared imports neither. Turns the "standard path =
vanilla upstream" invariant from a review convention into a failing test.

- Add com.lemonappdev:konsist 0.17.3 as a testImplementation dependency.
- ArchitectureBoundaryTest uses scopeFromProduction() so test-only cross-refs
  can't false-fail the boundary.
- Wire it into the ci-android.yml explicit --tests list (the broad aggregate
  hangs per issue #32, so the boundary test must be named or it never runs).

Verified: :app:testSideloadDebugUnitTest --tests "*ArchitectureBoundaryTest"
BUILD SUCCESSFUL — Konsist resolves cleanly on Kotlin 2.3.21.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 17:27:13 -04:00
Bailey DixonandClaude Opus 4.8 ea33bc9944 refactor(network): fence network/ into upstream/relay/shared packages
Physically separate vanilla-upstream network surfaces from Relay additions so
changes from either side have a contained blast radius (ADR 34). No behavior
change — pure package move plus import repointing.

- Split app/.../network/ into network/{upstream,relay,shared} (main + mirrored
  test sources). Upstream: Hermes/Gateway/Dashboard clients, chat payloads,
  ChatHandler, session models. Relay: ConnectionManager, ChannelMultiplexer,
  RelayHttp/Voice clients, BridgeCommandHandler, Envelope. Shared: connectivity/
  endpoint/LAN/profile-URL utilities + the voice routing seam.
- Split VoiceAudioClient.kt three ways: the VoiceAudioClient interface +
  AutoVoiceAudioClient router -> shared; StandardHermesVoiceClient -> upstream;
  RelayVoiceAudioClientAdapter -> relay. Co-locating them would force one file
  to import both worlds.
- Extract LocalDispatchResult to shared. The move surfaced the one real hidden
  upstream->relay coupling: ChatHandler (chat) renders phone-action bubbles from
  the bridge's LocalDispatchResult DTO via a same-package reference. As a passive
  DTO it belongs in shared; both sides now depend only on shared to speak it.
- ChatHandler placed in upstream (not shared): per ADR 3 chat never flows through
  the relay multiplexer; the handler is fed only by upstream transports.
- Update AndroidManifest GatewayKeepAliveService FQCN and the ci-android.yml
  RelayUrlDeriverTest path (it moved to network.relay).

Verified: :app:compileSideloadDebugKotlin and
:app:compileSideloadDebugUnitTestKotlin both BUILD SUCCESSFUL.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 17:20:12 -04:00
Bailey DixonandClaude Opus 4.8 d0b33140ce docs(decisions): ADR 34 — structural fence for upstream/relay isolation
Record the decision to physically separate vanilla-upstream network surfaces
from Relay additions via three net-additive changes: a package fence
(network/{upstream,relay,shared}), a Konsist import-rule JUnit test, and a
vanilla-upstream route-contract CI job. Documents the placement calls decided
by reading (ChatHandler -> upstream per ADR 3; VoiceAudioClient.kt split three
ways) and the rejected alternatives (ConnectionViewModel transport-strategy
split deferred as too risky; custom ktlint/detekt rule deferred in favor of a
Konsist test that reuses existing JVM test infra).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 16:45:26 -04:00
Bailey Dixon 0a34e73ab5 Merge pull request #80 from Codename-11/Codename-11/docs-site-mobile-hero-fix
fix(docs-site): keep hero sphere canvas backing store synced to its css box
2026-06-17 16:45:03 -04:00
Bailey DixonandClaude Opus 4.8 92e507206f fix(chat): scroll bounce/false-FAB on bubble growth + honest resume context
Bubble-growth scroll bugs (Telegram-style tail-follow without reverse layout):
- A bare isAtBottom flip from a growing streaming bubble was misread as "user
  scrolled away" — it popped the scroll-to-bottom FAB and aborted auto-follow
  though the user never touched the screen. Now userScrolledAway is driven only
  by a genuine scroll GESTURE (isScrollInProgress falling edge); content growth
  never sets that, so it can't false-trigger. Reaching the bottom re-arms follow.
- Tail-follow is now an atomic single scrollToItem(bottom) per growth instead of
  the multi-frame settle loop, which collectLatest cancelled mid-settle on the
  next token (~every frame) and stranded the viewport — the visible bounce.

Resume context: on a COLD resume the server's per-session token counters +
compressor are reset, so session.info reports context_used=0 until the first
turn rebuilds the prompt. Painting that would show a misleading 0% on a session
with real history, so only adopt a non-zero figure (warm resume / post-turn);
cold resumes fill on the first exchange. (Server has no pre-turn context to give.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 16:43:29 -04:00
Bailey DixonandClaude Opus 4.8 c53b7cabb9 fix(docs-site): keep hero sphere canvas backing store synced to its css box
On mobile the hero phone-preview morphing sphere rendered at ~1/3 size and
hugged the top-left of the frame. The canvas backing store (sized once in
resize() from a clientWidth snapshot, with the dpr transform) drifted from
drawSphere()'s live per-frame clientWidth reads, so the grid was drawn into a
coordinate space that no longer matched the store — and canvas drawing starts
at (0,0), hence the top-left pin. Mobile triggered it via late-resolving 88cqw
container-query width (resize() bailed on cw<=0, leaving the 300x150 default
store with no dpr transform that the truthy-width guard never retried) and via
the 88cqw->80cqw boot->chat width tween that never resized screenEl.

Add syncCanvasSize(): measure the real box with getBoundingClientRect(),
reallocate the backing store only on an actual pixel-size change (re-applying
the dpr transform), and return the css-px dims to draw against. drawSphere()
now calls it every frame and draws against that single measurement, so the
store and draw math can no longer diverge and a not-ready layout self-heals on
the next frame. Point the ResizeObserver at the canvas (not screenEl) so the
boot->chat width tween is tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 16:33:31 -04:00
Bailey DixonandClaude Opus 4.8 4fad1c5de5 fix(chat): stop scroll-to-bottom FAB flicker; subtle approvals-off marker
- Scroll-to-bottom FAB no longer blinks during streaming. It now hides while
  we're actively auto-pinning — a programmatic scroll in flight, or
  streaming-and-following (smoothAutoScroll on, not scrolled away) — since a
  content burst can momentarily make the list scrollable-forward for a frame
  before the re-pin. The FAB appears only once the user actually scrolls up.
- Subtle approval-bypass marker: when the server reports approvals effectively
  off (YOLO toggle, --yolo, or global approvals.mode=off — all folded into the
  session.info `yolo` boolean), the chat header subtitle carries a quiet amber
  "⚡ approvals off" so the risk is visible without opening the agent drawer.
  The loud toggle + warning stay in the agent drawer (desktop parity).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 16:30:26 -04:00
Bailey DixonandClaude Opus 4.8 75a9fdbbce feat(chat): on-resume context, ephemeral model-switch, opt-in recents
- Context bar on resume: session.info carries upstream's usage block
  (context_used/context_max), emitted on session resume. Parse it into a new
  serverContext flow and paint the context bar immediately instead of waiting
  for the first turn's usage event.
- Model switch is now ephemeral-only: drop the injected "Model switched to X"
  system bubble. The pill updating is the confirmation; server warnings/errors
  surface via a new transientNotice → snackbar channel (never a chat bubble,
  never dropped).
- Recent-prompt chips are now a config option, OFF by default
  (chatRecentPromptsEnabled in ConnectionViewModel + a toggle in Chat settings);
  the composer row is gated on it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 16:22:43 -04:00
Bailey DixonandClaude Opus 4.8 977566c70c feat(chat): live session.info sync (reasoning/credential/yolo/fast) + stale-state refreshes
Augment the gateway surface to match the official desktop without showing stale
state. Audit found we dropped most session.info fields and fetched several server
lists once; contract verified against upstream, parallel review confirmed the new
config.set calls match exactly.

- session.info interceptor now also surfaces reasoning_effort, credential_warning,
  yolo, fast → serverReasoningEffort/serverCredentialWarning/serverYolo/serverFast
  flows; startGatewayStateSync gains one guarded collector each. A /reasoning change
  on desktop/TUI reflects live (not just on turn-complete).
- credential_warning surfaced once per distinct warning as a system notice (dedup'd
  against the constant session.info echoes, cleared when the key is fixed) — turns
  with a missing provider key no longer fail silently.
- YOLO + Fast toggles in the agent sheet: config.set yolo (value 1/0, scope session)
  + config.set fast (value fast/normal), optimistic set+rollback, live state from
  session.info, reset across every session/profile/connection switch. YOLO renders
  loud (error caption + "Approvals are OFF" banner) and stays session-ephemeral.
- refreshSkills()/refreshModels() on agent-sheet open so server-side skill/model
  changes appear without an app reload.
- review fixes: activateGatewayProfile nulls yolo/fast (missing 5th clear site);
  setYolo/setFast rollback re-checks client identity after prewarm and only rolls
  back if it still owns the optimistic value.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:47:56 -04:00
Bailey DixonandClaude Opus 4.8 29793481a7 feat(chat): tool-card collapse persistence, recent-prompt recall, queue mgmt
Phase 2 (native-Chat desktop-TUI parity) — all enhancements to the existing
Compose chat, no TUI/xterm surface:

- 2.1 Tool-call cards keep their expand/collapse across scroll-off and
  re-render (rememberSaveable keyed per tool call, namespaced by the message
  item key). The chevron/rail tree affordance already existed.
- 2.3 Recent-prompt recall: a soft keyboard has no up-arrow, so the composer
  surfaces recent prompts as tappable chips while empty (recentPrompts flow,
  bounded 15, slash-commands excluded). Tap prefills for tweak-and-resend;
  hides on typing / when a queue or fresh chat shows.
- 2.5 Queue management: the queue was count-only. Each queued message is now a
  row — tap to edit (pull back into composer), ✕ to drop one (removeQueuedAt /
  takeQueuedForEdit). Reorder omitted.

2.2 (context bar) landed earlier; 2.4 (session picker) was already adequate.
Plan doc updated — both phases complete; only 1.7 (inline rename) deferred.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:34:40 -04:00
Bailey DixonandClaude Opus 4.8 02595210dd feat(terminal): unread-output dots + jump-to-latest pill
Two more mobile-ergonomics wins, completing Phase 1 of the terminal/chat
parity plan:

- Unread dots: background tabs keep rendering output (stacked WebViews) with
  no signal. TabState.unreadOutput is set when terminal.output lands on a
  non-active tab and cleared on selectTab; a small dot shows on the inactive
  tab chip.
- Jump-to-latest: xterm onScroll reports atBottom via a new onScrollPosition
  bridge method into TabState.scrolledUp; a tappable pill appears over the
  terminal while scrolled up and snaps back to the live tail.

Plan doc updated: 1.1 (history) reclassified — the toolbar up-arrow already
sends ESC[A so shell-native history works; 1.6 (render parity) reclassified —
font cascade + resize contract already present, WebGL addon not vendored.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:19:29 -04:00
Bailey DixonandClaude Opus 4.8 bb9ca52945 docs(play): listing copy + markdown tweaks
Wording ("server" -> "instance"), bullet spacing, and HTML-entity/link
escaping for the Play Console description. (WIP from the parallel session.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:11:03 -04:00
Bailey DixonandClaude Opus 4.8 e2d0a7b214 docs: terminal + chat-parity tracking plan
Companion checklist to the e2e UX audit, scoped from the 3-way comparison
(Android terminal vs upstream desktop TUI vs web dashboard). Phase 1 terminal
ergonomics, Phase 2 native-Chat parity (explicitly enhancement, not a TUI
replacement).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:11:02 -04:00
Bailey DixonandClaude Opus 4.8 c73567a74c feat(terminal): copy-selection and keyboard-toggle keys
Two mobile-ergonomics gaps in the remote shell. COPY reads the xterm
selection via a new window.getSelectionText() hook and commits non-empty
text to the system clipboard (WebView long-press copy is unreliable; pairs
with the existing PASTE). The new keyboard key toggles the soft keyboard via
WindowInsetsControllerCompat on the active tab, focusing xterm on show, since
tapping the terminal doesn't reliably raise the IME on phones.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:10:48 -04:00
Bailey DixonandClaude Opus 4.8 29693bba3d feat(chat): per-session context-usage bar + faster cold-open transport
Context bar: expose absolute per-session token counts (ContextWindowUsage +
contextWindow flow) from the gateway usage events, reset at all four
per-session points. Rewrite ContextMeterBar from an invisible <50% hairline
into a clean desktop-style gauge: filled bar + `NN% · used/max` readout,
color-graded green/amber/orange/red, shown whenever the server reports a
context window. Drop the redundant header "NN% ctx" suffix.

Cold-open: the effort chip (and transport-gated UI) could lag ~30s because
the dashboard probe that flips gatewayAvailability to Ready was only retried
on the 30s health tick. Add a bounded fast-probe on chat foreground while the
verdict is still Unknown, collapsing it to ~1-3s.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:09:31 -04:00
Bailey DixonandClaude Opus 4.8 6b15ae511f fix(ui): throttle idle animations to ~30fps; drop dead ARR workaround
The ambient orb (MorphingSphere) and the always-on ConnectionStatusBadge
heartbeat drove the whole Compose window at the panel refresh (120Hz)
forever — even idle — which on Android 15 makes the platform log
setRequestedFrameRate every frame and wastes battery. Replace the
infinite transitions with a shared frame-throttled driver:

- New rememberAmbientPhase() runs ~30fps and parks when not running.
- MorphingSphere advances on a manual withFrameNanos loop: full-rate while
  active (thinking/streaming/voice), ~30fps idle. dt-accumulation keeps the
  motion speed identical.
- ConnectionStatusBadge + the two pulse banners use rememberAmbientPhase.
- Remove ComposeArrWorkaround (+ its 4 call sites): it reflected a field
  `isArrEnabled` that became a hardcoded SDK>=35 method in Compose 1.11.2,
  so it had been a silent no-op. The NaN log is a platform log of every
  ARR vote and is not suppressible from app code; only redraw frequency is.

Measured idle: ~114fps -> ~43fps (~62% fewer draws/logs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 15:09:11 -04:00
Bailey DixonandClaude Opus 4.8 cdb6eb4b4f fix(chat): server-owned personality on the gateway + picker-command handling
/personality is a picker command upstream (model/skin/personality) — the
desktop/TUI never raw-forward it; a named/none value is applied via config.set,
persisted to display.personality + the live session, and echoed on session.info.
The app forwarded it to slash.exec (dead-end on mobile), had no `none` concept,
and never consumed session.info — so it kept injecting a stale per-turn persona
prompt that fought the server.

- preserve `system-notice-` bubbles across the post-turn reconcile so slash
  results (incl. the disappearing /personality bubble) no longer vanish
- GatewayChatClient: serverPersonality/serverModel/serverProvider flows,
  getPersonality()/setPersonality() (config.get/set), session.info interceptor
- ChatViewModel: selectPersonality() pushes config.set on the gateway and syncs
  _selectedPersonality + the model pill from session.info; bare /personality and
  /model intercepted as picker commands; refreshPersonalities() on sheet open so
  server-supplied changes need no app reload
- startStream: gateway sends no persona/profile prompt (server owns SOUL +
  overlay) — fixes profile-SOUL double-injection; SSE keeps client injection
- ConnectionInfoSheet: drop the synthetic "Default" row; show None + the
  server-provided personalities (server default tagged); AgentDisplay treats
  none/neutral as cleared aliases

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 14:38:29 -04:00
Bailey Dixon fc5b4d0522 Merge feature/ux-audit-wave-1 into dev
e2e UX audit + chat fixes: model-alias guard, gateway error surfacing,
searchable provider-aware model picker, agent-drawer provider line, Stop
polish, and the disappearing-reply (errored-turn reconcile) fix.
2026-06-16 22:55:22 -04:00
Bailey DixonandClaude Opus 4.8 f62bcc4fe3 fix(chat): don't reconcile (and wipe) the error bubble on a failed turn
The post-turn server reconcile (loadMessageHistory in onCompleteCb, added in the
1.1.0 session-UX pass) ran unconditionally for gateway/sessions turns. A turn that
ends in an error has NO assistant message persisted server-side, so the reconcile
replaced [user, assistant-error] with the server's [user] — the assistant error
bubble vanished while the user message stayed (the "disappearing reply" regression;
1.0.0 didn't reconcile gateway turns, hence was unaffected).

Skip the message reconcile when the turn carries the "Error" badge (gateway ❌
lifecycle), keeping the local error visible; still refresh the drawer + drain queue.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 22:48:44 -04:00
Bailey DixonandClaude Opus 4.8 a8247637aa feat(chat): show provider next to model in the agent drawer header
Detail views show "model · provider" (e.g. "gpt-5.5 · Codex"); the chat composer
pill stays model-only by design. Provider resolved from the live gateway current
provider via the model.options provider list.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 22:36:48 -04:00
Bailey DixonandClaude Opus 4.8 e9fba1f6f6 feat(chat): provider-aware model picker — searchable sheet + availability routing
- model.options now parses authenticated / unavailable_models / free_tier /
  total_models (the picker hints upstream already sends via build_models_payload).
- the picker is current-provider-first and DISABLES models the account can't use
  (free-tier / no-credits → "Not on your plan") and flags unauthenticated
  providers ("Needs setup") — matching the desktop picker, so a switch can't land
  on a model that 400s / credits-fails (e.g. nous gpt-5.5 with no balance).
- the model pill now opens a full searchable ModelPickerSheet (CommandPalette
  style: search + provider group headers + selected check) instead of the cramped
  inline dropdown. The composer pill itself stays model-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 22:20:01 -04:00
Bailey DixonandClaude Opus 4.8 42bcc01510 fix(chat): guard hermes-agent model alias, surface gateway errors, stronger Stop
- never send a generic agent alias ("hermes-agent" / "hermes_agent" / "hermes
  agent") as a model on any send-path (in-chat override, gateway setModel /
  reset-to-default, SSE modelOverride). The server 400s on it and falls back to
  a paid model the account can't afford — the real cause of "no replies."
  Resolving the alias to null sends no model, so the server uses its true
  configured default. Adds AgentDisplay.requestModelName().
- surface gateway status.update lifecycle (model fallback, retries, errors) as
  a live status line above the composer, and stamp an "Error" badge on a turn
  that ends in a ❌ error so a failure no longer reads as a normal answer.
- Stop: firm LongPress haptic + a persistent "Stopped" badge on the cancelled
  turn (was a near-imperceptible TextHandleMove + transient toast only).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 22:00:53 -04:00
Bailey DixonandClaude Opus 4.8 6cb18ad364 fix(release): wire Play Console "What's new" into the release flow
gradle-play-publisher reads the Play "What's new" from
app/src/<flavor>/play/release-notes/<locale>/<track>.txt, which never existed —
so the v1.1.0 Production draft uploaded with EMPTY release notes
(RELEASE_NOTES.md only feeds the GitHub Release body, not Play).

- Add app/src/googlePlay/play/release-notes/en-US/default.txt (Play "What's new",
  <=500 chars; seeded with the 1.1.0 text).
- bump-android-version.sh "Next steps" now reminds to update it + adds it to the
  git-add line.
- RELEASE.md section 2 documents it (separate from RELEASE_NOTES.md).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 21:32:07 -04:00
Bailey Dixon 79bc7eaf15 docs: add GitHub issue templates 2026-06-16 21:18:14 -04:00
Bailey DixonandClaude Opus 4.8 f2778f50c3 docs: add e2e UX audit and UX fix-tracking plan
High-level end-to-end UX / daily-use audit (first install -> pair -> standard vs
relay -> all surfaces -> Hermes management), benchmarked against the Hermex client
and the Hermes desktop dashboard, plus a phased fix checklist tracked as work lands.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 21:03:00 -04:00
Bailey DixonandClaude Opus 4.8 121da28767 fix(ux): apply audit waves 1-2 — onboarding, gates, safety, recovery & feedback
Wave 1 (quick wins):
- onboarding: replace "Standard/Advanced" tier cards with capability copy
- power-feature gate: name the server-side Relay-plugin prerequisite
- destructive-verb confirm: make Deny the dominant button, Allow low-emphasis amber
- bridge safety summary: reframe counts as protections, not capabilities
- notification companion: lead with Status + grant action
- chat: send suggestion chips on tap; disable the unimplemented Auto-TTS toggle

Wave 2 (recovery & feedback):
- add RelayUiState.Expired so a revoked/restarted relay session shows
  "Pairing expired — tap to pair again" instead of looping a doomed reconnect
  (wired through asBadgeState/statusText and the Settings relay pill)
- method-aware pairing-verify timeout copy
- camera-permission denial falls through to manual pairing
- "Stopped" acknowledgment on cancel; "Still working…" after a slow first token
- terminal PASTE key (clipboard -> PTY)

Verified: builds (assembleSideloadDebug) and installs to device.
See docs/audits/2026-06-16-e2e-ux-audit.md and 2026-06-16-ux-fix-plan.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 21:03:00 -04:00
324 changed files with 30757 additions and 4991 deletions
+90
View File
@@ -0,0 +1,90 @@
name: Bug report
description: Report a reproducible problem in Hermes-Relay.
title: "[Bug]: "
labels: ["bug"]
body:
- type: markdown
attributes:
value: |
Before submitting, remove secrets, access tokens, real hostnames/IPs, private deployment names, and personal names. Public example IPs such as `192.168.1.100` are fine.
- type: dropdown
id: area
attributes:
label: Affected area
description: Pick the closest surface.
options:
- Android app
- Standard Hermes chat or voice
- Relay plugin or server
- Desktop CLI or tray
- Dashboard plugin
- Docs or installer
- CI, release, or packaging
- Unsure
validations:
required: true
- type: textarea
id: summary
attributes:
label: What happened?
description: State the behavior you saw and what you expected instead.
placeholder: |
Observed:
Expected:
validations:
required: true
- type: textarea
id: steps
attributes:
label: Reproduction steps
description: Include the smallest sequence that reproduces the issue.
placeholder: |
1. Pair or configure...
2. Open...
3. Tap or run...
4. See...
validations:
required: true
- type: textarea
id: environment
attributes:
label: Environment
description: Include only the fields that apply.
value: |
- Hermes-Relay version/tag:
- Install surface: Google Play / sideload APK / local build / plugin / desktop CLI
- Android device and OS:
- hermes-agent version or commit:
- Connection mode: LAN / Tailscale / public TLS / other
validations:
required: true
- type: textarea
id: logs
attributes:
label: Sanitized logs, screenshots, or traces
description: Paste the smallest useful log excerpt. Remove tokens, private URLs, hostnames, IPs, and user-identifying data.
render: shell
- type: textarea
id: upstream
attributes:
label: Upstream or standard-path notes
description: If relevant, note whether this reproduces against unmodified upstream hermes-agent or only with the relay plugin enabled.
- type: checkboxes
id: checklist
attributes:
label: Checklist
options:
- label: I searched existing issues first.
required: true
- label: I removed secrets, tokens, private infrastructure, and personal names.
required: true
- label: I included the affected version or install surface where known.
required: true
+11
View File
@@ -0,0 +1,11 @@
blank_issues_enabled: true
contact_links:
- name: Security guidance
url: https://github.com/Codename-11/hermes-relay/blob/main/docs/security.md
about: Review the security model before posting sensitive vulnerability details publicly.
- name: User documentation
url: https://codename-11.github.io/hermes-relay/
about: Read setup, pairing, remote access, and troubleshooting docs.
- name: Contributing guide
url: https://github.com/Codename-11/hermes-relay/blob/main/CONTRIBUTING.md
about: Review local setup, branch, commit, changelog, and test conventions.
+64
View File
@@ -0,0 +1,64 @@
name: Documentation or setup issue
description: Report unclear, stale, or missing docs and setup guidance.
title: "[Docs]: "
labels: ["documentation"]
body:
- type: markdown
attributes:
value: |
Use this for docs, installer, setup, release-note, or contribution-guide problems. Remove private hostnames/IPs, tokens, and personal names before posting.
- type: dropdown
id: area
attributes:
label: Documentation area
options:
- README
- User docs site
- Android setup
- Relay plugin setup
- Desktop CLI or tray setup
- Release notes or changelog
- Contributor docs
- Other
validations:
required: true
- type: input
id: location
attributes:
label: Page, file, or section
description: Link the page or name the file and heading.
placeholder: user-docs/guide/getting-started.md, README install section, etc.
validations:
required: true
- type: textarea
id: issue
attributes:
label: What is wrong or missing?
description: Explain what was unclear, outdated, misleading, or absent.
validations:
required: true
- type: textarea
id: expected
attributes:
label: Suggested correction
description: Optional. Include the wording, command, screenshot need, or structure that would help.
- type: textarea
id: context
attributes:
label: Context
description: Optional. Include the version, install path, device, or command you were following.
- type: checkboxes
id: checklist
attributes:
label: Checklist
options:
- label: I checked that this is not already covered in current docs.
required: true
- label: I removed secrets, private hostnames/IPs, internal deployment names, and personal names.
required: true
@@ -0,0 +1,78 @@
name: Feature request
description: Propose a product, workflow, or platform improvement.
title: "[Feature]: "
labels: ["enhancement"]
body:
- type: markdown
attributes:
value: |
Keep requests focused on user-visible outcomes. Do not include private infrastructure, secrets, personal names, or branch/workspace plumbing.
- type: dropdown
id: area
attributes:
label: Affected area
options:
- Android app
- Standard Hermes chat or voice
- Relay plugin or server
- Desktop CLI or tray
- Dashboard plugin
- Docs or installer
- CI, release, or packaging
- Unsure
validations:
required: true
- type: textarea
id: problem
attributes:
label: Problem or workflow
description: What is hard, missing, slow, confusing, or unsafe today?
placeholder: Describe the concrete user workflow this would improve.
validations:
required: true
- type: textarea
id: proposal
attributes:
label: Proposed behavior
description: Describe the outcome, not just an implementation detail.
placeholder: After this change, a user should be able to...
validations:
required: true
- type: textarea
id: standard_path
attributes:
label: Standard upstream compatibility
description: If this touches chat, voice, dashboard, API routes, or server behavior, note whether it can work against unmodified upstream hermes-agent.
placeholder: This should work on vanilla upstream because... / This requires the relay plugin because...
- type: textarea
id: alternatives
attributes:
label: Alternatives considered
description: Optional. Mention current workarounds or related approaches.
- type: textarea
id: acceptance
attributes:
label: Acceptance criteria
description: What would make the request complete?
placeholder: |
- Users can...
- The app/server handles...
- Documentation covers...
- type: checkboxes
id: checklist
attributes:
label: Checklist
options:
- label: I searched existing issues first.
required: true
- label: I described the user outcome and affected surface.
required: true
- label: I removed private infrastructure details and personal names.
required: true
+13 -4
View File
@@ -6,11 +6,20 @@
-
## Verification
<!-- List the checks you ran, or explain why a check is not applicable. -->
-
## Checklist
- [ ] `./gradlew assembleDebug` succeeds
- [ ] `./gradlew test` passes
- [ ] Tested on emulator or device (if UI change)
- [ ] Target branch is `dev` unless this is a release PR
- [ ] Android changes: lint and focused unit tests ran, or rationale is listed above
- [ ] Server changes: focused `python -m unittest ...` checks ran, or rationale is listed above
- [ ] Desktop changes: `npm run build` or a narrower documented check ran, or rationale is listed above
- [ ] Docs/site changes: docs build or link check ran, or rationale is listed above
- [ ] UI changes were tested on emulator/device or desktop surface when applicable
- [ ] Commit messages follow [Conventional Commits](https://www.conventionalcommits.org/)
- [ ] CHANGELOG.md updated (if user-facing)
- [ ] No credentials or secrets in committed files
- [ ] Public writing hygiene checked: no secrets, private infrastructure, personal names, or AI/process narration
+22
View File
@@ -0,0 +1,22 @@
# GitHub Copilot instructions — Hermes-Relay
This file exists so GitHub Copilot (which reads `.github/copilot-instructions.md`,
not `AGENTS.md`) picks up the project's agent guidance.
**Read [AGENTS.md](../AGENTS.md) first — it is the single source of truth**
for agent guidance: the entry point, the non-negotiables, and the public-repo
writing hygiene. It links on to `CLAUDE.md` for the deep reference
(architecture, upstream Hermes API, repository layout, per-language code style,
the dev loop, and the Key Files map). Follow those; don't restate them here.
Quick non-negotiables (the full list and rationale are in `AGENTS.md`):
- **Standard path = vanilla upstream only.** The default no-plugin connection
must work against unmodified upstream hermes-agent; server-side needs go
through upstream PRs or the optional relay plugin, never fork patches.
- **Conventional Commits**, `main`/`dev` branching — feature branches off
`dev`, `--no-ff` merges, tags cut from `main`.
- **Android:** Jetpack Compose (no XML), kotlinx.serialization (no Gson),
OkHttp (no Ktor), `wss://` only; run `./gradlew lint` before pushing Kotlin.
- **Public repo:** no personal names, no private infrastructure, no
AI/assistant self-narration in committed prose.
+23 -19
View File
@@ -3,7 +3,9 @@
# Runs on pushes to main/dev and on PRs targeting main/dev, scoped to
# Android-affecting paths so Python-only changes don't spin up the JVM.
#
# Pipeline: lint -> build + test (parallel) -> upload artifacts
# Pipeline: lint, build, and focused tests run concurrently. PRs build debug
# APKs before merge; dev pushes keep lint/tests only to avoid duplicate
# post-merge packaging. Main pushes keep APK artifacts.
name: CI — Android
@@ -38,11 +40,12 @@ concurrency:
jobs:
# ──────────────────────────────────────────────
# Android Lint — gate for build and test jobs
# Android Lint
# ──────────────────────────────────────────────
lint:
name: Lint (Android)
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@v6
@@ -55,25 +58,20 @@ jobs:
- name: Setup Gradle
uses: gradle/actions/setup-gradle@v6
with:
cache-read-only: ${{ github.ref != 'refs/heads/main' && github.ref != 'refs/heads/dev' }}
# Prefer ktlintCheck if configured; fall back to Android lint
- name: Run lint checks
run: |
if ./gradlew tasks --all 2>/dev/null | grep -q "ktlintCheck"; then
echo "Running ktlintCheck..."
./gradlew ktlintCheck
else
echo "ktlintCheck not found, falling back to Android lint..."
./gradlew lint
fi
- name: Run Android lint
run: ./gradlew lint --console=plain
# ──────────────────────────────────────────────
# Android Build — assembleDebug + upload APK
# Android Build — assembleDebug for PRs and main pushes
# ──────────────────────────────────────────────
build:
name: Build (Android)
needs: lint
if: ${{ github.event_name == 'pull_request' || github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 25
steps:
- name: Checkout repository
uses: actions/checkout@v6
@@ -86,12 +84,15 @@ jobs:
- name: Setup Gradle
uses: gradle/actions/setup-gradle@v6
with:
cache-read-only: ${{ github.ref != 'refs/heads/main' && github.ref != 'refs/heads/dev' }}
- name: Build debug APK
run: ./gradlew assembleDebug
run: ./gradlew assembleDebug --console=plain
- name: Upload debug APK
uses: actions/upload-artifact@v7
if: ${{ github.ref == 'refs/heads/main' }}
with:
name: debug-apk
# Product flavors (googlePlay, sideload) nest APKs under
@@ -109,7 +110,6 @@ jobs:
# ──────────────────────────────────────────────
test:
name: Test (Android)
needs: lint
runs-on: ubuntu-latest
timeout-minutes: 20
# Advisory on dev, strict on main. Evaluates to false (= strict) for
@@ -128,6 +128,8 @@ jobs:
- name: Setup Gradle
uses: gradle/actions/setup-gradle@v6
with:
cache-read-only: ${{ github.ref != 'refs/heads/main' && github.ref != 'refs/heads/dev' }}
# The broad Gradle `test` aggregate currently hangs in deferred JVM test
# suites tracked by issue #32. Keep CI release-relevant until that suite is
@@ -136,14 +138,16 @@ jobs:
- name: Run focused Android unit tests
run: |
./gradlew :app:testSideloadDebugUnitTest \
--tests com.hermesandroid.relay.network.RelayUrlDeriverTest \
--tests com.hermesandroid.relay.network.ArchitectureBoundaryTest \
--tests com.hermesandroid.relay.network.relay.RelayUrlDeriverTest \
--tests com.hermesandroid.relay.viewmodel.ConnectionSwitchTest \
--console=plain
# Upload test reports even if tests fail, for debugging
# Upload reports only for failures. Successful PR report uploads add
# noticeable latency and are rarely inspected.
- name: Upload test reports
uses: actions/upload-artifact@v7
if: always()
if: failure()
with:
name: test-reports
path: app/build/reports/tests/
+89
View File
@@ -0,0 +1,89 @@
# Hermes-Relay — Vanilla-Upstream Route Contract (ADR 34)
#
# Proves the Android *standard path* (no-plugin) route surface exists on
# UNMODIFIED NousResearch/hermes-agent — the invariant CLAUDE.md asserts but
# that was never tested. Source-parses upstream's declared routes (no server
# boot, no pip install, no model keys); see scripts/check-upstream-route-contract.py
# for the design + tradeoff (catches renamed/removed routes; not runtime auth).
#
# PR/push runs check a pinned ref (non-flaky); the weekly schedule tracks
# upstream `main` as a drift siren so a route rename surfaces on our clock.
name: CI — Upstream Contract
on:
push:
branches: [main, dev]
paths:
- "scripts/check-upstream-route-contract.py"
- ".github/workflows/ci-contract.yml"
- "app/src/main/kotlin/com/hermesandroid/relay/network/upstream/**"
pull_request:
branches: [main, dev]
paths:
- "scripts/check-upstream-route-contract.py"
- ".github/workflows/ci-contract.yml"
- "app/src/main/kotlin/com/hermesandroid/relay/network/upstream/**"
schedule:
- cron: "0 6 * * 1" # Mondays 06:00 UTC — upstream-drift siren (tracks main)
workflow_dispatch:
inputs:
upstream_ref:
description: "NousResearch/hermes-agent ref to check (branch, tag, or SHA)"
required: false
default: ""
concurrency:
group: ci-contract-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' && github.ref != 'refs/heads/dev' }}
jobs:
route-contract:
name: Vanilla-upstream route contract
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- name: Checkout hermes-relay
uses: actions/checkout@v6
- name: Resolve upstream ref
id: ref
run: |
# PR/push runs use a known-good NousResearch/hermes-agent commit so
# normal CI is stable. The weekly schedule below intentionally tracks
# main as the upstream-drift siren.
DEFAULT_REF="ef4b897a1843cd32c4f141f55db60f0f0602cc98"
if [ "${{ github.event_name }}" = "schedule" ]; then
REF="main" # weekly drift siren
elif [ -n "${{ github.event.inputs.upstream_ref }}" ]; then
REF="${{ github.event.inputs.upstream_ref }}" # manual override
else
REF="$DEFAULT_REF"
fi
echo "ref=$REF" >> "$GITHUB_OUTPUT"
echo "Checking standard-path route contract against upstream ref: $REF"
- name: Checkout vanilla upstream (no plugin, no bootstrap)
uses: actions/checkout@v6
with:
repository: NousResearch/hermes-agent
ref: ${{ steps.ref.outputs.ref }}
path: _upstream
fetch-depth: 1
- name: Set up Python 3.11
uses: actions/setup-python@v6
with:
python-version: "3.11"
- name: Assert upstream checkout is vanilla (no relay bootstrap/plugin)
run: |
if [ -e "_upstream/hermes_relay_bootstrap" ] || \
[ -e "_upstream/plugin/hermes_relay_bootstrap" ] || \
find _upstream -name "hermes_relay_bootstrap.pth" 2>/dev/null | grep -q .; then
echo "FAIL: upstream checkout contains a relay bootstrap — not vanilla."; exit 1
fi
echo "OK: upstream checkout carries no relay plugin/bootstrap."
- name: Run route-surface contract
run: python scripts/check-upstream-route-contract.py "_upstream"
+5 -5
View File
@@ -5,15 +5,11 @@ on:
branches: [main, dev]
paths:
- "plugin/dashboard/**"
- "scripts/check-plugin-version-sync.py"
- "scripts/check-server-version-sync.py"
- ".github/workflows/ci-dashboard.yml"
pull_request:
branches: [main, dev]
paths:
- "plugin/dashboard/**"
- "scripts/check-plugin-version-sync.py"
- "scripts/check-server-version-sync.py"
- ".github/workflows/ci-dashboard.yml"
permissions:
@@ -55,7 +51,11 @@ jobs:
run: python scripts/check-plugin-version-sync.py
- name: Install dashboard API test deps
run: pip install -r relay_server/requirements.txt fastapi httpx pytest requests
# The suite imports the `plugin` package transitively: __init__ loads
# android_tool/desktop_tool (`import requests`), and one test imports
# `plugin.relay`, whose server.py needs `aiohttp` (+ pyyaml) from
# relay_server/requirements.txt. fastapi+httpx cover plugin_api itself.
run: pip install -r relay_server/requirements.txt fastapi httpx requests
- name: Run dashboard API tests
run: python -m unittest plugin.dashboard.test_plugin_api
+4 -8
View File
@@ -14,6 +14,10 @@ on:
permissions:
contents: read
concurrency:
group: ci-desktop-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' && github.ref != 'refs/heads/dev' }}
jobs:
typecheck-and-build:
name: Type-check + build
@@ -47,16 +51,8 @@ jobs:
# prebuilt dist/ that references a source file that moved.
run: node bin/hermes-relay.js --version
- name: Upload dist/
uses: actions/upload-artifact@v4
with:
name: desktop-dist
path: desktop/dist
retention-days: 7
smoke-help:
name: Smoke — --help + --version work on every target OS
needs: typecheck-and-build
strategy:
fail-fast: false
matrix:
+1 -11
View File
@@ -4,7 +4,7 @@
# plugin-affecting paths so Android-only changes don't spin up the
# Python toolchain.
#
# Pipeline: syntax-check -> focused plugin tests
# Pipeline: syntax-check and focused plugin tests run concurrently.
name: CI — Plugin
@@ -17,9 +17,6 @@ on:
- "plugin/cli.py"
- "plugin/pair.py"
- "plugin/plugin.yaml"
- "plugin/dashboard/manifest.json"
- "plugin/dashboard/package.json"
- "plugin/dashboard/package-lock.json"
- "plugin/relay/**"
- "plugin/tools/**"
- "plugin/tests/**"
@@ -39,9 +36,6 @@ on:
- "plugin/cli.py"
- "plugin/pair.py"
- "plugin/plugin.yaml"
- "plugin/dashboard/manifest.json"
- "plugin/dashboard/package.json"
- "plugin/dashboard/package-lock.json"
- "plugin/relay/**"
- "plugin/tools/**"
- "plugin/tests/**"
@@ -76,9 +70,6 @@ jobs:
with:
python-version: "3.11"
- name: Install dependencies
run: pip install -r relay_server/requirements.txt
- name: Syntax check (plugin relay — canonical location)
run: |
python -m py_compile plugin/relay/server.py
@@ -103,7 +94,6 @@ jobs:
# ──────────────────────────────────────────────
unit-tests:
name: Focused Plugin tests (Python)
needs: syntax-check
runs-on: ubuntu-latest
timeout-minutes: 10
# Advisory on dev, strict on main. Evaluates to false (= strict) for
+1 -1
View File
@@ -43,7 +43,7 @@ jobs:
cache-dependency-path: user-docs/package-lock.json
- name: Install dependencies
run: npm install
run: npm ci
working-directory: user-docs
- name: Build VitePress site
+104
View File
@@ -0,0 +1,104 @@
name: Play Store Listing
on:
pull_request:
paths:
- "assets/screenshots/**"
- "assets/play-store-icon-512.png"
- "assets/play-store-feature-1024x500.png"
- "docs/media/screenshots.json"
- "app/src/googlePlay/play/default-language.txt"
- "app/src/googlePlay/play/listings/**"
- "scripts/screenshots.py"
- ".github/workflows/play-listing.yml"
push:
branches:
- main
- dev
paths:
- "assets/screenshots/**"
- "assets/play-store-icon-512.png"
- "assets/play-store-feature-1024x500.png"
- "docs/media/screenshots.json"
- "app/src/googlePlay/play/default-language.txt"
- "app/src/googlePlay/play/listings/**"
- "scripts/screenshots.py"
- ".github/workflows/play-listing.yml"
workflow_dispatch:
inputs:
publish_listing:
description: "Publish Play Store listing metadata after validation"
required: true
default: false
type: boolean
permissions:
contents: read
jobs:
validate:
name: Validate Listing Assets
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v6
- name: Set up Python
uses: actions/setup-python@v6
with:
python-version: "3.12"
- name: Install image tooling
run: python -m pip install --upgrade rich Pillow
- name: Validate screenshots and listing metadata
run: python scripts/screenshots.py validate
publish-listing:
name: Publish Listing Metadata
needs: validate
# Auto-publish the listing when its assets change on `main` (the release
# branch; the path filters above already scope this to screenshot/graphic/
# text changes). `dev` pushes and PRs validate only. A manual dispatch with
# `publish_listing` still works as an on-demand republish.
if: >-
${{ (github.event_name == 'workflow_dispatch' && inputs.publish_listing)
|| (github.event_name == 'push' && github.ref == 'refs/heads/main') }}
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v6
- name: Set up JDK 17
uses: actions/setup-java@v5
with:
distribution: temurin
java-version: 17
- name: Setup Gradle
uses: gradle/actions/setup-gradle@v6
with:
cache-read-only: false
- name: Write Play service account
id: sa
env:
PLAY_SERVICE_ACCOUNT_JSON: ${{ secrets.PLAY_SERVICE_ACCOUNT_JSON }}
run: |
if [ -z "$PLAY_SERVICE_ACCOUNT_JSON" ]; then
# Skip gracefully (no red CI) when the secret isn't configured — e.g.
# an auto-publish push to main before the service account is set up.
echo "::notice::PLAY_SERVICE_ACCOUNT_JSON not configured — skipping listing publish."
echo "configured=false" >> "$GITHUB_OUTPUT"
else
printf '%s' "$PLAY_SERVICE_ACCOUNT_JSON" > play-service-account.json
echo "configured=true" >> "$GITHUB_OUTPUT"
fi
- name: Publish Play Store listing
if: ${{ steps.sa.outputs.configured == 'true' }}
run: ./gradlew publishGooglePlayReleaseListing
- name: Remove Play service account
if: always()
run: rm -f play-service-account.json
+4 -3
View File
@@ -60,9 +60,8 @@ jobs:
- name: Setup Gradle
uses: gradle/actions/setup-gradle@v6
- name: Build debug APK
run: ./gradlew assembleDebug
with:
cache-read-only: false
# Keep the tag release gate aligned with CI — Android's broad Gradle
# `test` aggregate currently hangs in deferred JVM suites tracked by
@@ -90,6 +89,8 @@ jobs:
- name: Setup Gradle
uses: gradle/actions/setup-gradle@v6
with:
cache-read-only: false
- name: Decode release keystore
env:
+4
View File
@@ -31,6 +31,10 @@ local.properties
/app/release/
*.apk
*.aab
# Scratch / working directory (local pet packs, generated test assets, etc.)
/tmp/
/build-*.log
*.jks
*.keystore
/captures
+7 -2
View File
@@ -13,11 +13,12 @@ then `docs/spec.md` and `docs/decisions.md`.
- Release process → **[RELEASE.md](RELEASE.md)**
- Contributor setup → **[CONTRIBUTING.md](CONTRIBUTING.md)**
- `android_*` toolset + MCP → **[docs/mcp-tooling.md](docs/mcp-tooling.md)**
- Follow-ups / deferred work / known gaps → **[TODO.md](TODO.md)** (the single home for "what's next" — never DEVLOG, never scattered code comments)
## Non-negotiables (the short list)
- **Standard path = vanilla upstream only.** The default (no-plugin) connection —
chat via the API server, standard voice via the Hermes dashboard — must work
- **Vanilla Hermes path = upstream-only.** The default (no-plugin) connection —
chat via the API server, Vanilla Hermes voice via the Hermes dashboard — must work
against unmodified upstream hermes-agent. Server-side needs go through upstream
PRs or the optional relay plugin, never fork patches.
- **Verify endpoints against upstream** (`gateway/platforms/api_server.py` /
@@ -26,6 +27,10 @@ then `docs/spec.md` and `docs/decisions.md`.
`--no-ff` merges, version bumps at release-prep on `dev`, tags cut from `main`.
- **Android:** Jetpack Compose only (no XML), kotlinx.serialization (no Gson),
OkHttp (no Ktor), `wss://` only. Run `./gradlew lint` before pushing Kotlin.
- **Plugin (Python 3.11+):** aiohttp + asyncio (no threading), type hints
everywhere, structured `logging` (no `print`). **Desktop CLI (Node ≥21):**
zero runtime deps, strict TS + ES modules, ship compiled `dist/`. Full
per-language style and the dev loop live in CLAUDE.md → "Code Style".
## Public-repo writing hygiene
+59
View File
@@ -6,6 +6,65 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this
## [Unreleased]
## [1.2.0] - 2026-06-20
### Added
- **Sensitive-media classification (relay).** The relay teaches the agent — server-side, via a removable system-prompt block — to mark private/NSFW media so the phone blurs it per your setting. **On by default for relay installs** (installing the relay is itself the opt-in); reversible from the "Agent context" toggle in the Relay dashboard, or `RELAY_AGENT_CONTEXT_ENABLED=0`. The exact injected instruction is visible in the chat "What the agent sees" sheet under "Relay context (server-side)". No on-device or relay-side classifier — sensitivity stays model-emitted. Vanilla upstream (no plugin) is unaffected. See `docs/plans/2026-06-20-relay-enhancement-layer.md`.
- **Transport path is visible (chat).** The chat status strip now shows which streaming path is actually in use — ⚡ Gateway (live thinking), 📡 Sessions, Completions, or Runs — instead of a generic "api online", and Chat Settings adds a basic→best tier ladder explaining the active path and its fallback.
- **Injected-context audit (chat).** Tap the context-usage meter in chat to open a "What the agent sees" sheet showing the exact extra context prepended to your next turn — persona/profile, phone status, and any per-turn (voice) hint. On the gateway path it notes the persona is applied server-side, so the audit is honest about what the phone does and doesn't send.
- **Spoken-turn badges (chat).** Voice-mode replies now carry a "Voice" chip and realtime replies a "Realtime Agent" chip — both with a speaker glyph — so spoken turns are distinguishable from typed ones in the scrollback.
- **App themes.** A new theme picker in Settings → Appearance ships eight looks: the signature Hermes Relay brand (with full light/dark) plus ports of the Nous Hermes baselines — Hermes Teal, Nous Blue (light), Midnight, Ember, Mono, Cyberpunk, and Rosé. The whole app — brand chrome, accents, and chat background — follows the chosen theme. Light/Dark/Auto applies to themes that ship both modes; fixed-mode themes show their own complete look.
- **Hot-swappable agent sphere.** The orb is now a pluggable "skin": an Adaptive skin that recolors to match your theme, built-in Classic / Aurora / Solar / Mono looks, and support for **user-authored skins** loaded from a small JSON spec. Each skin declares which live signals it reacts to (voice, tool bursts, activity), shown as capability badges in the picker. See `docs/sphere-spec.md`.
- **Connections separate features from routes (Android).** Connection settings now distinguish what a connection can *do* (a **Features** section) from how this phone *reaches* Hermes (a **Route** section), so you can enable Relay features over whichever transport you prefer. A plugin-provided **Secure proxy** route is surfaced alongside LAN, Tailscale, public, and custom routes. The standard direct-to-upstream path is unchanged and still needs no plugin. See `docs/plans/2026-06-18-native-secure-routes.md`.
- **Enhanced voice control (Gemini & xAI).** When the relay uses a Gemini or xAI voice provider, Voice Settings can now steer it: pick a Gemini voice and model and turn on expressive tone tags (with optional natural-language voice direction), or set an xAI voice with expressive speech tags. Expressive tags also apply to xAI on the streaming voice-output renderer. Standard (no-plugin) voice stays configured server-side.
- **Voice render-path visibility.** Voice Settings shows which path is rendering speech (streaming vs. basic), and Diagnostics records it each session, making voice issues easier to troubleshoot.
- **Agent pets — a living, swappable avatar.** The orb can be replaced with an animated "pet" that reacts to what the agent is doing: idle / thinking / writing / speaking / listening states, a distinct **working** pose during tool calls, one-shot **greet** / **celebrate** reactions, and a loop that quickens as output streams. Add or remove pets right in Settings → Appearance (no `adb` needed), with a live state preview, a playback-speed slider, and optional frame auto-stabilization; capability badges (Voice · Tools · Activity) show honestly what each pet actually reacts to. Pets are pure data — an AI authoring kit and a JSON schema let you generate one from sprite art. See `docs/pet-spec.md` and the custom-avatars guide.
- **Per-profile agent icon + single-image avatars.** Each agent profile can wear its own small icon beside its name (client-side, never sent to Hermes), shown in chat, the agent sheet, the top bar, and Settings. Importing an avatar now also accepts a single image (auto-wrapped as a one-frame pet) — no animated pack required.
- **In-app crash reporting.** If the app ever force-closes, the next launch shows a clean dialog with the stack trace — **Copy** it, or **Report** to open a pre-filled GitHub issue from the bug template. The report persists until you acknowledge it, and the handler re-raises so the OS still records the crash in Play vitals.
- **Clean text-flow mode (chat).** A distraction-free chat layout where your sent text slides up into a continuous flow, paired with the swappable-avatar/pet system.
- **Permissions review screen.** A central page makes the permission model explicit — standard Chat and Manage need no phone-control permissions, while voice, camera, notifications, and sideload Device Control stay opt-in — reading the same live grants Bridge does.
- **In-app attachment previews + richer capture.** Attachments preview inline before sending, sensitive media is blurred per your setting, and the capture flow is richer.
### Changed
- **Much faster cold start.** The app was building several hardware-keystore-encrypted stores at launch, which serialize on a process-global lock and stalled the chat header (model, personality, approvals) for seconds. It now builds a single keyset and the dashboard cookies share it, cutting measured time-to-connected from ~2.9 s to ~1 s after first frame, with the keystore lock contention gone. Existing sign-ins are migrated automatically on first launch.
- **Honest loading, never stale, never hidden.** Model, personality, and approvals now show a brief "checking…" state and fade in once the server confirms them, instead of popping in or showing a possibly-wrong value. Standard upstream controls (Model, YOLO, Fast, reasoning effort) are no longer hidden while loading or when unavailable — they always appear: a live control when ready, "checking…" while a value loads, or a cleanly disabled control with the reason (e.g. "available over the gateway transport") when this connection can't use them. The chat composer's reasoning-effort chip now shows alongside the model chip instead of lagging seconds behind the gateway check, and picker lists (models, personalities) show a brief, bounded "loading…" cue. The same fade-in is applied to the context meter, session drawer, and Manage panels.
- **Tidier chat header.** The LAN/Tailscale chip was dropped from the top bar (the bottom status strip already shows the route, and is now tappable to open Connections), and a `none` personality is no longer shown — leaving more room for the model name.
- **Connection toast reads like the cold-start screen.** The floating connection status toast now shows a live checklist — Route / API / Relay each with a spinner, ✓, or ✕ as the checks land — instead of flat text, matching the splash screen's stepper. Swiping it up now tracks your finger (slide + fade) rather than snapping, and connection problems get an explicit "Open Connections →" link at the bottom so the path to the detailed view is obvious.
- **Tidier chat header.** The "approvals off" warning moved out of the agent subtitle into a single amber ⚡ icon in the top bar (tap for the full explanation in the agent sheet), and Share folded into a ⋮ overflow menu — so the personality · model subtitle no longer gets clipped by the trailing action icons.
- **Voice replies are formatted for listening.** In voice mode the assistant is now guided to answer in short, conversational sentences without markdown, emoji, or raw URLs — without changing what is stored in chat history.
- **Leaner terminal screen (Android).** The extra-keys bar scrolls horizontally with compact, fully-legible keys (no more clipped "CTRL"), the header is a single compact row showing one inline connection-status dot plus state, and the tab strip is hidden for single-tab sessions — the new-tab "+" moves into the header — reclaiming vertical space for the terminal.
- **Relay terminals run on an isolated, TUI-tuned tmux.** Sessions now use a dedicated tmux server/socket with its own config — instant ESC (`escape-time 0`), truecolor `$TERM`, mouse and focus events on, and no status bar — so editors and full-screen tools behave correctly, without touching the user's personal tmux.
- **"Standard" is now "Vanilla Hermes" throughout.** The user-facing name for the no-plugin upstream path is now **Vanilla Hermes**, so it's clear the default path runs on a plain Hermes agent.
- **QR pairing degrades gracefully on unusual cameras.** On foldables and devices where the camera can't initialize, the scanner now shows a "camera unavailable — pair manually" card instead of force-closing.
- **Image & attachment viewers rotate to landscape.** The full-screen image / attachment viewers can rotate to landscape even though the rest of the app stays portrait-locked.
### Fixed
- **Clearer error when a feature needs a newer relay.** Toggling a setting an older relay plugin doesn't recognize (e.g. xAI expressive speech tags) now shows "Relay update needed" instead of a generic HTTP 400 with a dead Retry button. Genuine input errors are unaffected.
- **Connection status toast is no longer see-through.** The floating connection-lost/switching toast renders fully opaque so content behind it no longer bleeds through and hurts legibility.
- **Provenance badges survive the post-turn history reload.** "Voice", "Realtime Agent", "Stopped", and "Error" chips are now preserved when the conversation reloads after a turn, instead of silently vanishing.
- **Chat and Manage no longer stay dark in Light mode.** Brand-styled surfaces bypassed the theme and were effectively hardcoded dark; they now follow the selected theme and light/dark mode, and the glow/border flourishes key off the active theme rather than the system setting.
- **Realtime voice no longer drops the conversation mid-session with some providers.** A normal end-of-turn signal was being rejected on certain voice providers, ending the session every turn.
- **Relay voice synthesis no longer leaves temporary audio files behind** on the server.
- **Clearer voice errors and an oversize-recording guard.** Standard voice now rejects an over-long recording before uploading it and shows a helpful message for audio the server can't read, instead of a generic HTTP error.
- **Terminal paste no longer auto-runs multi-line text.** The key-bar PASTE now uses bracketed paste, so multi-line content lands intact in shells and editors instead of executing line by line.
- **Terminal on-screen arrows behave inside TUIs.** Arrow/Home/End keys follow the running app's cursor-key mode (application vs. normal), so they work correctly in vim, less, and fzf.
- **Terminal footer spacing.** A small gap now keeps the last terminal row clear of the key bar (it could previously look like the footer overlapped it), and a redundant navigation-bar inset that left empty space below the keys was removed.
- **In-chat model picker now actually applies on a new chat.** Picking a model and provider in the chat composer (e.g. Grok 4.3 via your xAI subscription) is bound to the new conversation, so the agent runs on the picked model instead of silently falling back to the account's global default. Switching profiles retires an explicit pick so the profile's own model takes over, and the picker label updates immediately instead of lagging a round-trip.
- **Server-generated images render in chat when paired to the relay.** An assistant image that points at a server-side file path is now fetched through the relay's media route and shown inline (tap to zoom), instead of degrading to an "image is on the server" notice. On the SSE chat path the agent is also told it can surface images and files by path when a relay route is configured (visible in the chat "What the agent sees" sheet). Standard (no-plugin) connections are unchanged.
- **Smoother profile switching.** Switching profiles no longer blanks the conversation to an empty/"Loading…" state before the new history loads; the previous transcript is held and cross-fades to the new one.
- **In-chat model switch now applies mid-conversation, not just on new chats.** Picking a model in an already-started chat switches the live session in place — the same path the desktop/TUI `/model` uses — instead of racing into a global-default write, so the turn runs the model you picked.
- **Server-side turn errors always surface.** A failed turn (e.g. a provider rejecting the request) now stays on screen as an error bubble with the message, instead of appearing for a moment and then vanishing when the conversation reconciled after the turn.
- **The model shown in chat matches the live session.** The chat header and the agent detail sheet now show the model the current session is actually running (reflecting a mid-session switch) rather than the profile/global default, and the agent sheet no longer pairs the global default model name with the session's provider — it now also names the host's "Server default" when the session runs something different.
- **Server steering markers no longer appear as chat bubbles.** The "[System: the active model/personality changed]" notes the server injects into history for the agent's benefit are hidden from the transcript by default (matching the desktop/TUI); a new "Show system messages" debug toggle in Chat Settings can reveal them.
- **Per-reply token counts (and other per-message details) survive the post-turn reload.** The input/output token subtext, provenance badges, tapped-card state, and voice/realtime sync traces are now preserved when the conversation reconciles against the server after a turn — previously a normal reply lost its token line once the turn finished (the error bubble kept it only because errored turns skip that reload). The reloader now preserves client-only message details by default instead of dropping any it doesn't re-derive from the server.
- **PDF viewer no longer crashes when the document closes mid-render.** A PDF preview that was torn down during a layout pass could read a closed renderer and throw `IllegalStateException: Document already closed`; the renderer is now guarded so it returns nothing instead of crashing.
- **No crash opening a chat with a server-local image.** Rendering a relay-fetched image could throw `ClassCastException: kotlin.Result cannot be cast to byte[]` because a `suspend` function returned `kotlin.Result` (which collides with the coroutine machinery's own wrapper); a purpose-built result type fixes it.
- **Side-loaded avatars and sphere skins are reachable again.** Both loaders read internal storage while the docs (correctly) pointed `adb push` at external app-scoped storage, so a side-loaded pet or skin never appeared. Both now resolve through one external-preferred location, so the documented install path works.
- **Reopened chats paint the session's real model** (not the profile/global default), the model-picker "Server default" caption shows the true default rather than the active override, and a chat's media badge shows only when paired — with the underlying server-image fetch-failure reason surfaced when a fetch fails.
## [1.1.0] - 2026-06-16
### Added
+21 -18
View File
@@ -4,26 +4,26 @@
## What This Is
A native Android app (Kotlin + Jetpack Compose) paired with an optional Python relay plugin/server (aiohttp) for the Hermes agent platform. Standard chat, Manage, and dashboard voice work against unmodified upstream Hermes. Relay adds phone control, terminal, remote desktop tooling, extra voice engines, and dashboard Relay management.
A native Android app (Kotlin + Jetpack Compose) paired with an optional Python relay plugin/server (aiohttp) for the Hermes agent platform. Vanilla Hermes chat, Manage, and dashboard voice work against unmodified upstream Hermes. Relay adds phone control, terminal, remote desktop tooling, extra voice engines, and dashboard Relay management.
**Current state:** v1.0.0 stable. The default no-plugin path supports chat, Manage, and voice on vanilla upstream Hermes. Chat auto-prefers the dashboard `/api/ws` gateway transport when Manage auth is ready, then falls back to API-server SSE routes. Standard voice uses dashboard `/api/audio/*` with the Manage session. Relay remains an additive power path for terminal, bridge/device control, notification companion, extra/provider-native voice, remote access, and desktop tooling. Two Android product flavors ship: `googlePlay` (conservative, no unattended Device Control surface) and `sideload` (full-capability).
**Current state:** v1.0.0 stable. The default no-plugin path supports chat, Manage, and voice on vanilla upstream Hermes. Chat auto-prefers the dashboard `/api/ws` gateway transport when Manage auth is ready, then falls back to API-server SSE routes. Vanilla Hermes voice uses dashboard `/api/audio/*` with the Manage session. Relay remains an additive power path for terminal, bridge/device control, notification companion, extra/provider-native voice, remote access, and desktop tooling. Two Android product flavors ship: `googlePlay` (conservative, no unattended Device Control surface) and `sideload` (full-capability).
## Architecture
```
Phone (WS) -> Hermes dashboard (:9119) [standard gateway chat, live thinking]
Phone (HTTP/SSE) -> Hermes API Server (:8642) [standard chat fallback, sessions, runs]
Phone (HTTP) -> Hermes dashboard (:9119) [standard Manage + voice]
Phone (WS) -> Hermes dashboard (:9119) [vanilla Hermes gateway chat, live thinking]
Phone (HTTP/SSE) -> Hermes API Server (:8642) [vanilla Hermes chat fallback, sessions, runs]
Phone (HTTP) -> Hermes dashboard (:9119) [vanilla Hermes Manage + voice]
Phone (WSS/HTTP) -> Relay plugin/server (:8767) [optional bridge, terminal, relay voice, remote tools]
```
The standard path must stay vanilla upstream only. API-server bearer auth and dashboard cookie auth are separate. Terminal and bridge require Relay pairing; standard chat, Manage, and dashboard voice must not.
The Vanilla Hermes path must stay upstream-only. API-server bearer auth and dashboard cookie auth are separate. Terminal and bridge require Relay pairing; Vanilla Hermes chat, Manage, and dashboard voice must not.
### Upstream Hermes API Reference
**IMPORTANT:** Always verify endpoints against the actual hermes-agent source (`gateway/platforms/api_server.py`). The upstream repo is the source of truth — not our docs, not our memory, not assumptions from other frontends.
**Standard endpoints (confirmed in hermes-agent source):**
**Vanilla Hermes endpoints (confirmed in hermes-agent source):**
| Endpoint | Purpose | Tool Call Format |
|----------|---------|-----------------|
@@ -65,9 +65,9 @@ The Android client probes per-endpoint capability via `HermesApiClient.probeCapa
**Dashboard web server (separate surface — standard Manage / Desktop remote gateway):**
hermes-agent ships a second web server at `hermes_cli/web_server.py` that hosts the React admin dashboard at `hermes_cli/web_dist/`. It has its **own** `/api/*` routes that **do not live on `api_server.py`** — notably: `GET/PUT /api/config` (full tree), `GET /api/config/schema`, `GET /api/config/defaults`, `GET/PUT /api/config/raw` (YAML text), `GET/PUT/DELETE /api/env` + `POST /api/env/reveal`, `PUT /api/skills/toggle`, `/api/cron/jobs/*` (different shape from `/api/jobs/*`), `/api/providers/oauth/*`, `/api/dashboard/themes`, `/api/dashboard/plugins`, `/api/model/info` + `/api/model/options` + `POST /api/model/set`, `/api/profiles/*` (CRUD, `POST /api/profiles/active`, per-profile soul/description/model), `/api/mcp/*`, `/api/logs`, `/api/analytics/usage`, and **`POST /api/audio/transcribe` + `POST /api/audio/speak`** (base64 data-url contract, built for hermes-desktop voice). The API server has **no audio routes** — its `/v1/capabilities` advertises `audio_api: false`; PR #8199 (`/v1/audio/*`) is the canonical future surface but is unmerged. Android's **standard (no-plugin) voice** therefore rides this dashboard surface via `StandardHermesVoiceClient` with the per-connection dashboard cookie session (Manage sign-in unlocks voice); `AutoVoiceAudioClient` prefers Relay when paired and falls back to standard.
hermes-agent ships a second web server at `hermes_cli/web_server.py` that hosts the React admin dashboard at `hermes_cli/web_dist/`. It has its **own** `/api/*` routes that **do not live on `api_server.py`** — notably: `GET/PUT /api/config` (full tree), `GET /api/config/schema`, `GET /api/config/defaults`, `GET/PUT /api/config/raw` (YAML text), `GET/PUT/DELETE /api/env` + `POST /api/env/reveal`, `PUT /api/skills/toggle`, `/api/cron/jobs/*` (different shape from `/api/jobs/*`), `/api/providers/oauth/*`, `/api/dashboard/themes`, `/api/dashboard/plugins`, `/api/model/info` + `/api/model/options` + `POST /api/model/set`, `/api/profiles/*` (CRUD, `POST /api/profiles/active`, per-profile soul/description/model), `/api/mcp/*`, `/api/logs`, `/api/analytics/usage`, and **`POST /api/audio/transcribe` + `POST /api/audio/speak`** (base64 data-url contract, built for hermes-desktop voice). The API server has **no audio routes** — its `/v1/capabilities` advertises `audio_api: false`; PR #8199 (`/v1/audio/*`) is the canonical future surface but is unmerged. Android's **Vanilla Hermes (no-plugin) voice** therefore rides this dashboard surface via `StandardHermesVoiceClient` with the per-connection dashboard cookie session (Manage sign-in unlocks voice); `AutoVoiceAudioClient` prefers Relay when paired and falls back to standard.
Current upstream supports two auth modes on this surface. Loopback dashboards still use the injected `window.__HERMES_SESSION_TOKEN__` path. Remote/non-loopback dashboards use the Desktop-style dashboard auth gate: `/api/status` advertises `auth_required` and providers, `/auth/password-login` handles password providers, `/auth/login?provider=...` handles Nous/OIDC redirects, `/api/auth/me` returns the verified session, and `/api/auth/ws-ticket` mints a short-lived ticket for `/api/ws` / `/api/pty`. This dashboard session is **not** an `API_SERVER_KEY`. Android uses it for Manage, standard voice, and the gateway chat transport. `/api/ws` is backed by `tui_gateway/server.py` (what hermes-desktop + the Ink TUI speak) and is the only upstream surface with **live** `reasoning.delta`/`thinking.delta` streaming; the api_server SSE paths remain the standard fallback. Relay-only capabilities remain behind Relay pairing. **Do not proxy dashboard auth or dashboard admin APIs over the relay.**
Current upstream supports two auth modes on this surface. Loopback dashboards still use the injected `window.__HERMES_SESSION_TOKEN__` path. Remote/non-loopback dashboards use the Desktop-style dashboard auth gate: `/api/status` advertises `auth_required` and providers, `/auth/password-login` handles password providers, `/auth/login?provider=...` handles Nous/OIDC redirects, `/api/auth/me` returns the verified session, and `/api/auth/ws-ticket` mints a short-lived ticket for `/api/ws` / `/api/pty`. This dashboard session is **not** an `API_SERVER_KEY`. Android uses it for Manage, Vanilla Hermes voice, and the gateway chat transport. `/api/ws` is backed by `tui_gateway/server.py` (what hermes-desktop + the Ink TUI speak) and is the only upstream surface with **live** `reasoning.delta`/`thinking.delta` streaming; the api_server SSE paths remain the SSE fallback. Relay-only capabilities remain behind Relay pairing. **Do not proxy dashboard auth or dashboard admin APIs over the relay.**
**Tool call rendering paths:**
1. **Runs API** — Emits `tool.started`/`tool.completed` as real SSE events → `ToolProgressCard` in real-time.
@@ -75,7 +75,7 @@ Current upstream supports two auth modes on this surface. Loopback dashboards st
3. **Annotation parser** — Fallback for servers emitting inline markdown annotations (`` `💻 terminal` ``).
## Key Instructions
- **Standard path = vanilla upstream only.** The default (no-plugin) connection path — gateway/API chat, Manage, and standard voice via the dashboard surface — must work against **unmodified upstream hermes-agent**: no fork patches, no bespoke server config as a dependency. The app ships on Google Play to users whose servers we don't control. Features that need server-side changes go through upstream PRs (with graceful degradation until merged) or live behind the opt-in relay plugin.
- **Vanilla Hermes path = upstream-only.** The default (no-plugin) connection path — gateway/API chat, Manage, and Vanilla Hermes voice via the dashboard surface — must work against **unmodified upstream hermes-agent**: no fork patches, no bespoke server config as a dependency. The app ships on Google Play to users whose servers we don't control. Features that need server-side changes go through upstream PRs (with graceful degradation until merged) or live behind the opt-in relay plugin.
- **Always verify upstream before assuming an endpoint exists.** Check `gateway/platforms/api_server.py` in hermes-agent. If an endpoint isn't there, document whether bootstrap injects it or it requires the fork.
- If we use a non-standard endpoint, ensure `probeCapabilities()` covers it and the auto-resolver degrades gracefully.
- **Bootstrap maintenance:** Retire `plugin/hermes_relay_bootstrap/` per surface. Sessions and read-only skills/toolsets now have native upstream replacements; config, memory, legacy skill detail/toggle, available-models, and slash middleware still need explicit replacement decisions before full removal.
@@ -131,9 +131,10 @@ hermes-android/
## Project Conventions
### File Structure
- **Root-level:** README.md, CLAUDE.md, AGENTS.md, DEVLOG.md, .gitignore
- **Root-level:** README.md, CLAUDE.md, AGENTS.md, DEVLOG.md, TODO.md, .gitignore
- **docs/** — spec, decisions, security, and any other long-form documentation
- **DEVLOG.md** — update at end of each work session with what was done, what's next, blockers
- **DEVLOG.md** — update at end of each work session with what was done + verification (the factual record of *what happened*). It churns; do NOT park forward work here.
- **TODO.md** — the single home for follow-ups / deferred work / known gaps ("what's next"). Record them here — never buried in DEVLOG or scattered through code/doc comments where they get lost.
- **CLAUDE.md hygiene:** Key Files entries must stay one line — implementation detail belongs in the file or `docs/`. Run `/revise-claude-md` after feature-heavy sessions to trim drift.
### Public-repo writing hygiene
@@ -327,6 +328,7 @@ This is a **public, distributed repo** — every committed file (CHANGELOG, DEVL
| `quest/` | [EXPERIMENTAL] Meta Spatial SDK Quest/XR app — gradle `includeBuild("quest")`; needs further development, not shipped |
| **Tooling — dev iteration (not shipped)** | |
| `ui-preview/` | Desktop Compose Hot Reload harness — JVM Compose for Desktop; source-shares `MorphingSphereCore` from `:relay-ui`; `Main.kt` gallery; see `ui-preview/README.md` |
| `app/src/test/.../screenshots/StoreScreenshotTest.kt` | Roborazzi host-side store/docs screenshot renderer — deterministic, no device, exact 1080×2160; reuses real components+chrome with mock data; `capture(name, themeId){…}` renders any view; see `docs/screenshot-automation.md` §Deterministic rendering (JDK-21 + no-plugin gotchas) |
## What NOT to Do
@@ -335,7 +337,8 @@ This is a **public, distributed repo** — every committed file (CHANGELOG, DEVL
- **Don't use Ktor for networking** — OkHttp for WebSocket
- **Don't use plaintext WebSocket** — `wss://` only, even in development
- **Don't put documentation in root** — long-form docs go in `docs/`
- **Don't forget DEVLOG.md** — update it
- **Don't forget DEVLOG.md** — update it (record *what happened*)
- **Don't bury follow-ups** — deferred work / known gaps go in `TODO.md`, never in DEVLOG or one-off code/doc comments
## MCP Tooling
@@ -374,7 +377,7 @@ Curls every bridge HTTP route via `localhost:8767`. Catches the silent-drop regr
1. **Edit locally** — Windows checkout. Both plugin (`plugin/`) and app (`app/`) live here.
2. **Python syntax check** — `python -m py_compile plugin/<file>.py`. Full tests run on the server.
3. **Kotlin changes** — do NOT run `gradle build`. Bailey builds via Android Studio's ▶ button. Never `adb install` from Claude.
4. **Before pushing Kotlin changes** — run `./gradlew lint` locally. It's the exact task CI runs (see `.github/workflows/ci.yml` → `gradlew lint` fallback) and catches errors Android Studio's live inspections miss — e.g. `UnsafeOptInUsageError` with `kotlin.OptIn` vs `androidx.annotation.OptIn`, `FlowOperatorInvokedInComposition` (mapped flows inside Composables), Media3 `@UnstableApi` propagation. Lint is a hard blocker in CI: Build + Test show "skipping" until lint passes, and lint prints only the **first failure** before aborting — so CI iterations reveal errors one at a time while a single local lint run surfaces all of them.
4. **Before pushing Kotlin changes** — run `./gradlew lint` locally. It's the exact task CI runs and catches errors Android Studio's live inspections miss — e.g. `UnsafeOptInUsageError` with `kotlin.OptIn` vs `androidx.annotation.OptIn`, `FlowOperatorInvokedInComposition` (mapped flows inside Composables), Media3 `@UnstableApi` propagation. Android CI runs lint alongside build/test for faster feedback, but a local lint run still surfaces issues before the workflow spends runner time compiling and packaging.
5. **Commit + push** — feature branch off `dev`, merged back to `dev` via PR. `main` is reserved for release merges.
6. **Pull + restart on server** — see Server Deployment below.
7. **Test on phone** — Bailey builds from Studio, installs to Samsung device, pairs via `/hermes-relay-pair`.
@@ -396,7 +399,7 @@ Server is a Linux box running hermes-agent with hermes-relay editable-installed
**Compat hook:** `hermes relay compat status/install/remove` manages only the
optional `hermes_relay_bootstrap.pth` startup hook. New installs load the
plugin-owned bootstrap from `plugin/hermes_relay_bootstrap/`; the repo-root
package is only a legacy import shim. Standard chat, Manage, and dashboard voice
package is only a legacy import shim. Vanilla Hermes chat, Manage, and dashboard voice
must not depend on this hook.
**Key conventions:**
@@ -429,13 +432,13 @@ See [RELEASE.md](RELEASE.md) for the full recipe.
| Surface | Endpoint | Notes |
|---------|----------|-------|
| Chat (gateway) | Dashboard `POST /api/auth/ws-ticket` -> WS `/api/ws` | Standard upstream dashboard/tui_gateway path; live thinking/reasoning; requires dashboard auth |
| Chat (gateway) | Dashboard `POST /api/auth/ws-ticket` -> WS `/api/ws` | Vanilla Hermes dashboard/tui_gateway path; live thinking/reasoning; requires dashboard auth |
| Chat streaming | `POST /v1/runs` → `GET /v1/runs/{id}/events` | Structured tool events; async run-control path |
| Chat (sessions) | `POST /api/sessions/{id}/chat/stream` | Native upstream session-persisted SSE; preferred when capability probe finds it |
| Chat (compat) | `POST /v1/chat/completions` (stream=true) | Inline tool annotations only |
| Session CRUD | `GET/POST/PATCH/DELETE /api/sessions` | Native upstream (#33134); bootstrap fallback only for old builds |
| Manage | Dashboard `/api/status`, `/api/auth/me`, `/api/config`, `/api/profiles/*`, `/api/env`, `/api/model/*`, `/api/mcp/*` | Standard upstream dashboard surface; do not proxy through Relay |
| Standard voice | Dashboard `POST /api/audio/transcribe`, `POST /api/audio/speak` | Standard no-plugin voice; uses dashboard session from Manage |
| Manage | Dashboard `/api/status`, `/api/auth/me`, `/api/config`, `/api/profiles/*`, `/api/env`, `/api/model/*`, `/api/mcp/*` | Vanilla Hermes dashboard surface; do not proxy through Relay |
| Vanilla Hermes voice | Dashboard `POST /api/audio/transcribe`, `POST /api/audio/speak` | Vanilla Hermes no-plugin voice; uses dashboard session from Manage |
| Pairing (QR) | `POST /pairing/register` (loopback only) | Via `/hermes-relay-pair` or `hermes-pair` shim; accepts optional `endpoints` for multi-endpoint QRs |
| Pairing (multi-endpoint) | QR `endpoints` array (ADR 24) | `hermes: 3` schema; ordered `lan`/`tailscale`/`public`/... candidates; phone re-probes on network change |
| Pairing auth | WSS `auth.ok` payload | Includes `expires_at`, `grants`, `transport_hint` |
+295
View File
@@ -1,5 +1,300 @@
# Hermes-Relay — Dev Log
## 2026-06-20 — Release-prep: android-v1.2.0 + plugin-v1.2.0
**Why.** Cut a combined 1.2.0 across both lockstep surfaces (both were at 1.1.0). The accumulated `[Unreleased]` block had captured the major feature arcs but a second wave had landed undocumented — audited every commit since the `*-v1.1.0` tags and backfilled the changelog before promoting it.
- **Versions.** `bump-android-version.sh 1.2.0` (`appVersionName 1.1.0→1.2.0`, `appVersionCode 13→14`); `bump-plugin-version.sh 1.2.0` (pyproject + `plugin/relay/__init__.py` + plugin.yaml + dashboard manifest/package/lock, all in sync). `check-version-tracks.py` + `check-plugin-version-sync.py --expect 1.2.0` green.
- **CHANGELOG backfill.** Promoted `[Unreleased]` → `[1.2.0] - 2026-06-20` with a fresh empty `[Unreleased]`. Added the missing shipped features the accumulator had skipped: **agent pets** (swappable animated avatar + reactivity + in-app add/remove + AI authoring kit), per-profile agent icons + single-image avatars, **in-app crash reporting**, clean text-flow mode, the permissions-review screen, attachment previews; **Changed**: "Standard"→"Vanilla Hermes" rename, QR camera hardening for foldables, viewer landscape rotation; **Fixed**: PDF mid-render crash, the `kotlin.Result`-in-suspend `ClassCastException` on server images, side-loaded avatar/skin storage path, reopened-session model + media-badge fixes.
- **Release notes.** Rewrote `RELEASE_NOTES.md` (Android, "Make it yours" framing) and `PLUGIN_RELEASE_NOTES.md` (enhancement-layer + enhanced voice). Updated in-app `whats_new.txt`, Play `release-notes/en-US/default.txt` (438/500 chars), and the `docs/play-store-listing.md` What's-new block.
- **Scrub.** Grep'd the `[1.2.0]` block for names / private infra / fork plumbing — clean (only pre-existing released blocks carry the LAN host IP from old desktop-alpha entries; out of scope for this cut, flagged separately).
- **Verification.** `python -m unittest plugin.tests.test_enhancements plugin.tests.test_terminal_channel` (20 pass). Android AAB build + `keytool` cert verify is Studio-side (Bailey) per the dev loop. Prep committed on `dev` in two commits (`release(android)` / `release(plugin)`); merge-to-`main` + tags deferred to operator.
## 2026-06-20 — Static-image avatars + per-profile agent icon (Android)
**Why.** Two requests: a custom avatar shouldn't require authoring an animated pack (a single image should work), and each agent profile should be able to wear its own small icon beside its name — client-side, mirroring the existing local-name override.
- **Static-image import (`PetImporter`).** "Add a pet" now accepts a single image (`.png`/`.jpg`/`.gif`/`.webp`), not just a `.zip` — detected by **magic bytes**, not the file name. An image is auto-wrapped as a one-frame static pet (written as `idle.png` with a synthesized minimal `pet.json`), so a static avatar needs no manifest authoring. The renderer already supported a one-frame `idle`; this is purely import ergonomics. `importZip` → `importUri`; `ConnectionViewModel.importPetFromZip` → `importPet`. Test covers the image-wrap path.
- **Per-profile agent icon (client-side).** A direct twin of `ProfileDisplayAliasStore`: new `ProfileIconStore` (own DataStore `profile_icons`, keyed per `(connection, profile)`, **never sent to Hermes**). It stores a **path** to an image copied into `files/profile-icons/` (not a SAF URI, so it survives without persistable permission). Wired through `ProfileController` next to `profileDisplayAlias` (`profileIcon` StateFlow + `setProfileIcon`/`clearProfileIcon` + the copy), exposed on `ConnectionViewModel`, provided at the app root as `LocalAgentIconPath`, and rendered beside the agent name in `MessageBubble` **and** as the **header avatar** in the agent sheet, the chat top bar, and Settings — via a shared `AgentAvatarFace` that shows the icon (Coil from the file path) or falls back to the name's initial. The picker (`AgentIconRow`) sits right under the local-name row in `ConnectionInfoSheet`. Scope: small name-adjacent icon only — the big empty-chat/voice avatar stays global. `ProfileIconStore` cleared alongside the alias on connection removal; test mirrors the alias store's.
- **Verification.** `:app:assembleSideloadDebug` + `PetImporterTest`/`ProfileIconStoreTest` <pending>. Installed via `adb install -r`. On-device check (import an image as a pet; set a profile icon and see it by the name) in TODO.
## 2026-06-20 — Pet state preview in Appearance (Android)
**Why.** Testing a pet meant *inducing* each state by driving the agent (run a tool to see `working`, fail a turn for `error`, start voice for `speaking`/`listening`) — painful. An in-app preview turns Appearance into a pet test harness.
- **Live preview (`AppearanceSettingsScreen`).** Under the speed/stabilize controls (pet selected only): a ~140 dp canvas rendering the active pet, a `FilterChip` row for the seven sustained states (`Idle · Thinking · Working · Writing · Speaking · Listening · Error`), and `Greet`/`Done` buttons that replay the one-shots. Pure UI on the existing `AgentAvatar` seam — no new ViewModel/pref/renderer; it just calls `activeAvatar.Render(AvatarRenderState(state=…))` with a user-picked state, so it also reflects the live speed and stabilize settings.
- **State→render mapping.** Base states feed `state=…`; **Working** feeds `state=Thinking, toolCallBurst=1f` (lights the overlay); **Greet** remounts the preview via a `key` (re-fires the on-appear reaction); **Done** drives a momentary `Speaking → Idle` transition on the live instance (fires the celebrate reaction), then returns to the selected chip.
- **Verification.** `:app:assembleSideloadDebug` BUILD SUCCESSFUL; installed via `adb install -r`. On-device visual check recorded in TODO.
## 2026-06-20 — Pet frame auto-stabilization (Android)
**Why.** On-device audit of a 4×4 AI pet (Lucy) found the character's vertical center drifting 34 px across the 16 cells, with 8/16 frames touching the cell edge — the image model held *appearance* but not *position/scale*, so the pet floated upward and bled the next frame in. The renderer slices/centers exact cells faithfully, so the drift can't be cured per-frame there — but it can be neutralized by re-centering each frame on its own content.
- **Decode-time recenter (`PetAvatar`).** With stabilization on, `decodeClip` scans each frame's opaque pixels (alpha bbox) and stores a per-frame offset that moves the content's bbox center to the cell center; `drawPetFrame` applies it (source px → dest px, scaled). Works for sprite sheets (per-cell) and frame sequences (per-bitmap); empty/transparent frames get a zero offset. The scan is one-time per clip decode on `Dispatchers.IO` with a reused scratch buffer, so steady-state cost is nil.
- **Global toggle, default on (`ConnectionViewModel`, `LocalPetStabilize`, `AppearanceSettingsScreen`).** A `pet_stabilize` pref → `LocalPetStabilize` provided at the app root → read in `PetAvatar.Render` (keys the decode `produceState`, so flipping re-decodes). A "Stabilize frames" Switch sits under the playback-speed slider when a pet is selected. Default on because AI sheets nearly always need it; a hand-authored pet with intentional motion can switch it off.
- **Authoring (docs).** The prompt kit now also stresses *registration* (lock head/shoulders, same position + scale, only secondary motion) so the art improves at the source — stabilization is the safety net for what the model still gets wrong.
- **Verification.** `:app:assembleSideloadDebug` BUILD SUCCESSFUL; installed via `adb install -r`. Fixes the *already-installed* Lucy at render time (no re-import). On-device visual confirmation recorded in TODO.
## 2026-06-20 — Pet playback-speed control + cell-resolution guidance (Android + docs)
**Why.** Two more on-device tuning gaps: a pet that still felt fast needed re-authoring/re-importing to slow down (slow loop), and a 128 px-celled pet looked pixelated blown up to the full-screen chat background (while crisp in the small voice overlay — same frames, different scale).
- **Playback-speed control (`ConnectionViewModel`, `LocalPetPlaybackSpeed`, `PetAvatar`, `AppearanceSettingsScreen`).** A global multiplier pref (`pet_speed`, 0.5×–1.5×, default 1.0) surfaced as a **Slider in Appearance** when a pet is selected. Provided at the app root via a new `LocalPetPlaybackSpeed` composition local and read **live** in `PetAvatar.Render` (`rememberUpdatedState`), so dragging the slider re-times the pet instantly with no restart. Applies to every clip (including one-shots) and composes with intensity (`baseFps × speed × intensityFactor`, clamped 1–60). The sphere ignores it.
- **Cell-resolution guidance (docs).** Pixelation is a *resolution* axis (cell px) distinct from smoothness (frame count): one frame set is contain-fit into every surface, so author for the **largest** (the chat background). Bumped the kit default to **256 px cells** (a 1024×1024 sheet for a 4×4 grid), noted 512 px is fine for a sprite sheet (decodes as one bitmap), and that the old "≲256 px" note applied to frame-*sequences*. Updated `custom-avatars.md`, `pet-prompt-kit.txt`, `pet-spec.md`.
- **Verification.** `:app:assembleSideloadDebug` BUILD SUCCESSFUL; installed to the sideload build via `adb install -r`. Slider behavior is on-device-visual (recorded in TODO).
## 2026-06-20 — Pet kit defaults to 4×4 (16-frame) sheets (docs)
**Why.** A 4-frame (2×2) sheet reads steppy no matter the fps — the on-device Lucy made that obvious. The renderer already slices any N×M grid (`decodeClip` derives `cols`/`rows` from sheet size ÷ cell size; `drawPetFrame` indexes `col = i%cols`, `row = i/cols`), so "support 4×4" is an authoring-default change, not a renderer one.
- **Kit + spec default to a 4×4 grid (16 frames).** The prompt template, the manifest example, and `pet-prompt-kit.txt` now use `frameCount: 16` with fps matched to the higher count (idle ~8 → a calm ~2 s loop); 2×2 / 4 frames stays documented as the easier-to-keep-consistent fallback. `docs/pet-spec.md` states any rectangular grid works (a 4×4 sheet holds 16 frames, decoded as one bitmap regardless of cell count). Added a `PetLoaderTest` case for a 16-frame sheet.
- **Diagnosis note.** The "still fast" report was tracked to the *installed* `pet.json` still carrying `fps 6` (the tuned `fps 3` zip post-dated the import); `intensity` was ruled out by tracing `streamingIntensity` → `0f` at idle. Audited by `adb shell cat`-ing the on-device manifest, not the repo copy.
## 2026-06-20 — Pet frame-loop smoothness fix (Android)
**Why.** First on-device pet (Lucy) animated with a periodic hitch and felt a touch fast. Root cause: `PetAvatar.Render`'s frame loop awaited `withFrameNanos` (one vsync ≈16ms) **and** `delay(1000/fps)` each iteration, so every frame waited ~16ms longer than its `frameDurSec`; the surplus accumulated until the loop forced a 2-frame skip to catch up — a visible hitch, worst at low fps.
- **Vsync-paced loop (`PetAvatar.Render`).** Removed the per-frame `delay`. `withFrameNanos` already suspends until the next frame, so the loop is now purely vsync-paced (~60fps) and advances the sprite only when `frameDurSec` of real time has accumulated — no double-count, no periodic skips. Intensity modulation still recomputes fps each tick; the accumulator absorbs the variable rate without skipping.
- **Authoring guidance (docs).** Clarified that smoothness comes from frame **count**, not fps: 4 frames (2×2) is the consistent-but-steppy minimum, 8–16 (3×3 / 4×4) for fluid motion; match fps to count (calm states 3–4, not 6+). Added to `docs/pet-spec.md`, the user-docs kit, and `pet-prompt-kit.txt`; lowered the example/kit `idle`+`listening` fps to 4.
- **Verification.** `:app:assembleSideloadDebug` BUILD SUCCESSFUL. On-device re-check pending: the test device's wireless adb dropped mid-deploy; the rebuilt APK + a tuned `lucy.zip` (idle/listening fps lowered) are staged to install + re-import once it reconnects.
## 2026-06-20 — In-app pet avatar add/remove/refresh (Android)
**Why.** The Appearance screen could *select* avatars but offered no way to **add or remove** a pet from inside the app — the only path was `adb push` into app-scoped external storage, which scoped storage stalls on (confirmed hanging on a Samsung device: a push into `/sdcard/Android/data/<pkg>/files/pets/` wrote nothing, though `adb shell ls` of the dir worked). And the avatar list loaded once at startup, so even a successfully-pushed pet never appeared without a restart. Net effect: users saw only the Sphere.
- **In-app import (`PetImporter.kt`, new).** "Add a pet" launches a SAF document picker; the chosen `.zip` unpacks into `pets/`. Hardened: zip-slip guard (every entry confined to a staging dir under `cacheDir`), per-file / total-size / entry-count ceilings (zip-bomb), and post-extract validation through the same `PetSpec.toAvatar` the loader uses — an archive that wouldn't render is rejected up front. Accepts either shape (pet.json at the root, or one folder deep — shallowest wins); installs under the manifest `id` (sanitized), replacing a same-named pack.
- **In-app remove (`PetLoader.deletePet`).** Resolves a pack by manifest `id` (not directory name) and deletes it, behind a confirm dialog. If the deleted pet was the selected avatar, the selection falls back to the Sphere.
- **Live refresh (`ConnectionViewModel`, `RelayApp`).** A `avatarsRefreshTick` StateFlow keys the avatar `produceState`, so import/delete — and opening the Appearance screen — re-scan `pets/` and update every surface (chat, clean mode, voice, splash) without an app restart. This also resolves the standing "pet load is process-scoped" TODO. Add/remove results surface as snackbars via a one-shot `avatarEvents` flow.
- **Appearance UI (`AppearanceSettingsScreen`).** Added an "Add a pet" button + "Rescan", an "Installed pets" management list with per-pet remove, and a remove-confirm dialog; replaced the static "drop a pack into pets/ via adb" hint with the in-app flow.
- **Tests.** `PetImporterTest` (root + nested import, no-manifest, missing-idle, **zip-slip refused writes nothing outside staging**); `PetLoaderTest` delete cases (by id, id≠dirname, no-match). Both pass under `:app:testSideloadDebugUnitTest` (the 12 build failures are the pre-existing DataStore/`FileStorage.kt:114` JVM cases — `BargeIn`/`ProfileSelection`/`ProfileSession`Store).
- **Verification.** `:app:assembleSideloadDebug` built `hermes-relay-1.1.0-sideload-debug.apk`; installed to the Samsung sideload build via `adb install -r` (Success). A `lucy` test pet (9 sprite-sheet states, 256×256 RGBA with real transparency, schema-validated) staged at `/sdcard/Download/lucy.zip` for an import smoke test (Add a pet → pick from Downloads). On-device import/delete smoke recorded in TODO.md.
## 2026-06-20 — Pet AI authoring kit + JSON schema (docs)
**Why.** Pets are pure data, so the only real barrier to making one is sourcing the art. Documented an AI-generation workflow and a machine-readable contract so both humans and AI agents can author and validate a pet without hand-drawing.
- **AI prompt kit (`user-docs/features/custom-avatars.md`).** A reference-image-first, character-agnostic prompt template (`{character}`/`{style}`/`{accent}`), a per-state motion table mapping image generation to our state vocabulary, a full 9-state manifest, transparency/consistency caveats, and a "fastest first pass" (one 3×3 sheet → nine stills). Mirrored as a one-click `user-docs/public/pet-prompt-kit.txt`. A vendor-neutral "let an AI agent build the pack" callout names Codex/Claude Code as examples and states the acceptance criteria + image-gen prerequisite.
- **JSON Schema (`user-docs/public/pet.schema.json`).** Draft-07 schema mirroring the loader's structural rules (required `idle`, frames-XOR-sheet via `anyOf`, positive sheet dims, `schemaVersion` ≤ 1); `$schema` wired into the manifest examples and tolerated by the lenient loader (`ignoreUnknownKeys`). `docs/pet-spec.md` gained an "Editor validation" section, honest that file-existence/decodability remain load-time checks. Validated: legal draft-07 + accepts good / rejects no-idle, empty-clip, schemaVersion-2, sheet-missing-dims.
## 2026-06-20 — Relay enhancement layer + agent-context injection
**Why.** Teaching the agent to mark sensitive media (and, more generally, to know things only the relay can teach it) needs a way to inject context into the agent's system prompt — but hermes-agent exposes no plugin context hook (`system_prompt_block()` is memory-provider-only; lifecycle hooks are observers). The one transport-agnostic seam is `AIAgent._build_system_prompt`. Rather than a one-off patch, we built a reusable, removable **enhancement layer** so the relay can apply such patches cleanly and retire them per-surface as upstream catches up — the same pattern as the bootstrap route shims.
- **`plugin/enhancements/` (registry + contract).** Each enhancement declares `name · phase · enabled() · apply() · retirement note`. Config-gated, **default ON for relay installs** (the relay install is the opt-in; `RELAY_AGENT_CONTEXT_ENABLED=0` opts out), no-op on vanilla, removable with the plugin.
- **`context_injection` enhancement (fail-open).** Wraps `AIAgent._build_system_prompt` at plugin-load; appends auditable fenced blocks (`<!-- hermes-relay:<name> -->`). Fail-open at every step — seam absent / block build throws / setattr fails ⇒ returns the base prompt unchanged. With `RELAY_AGENT_CONTEXT_ENABLED` off, the prompt is byte-for-byte unchanged. Works on BOTH gateway and SSE (agent core).
- **First block: media-sensitivity.** Teaches the agent to mark private/NSFW media with the client's spoiler convention (`||![alt](path)||` / sentinel alt) — the bit the client already blurs. No soul/memory touched.
- **`GET /context/injected` audit route + client audit.** The relay exposes exactly what it would inject; the chat "What the agent sees" sheet gained a "Relay context (server-side)" section. Server-side injection is never hidden.
- **Sensitivity re-thread (client).** `ServerImageResult.Success` carries the fetched `sensitive` bit again; `RelayServerImage` blur ORs it with the markdown-parsed flag.
- **Transport-path UI.** New `ChatTransportStatusBadge` + `RelayStatusStrip` surface the ACTUAL chat tier (⚡ Gateway / 📡 Sessions / Completions / Runs / offline) instead of a bare "api online"; Chat Settings gained a basic→best tier ladder + a gateway sign-in callout.
- **Dashboard.** Relay management tab gained Agent-context master + per-block toggles (off by default; labeled experimental/server-side/removable).
- **Verification.** Plugin: `python -m unittest plugin.tests.test_enhancements` (16 pass). Android: `:app:lintSideloadDebug :app:assembleSideloadDebug --no-daemon` green. Built by a 2-worker Orca orchestration (server + client slices), coordinator-integrated. Design: `docs/plans/2026-06-20-relay-enhancement-layer.md`.
- **Follow-ups (TODO).** Confirm the `AIAgent` module on the live host when enabling; structured media channel (`docs/plans/2026-06-20-structured-media-channel.md`); incremental bootstrap migration into the enhancement layer; retire the wrap when upstream ships a context hook.
## 2026-06-20 — Pet intensity modulation (Android)
**Why.** The last continuous-reactivity gap: a pet's clip looped at a fixed rate regardless of how hard the agent was working. `intensity` (the activity ramp already fed to every avatar — ~0.7 while streaming) was plumbed to `PetAvatar.Render` but ignored. Wiring it completes the reactivity story (voice ✓ · tools ✓ · activity ✓) and un-clamps the last reserved badge flag — and unlike the deferred `attention`, the signal needed no host plumbing.
- **Live playback-rate modulation (`PetAvatar.Render`).** Opt-in via `reactive.intensity`. The active base/working loop's fps is scaled by `1 + intensity·PET_INTENSITY_RATE` (0.6 → ~1.4× at typical streaming, 1.6× peak, capped at `PET_MAX_FPS`), so it visibly "works harder" as output streams. Read **live** inside the frame loop via `rememberUpdatedState(state.intensity)` so the speed tracks the agent mid-clip without restarting the long-lived loop (re-keying on a continuously-animated float would thrash). One-shot reactions are excluded (`!playOnce`) so `greet`/`done` play at their authored rate.
- **Badge un-clamp (`PET_RENDERER_CAPABILITIES.intensity` → true).** The loader's existing `reactive.intensity && capability` formula now lets a declared `intensity:true` through, so the pet advertises **Activity** honestly. No loader change needed beyond the flag.
- **Tests.** `PetLoaderTest`: a declared `intensity:true` is now honored (`Voice · Activity`); the prior clamp test was split — `tools` without a `working` clip still stays off the badge.
- **Docs (`docs/pet-spec.md`).** Reactivity table's `intensity` row rewritten from "Reserved" to the speedup behavior; removed from "Forthcoming" (now only `attention` remains there). Reactivity is now Voice · Tools · Activity complete.
- **Verification.** Code + loader tests authored to the established patterns; not run here (Studio-side). On-device check (a writing/working loop quickening while streaming) recorded in TODO.md.
## 2026-06-20 — Pet one-shot reaction layer (Android)
**Why.** The behavior model's event tier: transient "reactions" that play once over the base loop, then return — the touch that turns a status display into a character (cf. the Peon Pet's celebrate-on-finish). Distinct from the sustained per-state loops and the `working` overlay.
- **Pet-local triggers, zero host plumbing (`PetAvatar`).** One-shots are derived from the activity-state transitions the avatar already observes each frame — no new `AvatarRenderState` edge from the host. `PetOneShot.Greet` fires on first composition (the pet appears); `PetOneShot.Done` fires when a *productive* turn ends (a `Streaming`/`Speaking` → `Idle` transition; `Thinking → Idle` and `Error → Idle` don't celebrate). Both are opt-in (only if the pet ships the clip) and require ≥2 frames.
- **Play-once-then-revert (`PetAvatar.Render`).** The frame loop gained a `playOnce` mode: a reaction clip plays 0→end (no modulo wrap), parks on its last frame, clears `activeOneShot`, and recomposition hands back to the base loop. A live reaction overlays everything (including `working`). Suppressed under reduced motion (`paused`). An `ONE_SHOT_MAX_MS` (4s) backstop guarantees a reaction never lingers on decode failure / single frame / pause.
- **Friendly aliases (`PetLoader.toAvatar`).** Resolves `greet`/`wake` → `PetOneShot.Greet` and `done`/`celebrate` → `PetOneShot.Done` from explicit `states` keys only (no fallback); absent reactions just don't play. One-shots are reactions, **not** a reactivity signal, so they don't touch the picker badge.
- **Tests.** `PetLoaderTest`: a pack with `greet`/`done` keys loads cleanly and the badge stays `Voice` (no accidental Tools/Activity coupling). Render-time playback (the actual one-shot animation) is on-device/Compose-test territory — flagged in TODO.
- **Docs (`docs/pet-spec.md`).** New "One-shot reactions" section (Greet/Done table, opt-in, play-once, reduced-motion), an Expressive tier on the authoring ladder, and the Loop-vs-one-shot note updated. `attention`-on-notification stays in "Forthcoming" — it needs a host event the avatar doesn't receive yet.
- **Verification.** Code + loader test authored to the established patterns; not run here (Studio-side). On-device checks (greet on appear, celebrate on turn-finish, overlay-over-working) recorded in TODO.md.
## 2026-06-20 — Pet `working`/tool-use behavior (Android)
**Why.** The behavior-model spec called for a distinct "agent is running a tool" pose — the strongest cross-system convention (Microsoft Agent splits `Think` from `Process`/`Search`; the `pi-animations` indicator splits Thinking · Working · Tool) is that *acting* should look different from *thinking*. Our six `SphereState`s folded tool-use into thinking/streaming.
- **Pet-local tool overlay (`PetAvatar`).** Implemented as a sub-state derived from the already-plumbed `toolCallBurst`, **not** a 7th `SphereState` — zero blast radius on the Sphere or the call sites. `Render` swaps to an optional `workingClip` when `toolCallBurst ≥ WORKING_BURST_THRESHOLD` (0.5) during a `Thinking`/`Streaming` turn, and returns to the base-state clip as the burst decays (the signal ramps to ~1 in 200ms and decays over 1200ms, so 0.5 activates fast and lingers ~600ms — smoothing back-to-back tool calls). Error keeps its own clip; `toolCallBurst` is ~0 outside tool activity, so it never fires spuriously.
- **Opt-in, clip-driven capability (`PetLoader.toAvatar`).** `workingClip` resolves only from an explicit `working` key (no fallback) — a pet without one keeps its base-state clip during tool use, exactly as before. Shipping a usable `working` clip is *itself* the tool-reactivity capability: it drives both the swap and the **Tools** badge (`reactivity.tools = (workingClip != null) && PET_RENDERER_CAPABILITIES.tools`), so the declared `reactive.tools` flag is no longer needed and can't over-promise. Flipped `PET_RENDERER_CAPABILITIES.tools` to `true` (the renderer now consumes the signal).
- **Tests.** `PetLoaderTest`: a `working` clip lights the Tools badge (`Voice · Tools`); a `working` clip with missing files does not; the existing declared-but-no-clip case still clamps to `Voice`.
- **Docs (`docs/pet-spec.md`).** `working` moved from "Forthcoming" into the implemented model: a `Working` row in the state table, a "The `working` overlay" subsection (opt-in, tool-use vs. thinking), the authoring ladder's Rich tier now 7 clips, and the reactivity table's `tools` row now "driven by the `working` clip." Forthcoming trimmed to one-shot reactions + intensity modulation.
- **Verification.** Code + tests authored to the established patterns; not run here (Studio-side). On-device check (clean mode, a `working` clip swapping in during a tool run) recorded in TODO.md.
## 2026-06-19 — Pet reactivity: honest badge + behavior-model spec (Android)
**Why.** Follow-up to the custom-avatar audit. The pet picker badge read `reactivity.summary()` straight from `pet.json`, so a pet declaring `reactive:{tools:true,intensity:true}` advertised "Voice · Tools · Activity" while `PetAvatar.Render` only ever consumed voice — the badge could lie. Separately, the goal was to let pets *associate behavior with agent activity* (show "thinking" vs "writing" vs "speaking"), which needed a documented behavior model rather than a fallback table buried in the clip docs.
- **Honest capability badge (`PetAvatar`, `PetLoader`).** Added `PET_RENDERER_CAPABILITIES` (the live signals `Render` actually consumes today: voice only) and clamp a pet's effective `reactivity` to `declared AND supported` in `toAvatar`. One forward-compat switch: flip a flag there the day the renderer learns a signal and every manifest that already declared it lights up. `PetLoaderTest` gained a case asserting declared tools/intensity are dropped from the badge.
- **Friendly `writing` alias (`PetLoader.STATE_CLIP_CHAIN`).** The Streaming (output-producing) state now resolves `writing` → `streaming` → … so authors can target it with the intuitive key; tidied the Speaking/Error chains to fall back through related activity clips before idle. Backward compatible (existing `streaming`/`speaking`/`error` keys still resolve).
- **Behavior-model spec (`docs/pet-spec.md`).** New "Agent states & pet behavior" section: what each of the six activity states means, the friendly clip-key vocabulary + fallback chains, a loop-vs-one-shot note, and a Minimal→Basic→Standard→Rich authoring ladder. A "Forthcoming behavior" subsection specifies the designed-not-yet-rendered tier — a distinct `working`/tool-use clip, one-shot reactions (greet/celebrate/attention), and continuous tool/intensity modulation — so authors can plan. Reactivity table updated to state the clamp.
- **Prior-art grounding.** Researched the convention (no first-party "Codex pet" exists — the agent-state→mascot pattern is third-party only: `pi-animations`, Peon Pet; canonical spec lineage is Microsoft Agent's `.acs` animation set, with Live2D/VRM/VTuber lip-sync and game-dev FSMs converging on idle-base + thinking/working/output split + amplitude-driven talking + one-shot reactions). The thinking≠tool-use split is the strongest cross-system signal and drives the recommended `working` state. Sources cited in TODO follow-up context.
- **Verification.** Code + test authored to the established patterns; not run here (Studio-side). Behavior roadmap (working state, one-shots, intensity modulation) recorded in TODO.md.
## 2026-06-19 — Custom-avatar audit: storage fix, loader unification, tests, docs (Android)
**Why.** An audit of the just-shipped swappable agent-avatar / "pet" feature found one blocking bug and a set of clarity/coverage gaps. The only documented way to install a pet (and a sphere skin) was `adb push … /sdcard/Android/data/<pkg>/files/{pets,spheres}/` — i.e. external app-scoped storage — but both loaders read from `context.filesDir` (internal, `/data/data/<pkg>/files/`), which is not `adb push`-able on a non-rooted device. So the documented side-load path could never work, on either flavor. Secondary gaps: no in-app hint that pets exist, no user-docs coverage of avatars/skins at all, and the pure loader logic was Context-coupled and therefore untested.
- **Shared storage layer (`UserContentDir.kt`, new).** Both `PetLoader.userDir` and `SphereSkinLoader.userDir` now resolve through one helper that prefers external app-scoped storage (`getExternalFilesDir(null)` = `/sdcard/Android/data/<pkg>/files/<name>/`, reachable by `adb push`, no runtime permission on API 19+) and falls back to internal `filesDir` only when external is unmounted. Single source of truth for "where side-loaded customization content lives" — fixes the bug once for pets and sphere skins together. The `adb push` commands the docs already showed are correct against this location.
- **Testability refactor (`PetLoader`, `SphereSkinLoader`).** Added pure `loadPets(dir: File)` / `loadUserSkins(dir: File)` overloads (no Android Context); the Context overloads delegate. The validation/resolution/skip-invalid path is now unit-testable against a temp directory, mirroring `DashboardManageDiskCache`'s `dir: File` shape.
- **Tests (`PetLoaderTest`, `SphereSkinLoaderTest`, new).** 17 + 6 JUnit cases covering parse, id/label fallbacks, schema-version + missing-`idle` + missing-file rejection, the `safeChild` **path-traversal guard** (a real `../escape.png` outside the pack is refused), fps clamping, one-bad-pack-doesn't-break-the-rest, sort order, and empty/absent dirs. No real bitmaps needed (the loader only checks `isFile`); `unitTests.isReturnDefaultValues=true` makes `Log.w` a no-op so no `mockkStatic`.
- **In-app discoverability (`AppearanceSettingsScreen`).** Added an "Add your own pet" pointer line under the Agent-avatar chips (mirrors the existing sphere-skin pointer) so users learn the feature exists even with no pets installed.
- **Docs (`docs/pet-spec.md`, `docs/sphere-spec.md`).** Corrected the storage prose (app-scoped external, external-preferred / internal-fallback, both flavor paths), cross-linked the two specs under one "customize the agent avatar" framing, fixed a "Agent sphere" → "Agent avatar → Sphere skin" naming drift, and added two authoring caveats: a present-but-undecodable image renders blank (not caught at load), and frame-sequence pets decode every frame at full resolution into RAM (keep frames small / prefer sprite sheets).
- **User docs (`user-docs/features/custom-avatars.md`, new + nav).** New user-facing page covering the two-level avatar→skin model, the reactivity badges, how to add skins and pets, reduced-motion behavior, and troubleshooting; wired into the Features sidebar.
- **Verification.** Tests authored to the established temp-dir pattern and validated against the actual `toAvatar`/`toSkin` contracts (read from source); not run here (Android build/test is Studio-side). Follow-ups (per-frame memory cap/downsample, decoded-clip cache to kill re-decode churn) recorded in TODO.md.
## 2026-06-18 — Chat: mid-session model switch, error surfacing, model-scope UI (Android)
**Why.** On-device testing surfaced four linked issues. (1) Switching the in-chat model in an *existing* conversation showed the pick but the turn still ran the old model. (2) A failed turn (grok-4.3 hitting xAI's 200-tool cap) flashed an error bubble then vanished. (3) The chat header and agent drawer showed the global/profile model, not the session's live model — the drawer even paired the global model *name* with the session *provider* (`gpt-5.5 · xAI Grok`). (4) An upstream-injected `[System: the active model changed …]` marker rendered as a chat bubble.
- **Session-scoped model switch (`GatewayChatClient.prewarmAwait`, `ChatViewModel.selectModel`).** `selectModel` called fire-and-forget `prewarm()` then immediately `setModel()`, so `config.set {key:"model"}` ran with `liveSessionId == null` and upstream applied it as a GLOBAL write — never touching the live session (verified: `_apply_model_switch`, `tui_gateway/server.py:2134`, is the session-scoped in-place swap the CLI/TUI `/model` uses). Added a suspending `prewarmAwait()` that resolves/resumes the live session before returning; `selectModel` awaits it then applies `setModel` session-scoped, or skips the global write and defers to the next `session.create` override when there's genuinely no session. Confirmed on-device: `sessionScoped=true` → `session.info` flips → turn runs the picked model.
- **Errors never swallowed (`GatewayChatClient.dispatchOn`, `ChatHandler.loadMessageHistory`).** Root cause: `dispatchOn` (marshals turn callbacks to the main thread) omitted `onStatusUpdate`, so it fell back to the data class's default no-op — the server's `❌` terminal-error lifecycle line never reached `markError`, the turn wasn't badged `Error`, and `onComplete`'s post-turn history reload (which a non-errored turn runs) wiped the client-only error bubble. Wired `onStatusUpdate` through `dispatchOn` (also restores live status lines, previously dead on the gateway), and hardened `loadMessageHistory` to re-inject local `Error`-badged assistant messages the server transcript lacks, so no reload path can swallow a failure.
- **Model display scoped to the session (`ChatScreen` header, `ConnectionInfoSheet.AgentSheetHeader`).** The header subtitle resolved `profile.model ?? serverModelName`; now mirrors the input chip (`selectedModelOverride ?? gatewayCurrentModel ?? profile ?? server`). The agent-sheet header was pairing the global model name with the session provider; it now takes a `sessionModelName` so model+provider come from one scope, and adds a quiet "Server default: …" caption only when the session diverges (the always-visible global-vs-session split; the redundant in-section split was removed).
- **System steering markers hidden (`ChatHandler`, `ConnectionViewModel`, `RelayApp`, `ChatSettingsScreen`).** Upstream injects `[System: …]` model/personality-change markers into history for the LLM (`tui_gateway/server.py:1769`); we rendered them as bubbles. `loadMessageHistory` now drops `role:system` `[System:`-prefixed rows by default (desktop/TUI parity), gated by a new `ChatHandler.showSystemMarkers` flag wired from a default-off "Show system messages" debug toggle in Chat Settings (DataStore-backed, mirrors `parseToolAnnotations`).
- **Verification.** `:app:assembleSideloadDebug` BUILD SUCCESSFUL; deployed to device; each fix confirmed on-device via filtered logcat traces. Separate known item (not an app fix): xAI/grok models exceed the provider's 200-tool cap with the full relay toolset (server-side).
## 2026-06-18 — Chat UX: model-picker apply, relay inbound images, smooth profile switch (Android)
**Why.** An audit of profile switching and the chat composer surfaced three issues: (1) the in-chat model picker showed the picked model but the agent ran on the account's global default; (2) an agent-returned server-local image showed a path/"on server" notice instead of rendering, even when paired to the relay; (3) switching profiles visibly tore down and rehydrated the conversation.
- **Model picker binds to `session.create` (`GatewayChatClient`, `ChatViewModel`, `GatewayModels`).** Verified against upstream `tui_gateway/server.py`: a model is a per-session override, applied via `config.set {session_id,…}` on a live session or `model`/`provider` params on `session.create` for a fresh one. The app only did the first; on a brand-new chat the `config.set` carried no `session_id` (upstream treats it as a no-op) and `session.create` omitted the model, so the agent built from the global default. Added a live `sessionModelProvider` (mirrors `sessionProfileProvider`) so the picker's model+provider bind onto each `session.create`; mid-session switches still go through `config.set`. A profile switch now retires an explicit pick (the profile defines its own model) and seeds the picker pill from the profile model so the header doesn't lag the round-trip. SSE paths already carried the model in the request body. Tests added for the new binding.
- **Relay-backed inbound images (`ChatImageContent`, `ChatScreen`, `ChatViewModel`).** Markdown images (`![](src)`) flowed through a renderer that only understood `http(s)` → Coil; a server-local path fell to a static "image is on the server" notice and never consulted the relay (the relay media route was only wired to the `MEDIA:` marker path). Added a `RelayServerImageResolver` CompositionLocal, provided by ChatScreen from `ChatViewModel.resolveServerImage`, which fetches an absolute server path through the relay's bearer-auth `/media/by-path` route (same route + sandbox the `MEDIA:` path uses), decodes, caches (bounded LRU keyed by path), and renders inline with tap-to-zoom. Null when unpaired → unchanged standard (no-plugin) behavior. Complementary nudge: `composeInjectedContext` appends a one-line media-capability hint to the SSE `system_message` when a relay route is configured (`RelayHttpClient.mediaUrlConfigured()`), surfaced in the "What the agent sees" audit sheet. SSE-only — the gateway has no per-turn system slot, so there the client render fallback (and upstream's own `MEDIA:` instruction) carry it.
- **Profile-switch transition (`ChatViewModel.switchProfileContext`, `ChatScreen`).** Stopped clearing the message list synchronously before the async history fetch; the previous transcript is held and swapped atomically when the new history resolves, so the `LazyColumn`'s per-item `animateItem()` cross-fades old→new instead of blanking to an empty/"Loading…" state. The top loading row is suppressed while held content is on screen.
- **Verification.** Rebased `Codename-11/fix-ui-ux-issues` onto `dev` first (its only unique change was already on `dev`). `:app:compileSideloadDebugKotlin` + `:app:compileSideloadDebugUnitTestKotlin` BUILD SUCCESSFUL (no new warnings in the changed files); `GatewayChatClientTest` extended with model-binding cases. On-device verification via Studio.
## 2026-06-18 — Cold-start keystore contention + honest loading states (Android)
**Why.** A cold-start logcat trace showed the chat header's identity/model/approvals lagging seconds behind first frame. The cause was on-device, not the network: `EncryptedSharedPreferences.create()` decrypts a Tink keyset via a KeyStore op (~0.6–1 s on StrongBox) and Tink serializes those process-globally, and the app was building **three** keysets at startup (a throwaway legacy-sentinel `AuthManager`, the active connection's token store, and the dashboard cookie store) — they thrashed the lock (`Long monitor contention … AndroidKeysetManager.build()`, `waiters` up to 4; a `by lazy` held for **2.369 s**). The relay auth round-trip itself was ~150 ms. Separately, a design constraint surfaced: never display unconfirmed server state (model/provider/approvals) as if confirmed — show an honest loading state and make the load fast, don't cache a maybe-wrong value.
- **Process-global store cache + raw-build factory (`SessionTokenStore.kt`).** `SecureStoreCache.getOrBuild(prefsName){…}` (a synchronous `ConcurrentHashMap.computeIfAbsent`) builds each prefs file's keyset once process-wide, and `buildRawTokenStore()` is the shared backend factory. Synchronous on purpose so the SAME instance serves both the suspend token path (wrapped in IO) and the synchronous OkHttp cookie-jar path.
- **Sentinel deferral (`AuthManager.kt` + `ConnectionViewModel.kt`).** Added `eagerHydrate` (false for the legacy sentinel that `ConnectionViewModel` builds at field-init and replaces the moment the active connection hydrates), so it no longer decrypts a keyset just to be discarded. Re-gated the pre-StrongBox `hermes_companion_auth → _hw` migration on the **file name** (not the sentinel's connection id) so deferral stays correct, and marker-gated it (`legacy_migrated`) so the legacy file is read at most once ever.
- **Cookie keyset unification (`DashboardApiClient.kt`, `UpstreamTransportController.kt`, `ConnectionViewModel.kt`, `DataManager.kt`).** The dashboard cookie store now rides the connection's **token** file (threaded `tokenStoreKey` via a `tokenStoreKeyProvider`) instead of its own `hermes_dashboard_<id>` keyset — eliminating the second build. One-shot, marker-gated migration (`dashboard_cookies_migrated`) copies existing cookies across on first access; failure just means a one-time Manage re-login (cookies are re-obtainable, unlike the relay token). Also fixed a double `store.load()` per cookie request.
- **Honest loading + fade-ins (`LoadedFadeIn.kt` + 5 call sites).** New shared `LoadedFadeIn`/`RelaySkeletonLine` mirroring the header's skeleton→identity spec. Wired into the header subtitle (model fades in when confirmed), the agent sheet's model line, `ContextMeterBar` (fade+expand), the session drawer (loading→list crossfade), and Manage (Loading→Loaded crossfade). The agent sheet's YOLO switch now shows "Checking…" instead of rendering the unknown (null) state as a definitive "off". No model/provider/approvals value is ever cached and shown as confirmed.
- **Also (header/chrome).** Dropped the redundant LAN/Tailscale endpoint chip from the chat top bar (the footer status strip already shows `<status> / <route>`) and made that footer strip tappable → Connections; subtitle no longer renders a `none` personality.
- **Verification.** `:app:assembleSideloadDebug` BUILD SUCCESSFUL; deployed to device. Cold-start logcat before→after (warm launch, same device): **3 keyset builds → 1**; the sentinel's "no stored session_token → Unpaired" build is gone; first-frame→Paired **~2.9 s → ~0.95 s**; steady-state (post-migration) trace shows **zero** `Long monitor contention` events. On-device check still wanted: Manage/voice remain signed in after the one-time cookie migration.
## 2026-06-18 — Connection toast stepper + chat header de-clutter (Android)
**Why.** Two adjacent UI/UX gaps. The floating connection-status toast already slid in from the top and supported swipe-up dismiss, but it rendered its trace entries as flat `label: detail` text and the swipe silently accumulated to a threshold then snapped — while the cold-start sphere right next to it had a far nicer live stepper (`·`/spinner/`✓`/`✕`). The two were built separately and never unified. Separately, the chat header subtitle (`personality · model · ⚡ approvals off`) was being clipped because the trailing actions row (endpoint chip + Share + Terminal + Settings) won the width fight; the appended approvals text made it worse.
- **Stepper state model (`RelayUiState.kt`).** Added `ConnectionStepState { Pending, Active, Done, Failed }` and an optional `state` on `ConnectionHandoffTraceEntry` (defaulted null so every producer keeps compiling). Null means "infer from position + snapshot"; producers that know a surface's verdict stamp it explicitly.
- **Probe entries stamp real states (`ConnectionViewModel.kt`).** `buildGlobalConnectionProbeEntries` now tags Route=Done, and API/Relay as Active (Probing) / Done (Reachable) / Failed (Unreachable), plus the relay-socket "Session" step as Active.
- **Toast redesign (`ConnectionHandoffBanner.kt`).** `ConnectionStatusToast` renders entries as a live stepper (fixed-width monospace glyph + spinner via `LaunchedEffect`, mirroring the splash's `StartupCheckRow` vocabulary), reusing `ConnectionStatusBadge`'s green for Done and `colorScheme.error` for Failed. Swipe-dismiss now tracks the finger with an `Animatable` offset + fade, flinging off-screen past threshold (then firing `onDismiss`) or springing back; keyed on status identity (title+tone) so frequent `updatedAtMs` trace bumps don't reset an in-flight swipe. Warning/Error poses get a bottom divider + "Open <destination> →" link for discoverability. The legacy edge-variant `ConnectionStatusBanner` was left as-is (only referenced within its own file).
- **Chat header (`ChatScreen.kt`).** Dropped the inline ` · ⚡ approvals off` annotation from the subtitle (now a plain single line). Added an amber ⚡ `Icons.Filled.Bolt` to the app-bar actions, shown only when approvals are effectively off, tapping into the agent sheet where the full explanation already lives. Share moved into a `⋮` `DropdownMenu` (rendered only when there's a conversation to share), leaving Terminal + Settings + the endpoint chip as the visible actions. `RelayChromeIconButton` gained optional `tint` / `borderColor` params for the amber treatment.
- **Verification.** Pending — Kotlin-only UI changes; per the project dev loop these build via Android Studio's run button (not `gradle build` from here). `./gradlew lint` recommended before push.
## 2026-06-18 — Terminal: TUI input correctness + chrome cleanup (Android + relay)
**Why.** On-device terminal use surfaced input bugs and wasted chrome, benchmarked against Orca's mobile terminal. The extra-keys bar clipped labels ("CTRL" → "CTR") because it was weight-distributed across a fixed width; the on-screen arrows and PASTE bypassed the emulator and sent fixed/raw bytes (wrong inside TUIs and unsafe for multi-line paste); and the header + tab strip + a stray inset ate vertical space, especially in the common single-tab case. Separately, the relay wrapped each PTY in the user's default tmux, inheriting tmux's 500ms `escape-time` and `screen` `$TERM` — the classic source of laggy ESC, mangled Alt, and degraded color in vim/htop.
- **Extra-keys bar (`ExtraKeysToolbar.kt`).** Rewrote from a weight-divided `Row` to `horizontalScroll` with fixed-min-width keys, so labels never clip and the cluster scrolls when wider than the screen (Orca's strategy). Then compacted to Orca's proportions — 32dp height, 36dp min width, 12sp, tighter padding/spacing — and centralized the key haptic.
- **Mode-aware special keys (`index.html` + `TerminalScreen.kt`).** Added `window.termSendKey(name)` that reads xterm's `applicationCursorKeysMode` and encodes arrows/Home/End as SS3 (`\eOA`) vs CSI (`\e[A`); the toolbar arrows now route through it. PASTE routes through `term.paste()` (bracketed paste) instead of a raw `sendInput`, so multi-line paste no longer auto-executes.
- **Compact header (`TerminalScreen.kt`).** Replaced Material's fixed 64dp `TopAppBar` with a ~52dp custom `Row` + `statusBarsPadding`: status is shown once, inline with the title (a `ConnectionStatusBadge` dot + one concise word, ellipsized via `weight(1f, fill = false)` so it can't push the actions off-screen), the whole block opens the info sheet. The subtitle no longer renders the full `hermes-<deviceId>-tabN` wire id (it was wrapping to two lines).
- **Less wasted vertical space.** The tab strip + its divider now render only for 2+ tabs; with one tab the new-tab "+" lives in the header instead. Dropped a redundant `navigationBarsPadding()` on the keys bar (the app Scaffold's bottomBar already owns the nav-bar inset, leaving a dead gap), and added an 8px bottom gap in `#terminal` so the last row clears the keys bar.
- **Isolated, TUI-tuned tmux (`plugin/relay/channels/terminal.py`).** Sessions now spawn on a dedicated `-L hermes-relay` socket with a generated `~/.hermes/hermes-relay-tmux.conf` (written lazily, best-effort): `escape-time 0`, `default-terminal "tmux-256color"`, truecolor `terminal-features/overrides`, `mouse on`, `focus-events on`, `set-clipboard on`, `aggressive-resize on`, `status off`. A dedicated socket is the only safe way to set server-global `escape-time` without altering the user's own tmux; persistence is unchanged (the socket's server outlives the relay process).
- **Verification.** `:app:assembleSideloadDebug` BUILD SUCCESSFUL and deployed to device; `python -m unittest plugin.tests.test_terminal_channel` (4 tests) + `py_compile` pass. Relay change hand-deployed to the staging box and verified live on the `hermes-relay` socket (`escape-time 0`, `default-terminal tmux-256color`, `status off`, `mouse on`, `focus-events on`). tmux 3.4 with `tmux-256color` terminfo present.
## 2026-06-18 — Native secure routes: split connection Features from Routes (Android)
**Why.** Connection setup conflated two separate questions — what a Hermes connection can *do* (features) and how this phone *reaches* it (route) — which coupled Relay features to a single transport. Modeling them separately lets a user enable Relay tools over any route (LAN, Tailscale, public HTTPS, VPN, or a plugin-provided secure proxy) and sets up a plugin-assisted native encrypted route that does not require Tailscale. The standard path stays direct-to-upstream and plugin-free. (Backfilled log entry — the work landed in PR #88; full design in `docs/plans/2026-06-18-native-secure-routes.md`.)
- **Split connection model.** `ConnectionsSettingsScreen` / `ActiveConnectionSections` now render distinct **Features** and **Route** sections; `Endpoint.kt` + `ConnectionData.kt` carry the route/role model and `QrPairingScanner` threads it through pairing.
- **Plugin secure proxy route.** A `plugin_proxy` route role surfaces as a "Secure proxy" option with encrypted / pinned-TLS treatment (recommended, not forced) alongside the existing LAN / Tailscale / public / custom roles.
- **Docs.** Added the `2026-06-18-native-secure-routes` plan and a connections split-model mockup; fixed the docs-site hero sphere to keep its canvas backing store synced to the CSS box (`HeroDemo.vue`).
- **Verification.** CI green on PR #88 (Android Build + Lint + Test).
## 2026-06-18 — Android onboarding permissions review surface
**Why.** Android onboarding already kept the standard path clean, but permissions were scattered between feature-specific prompts, Bridge, and Android Settings. A central review page makes the model explicit: standard Chat and Manage do not need phone-control permissions, while voice, camera, notifications, and sideload Device Control remain opt-in.
- **Shared permission snapshot.** Added `AppPermissionStatusProbe` so Bridge and Settings read the same runtime grants and special-access switches: notifications, microphone, camera, notification listener, accessibility, screen capture, overlay, contacts, SMS, phone, and location.
- **Permissions screen.** Added `PermissionsStatusScreen` with Standard Hermes, On demand, and flavor-aware Device Control sections. Rows show required/optional/session status and link to the relevant Android Settings surface or Bridge session grant.
- **Onboarding and Settings entry points.** The Power Tools onboarding page now has a "Review permissions" action, and Settings -> App includes a Permissions row. Google Play builds show the sideload Device Control section as unavailable rather than implying hidden phone-control permissions.
- **Verification.** `:app:compileGooglePlayDebugKotlin`, `:app:compileGooglePlayDebugAndroidTestKotlin`, and `:app:compileSideloadDebugKotlin` pass with `ANDROID_HOME` pointed at the local SDK.
## 2026-06-17 — Chat transparency + provenance polish: injected-context audit sheet, spoken-turn badges, version-skew error
**Why.** On-device voice testing surfaced two transparency gaps and two papercuts. The per-turn system context the phone injects (persona + phone status + voice hint) was invisible — no way to audit what the agent actually receives. Spoken voice-mode turns were indistinguishable from typed ones in the scrollback, while realtime turns were already badged. A field an older relay plugin doesn't accept produced a misleading "Network error · HTTP 400" with a dead Retry. And the floating connection toast's Warning tone was semi-transparent, letting content bleed through.
- **Injected-context audit sheet.** `ChatViewModel.composeInjectedContext()` extracts the per-turn system-message composition (persona/profile precedence + phone-status block + per-turn interface context) into one builder used by both `startStream` (which sends `combinedSystemMessage`) and the new `previewInjectedContext()` — so the audit can't drift from what is sent. `combinedSystemMessage` is built byte-for-byte as before (`listOfNotNull` over the raw blocks); per-block fields null out blanks only for display, and the resolved profile is passed in to preserve the no-skew invariant with `modelOverride`. Tapping the `ContextMeterBar` (now with an ⓘ affordance) opens `InjectedContextSheet` ("What the agent sees"); on the gateway path the persona block is labeled "added server-side — not sent from this device", since the server owns the soul + personality overlay there.
- **Spoken-turn badges.** Voice-mode replies are tagged "Voice", realtime replies keep "Realtime Agent"; both share a speaker glyph as the modality marker (`MessagePathBadge` gained an optional leading icon). The Voice tag is set on the assistant placeholder from the per-turn interface-context signal and rides the id-swap + content updates.
- **Badges survive history reload.** `ChatHandler.loadMessageHistory` now carries provenance badges forward by message id when it rebuilds the list from server data — the post-turn reload previously wiped them (this also fixes the pre-existing loss of "Stopped"/"Error").
- **Version-skew error.** `RelayErrorClassifier` maps a 400 whose body names an unsupported *field* to "Relay update needed" (non-retryable), distinct from a bad *value* like "unsupported codec" (which keeps its normal classification).
- **Connection toast opacity.** `ConnectionStatusToast` composites its container color over the theme `surface` so the floating overlay is always opaque while keeping each tone's tint; the in-flow `ConnectionStatusBanner` is intentionally left translucent (it blends with a known backdrop).
- **Verification.** `:app:compileSideloadDebugKotlin` BUILD SUCCESSFUL; `./gradlew lint` clean. On-device confirm via Studio/sideload.
## 2026-06-17 — App theming: theme-aware brand tokens, app themes, hot-swappable sphere
**Why.** Chat (and other brand-styled surfaces) were effectively hardcoded dark: the `RelayRefresh` brand palette was a single dark-only `object` of `val Color(...)` constants that ~150 call sites referenced directly, bypassing the Material light scheme. The goal was three-fold: fix that alignment, add real app themes (light/dark plus the Nous Hermes baselines), and make the agent sphere a hot-swappable component with user-authored skins.
- **Theme-aware brand tokens (the fix).** New `ui/theme/BrandPalette.kt` defines a 21-token `BrandPalette`, the `LocalBrand` CompositionLocal, a `toColorScheme()` derivation (Material scheme is now derived from the palette, never authored separately), and the `AppTheme`/`AppThemes` registry. `RelayRefresh` was converted from constants into a **snapshot-backed façade** over an `activePalette`: every `RelayRefresh.X` is now a getter reading a `mutableStateOf`, so reads in composition and draw phases subscribe and repaint on theme change — the ~150 existing call sites became theme-reactive with no edits. `HermesRelayTheme(appThemeId, themePreference, fontScale)` resolves the theme + light/dark/auto mode into one palette, drives the Material scheme + `LocalBrand`, and mirrors it into the façade via `SideEffect`.
- **App themes (8).** Hermes Relay (the brand, with full light + dark) plus ports of the canonical Nous dashboard baselines from upstream `web/src/themes/presets.ts` — Hermes Teal, Nous Blue (light), Midnight, Ember, Mono, Cyberpunk, Rosé. Per the hybrid model, the brand honors Light/Dark/Auto while the character themes are fixed-mode looks (matching how Nous ships themes; light mode lives as the Nous Blue theme). `ConnectionViewModel` gained an `appTheme` id pref (DataStore `app_theme`); the picker is a swatch gallery in `AppearanceSettingsScreen`, and the Light/Dark/Auto control disables + explains itself for fixed-mode themes.
- **Flourish alignment.** The glow/gradient/markdown-highlight flourishes across 13 files keyed off `isSystemInDarkTheme()`; they now read `LocalBrand.current.isDark`, so a fixed dark theme keeps its flourishes on a light-mode phone (and vice-versa). `isSystemInDarkTheme()` now lives only in `Theme.kt` (resolving Auto).
- **Hot-swappable sphere.** The core algorithm (`MorphingSphereCore.kt`, mirrored in `preview/web/sphere.js`) was left untouched — parity contract preserved. A new `SphereSkin` layer supplies per-state colors/params and declares its reactivity (voice/tools/intensity/gaze), gated inside `MorphingSphere` so reactive inputs are optional + detectable. `SphereRegistry` ships Adaptive (recolors to the active theme via `LocalBrand`), Classic (the original look), Aurora, Solar, and Mono. `MorphingSphere` gained a `skin` param defaulting to `LocalSphereSkin.current`, so all ~11 call sites are untouched; `RelayApp` provides the resolved skin + available set. User skins load from app-private `spheres/*.json` via `SphereSpec` (kotlinx.serialization, data-only, validated, invalid files skipped) → `SphereSkinLoader`; surfaced in the picker with capability badges and a "Custom" tag. Format documented in `docs/sphere-spec.md`.
- **Verification.** Reviewed-but-not-compiled — no Android SDK in this worktree and Studio owns builds. Brace/paren balance checked on all new/edited files; import resolution and call-site compatibility reviewed by hand. `./gradlew lint` + on-device confirm pending (via Android Studio). Palettes are single `BrandPalette` literals, easy to tweak after an on-device pass.
## 2026-06-17 — ConnectionViewModel decomposition: extract transport/pairing/profile collaborators (ADR 34 follow-up)
**Why.** `ConnectionViewModel` was the one god-object the ADR 34 fence deferred — ~5,531 lines reaching across both sides of the upstream/relay package boundary (the `HermesApiClient`/`GatewayChatClient`/`DashboardApiClient` zoo *and* the relay `ConnectionManager` *and* pairing *and* profiles). Goal: move cohesive concerns into named, testable collaborators in a new `viewmodel/connection/` package, behind the ViewModel's existing public surface, so the standard-vs-relay wiring lives in explicit seams. Pure mechanical, behavior-preserving extraction — the whole-module compile (production + every test source unchanged) is the public-API-preservation guard. Continues the `ConnectionSwitchCoordinator` precedent.
- **`PairingController`** (229 lines). Owns the paired-devices list (`GET /sessions`) + management (`loadPairedDevices`/`revokeDevice`/`extendDevice`/`revokeChannelGrant`, incl. the optimistic local removal and the full-grants-rebuild encoding) and the insecure-ack DataStore flags (`insecureAckSeen`/`insecureReason`/`setInsecureAckComplete`). Self-contained — nothing else reads its state except `applyPairingPayload`'s `insecureReason.value`. The ViewModel delegates unchanged.
- **`UpstreamTransportController`** (269 lines). Owns the per-connection encrypted `DashboardCookieStore` cache, a single consolidated `DashboardApiClient` factory (was 4+ scattered build sites — the plan's headline), the cached `GatewayChatClient` (lazy build + mid-turn LAN/Tailscale retarget) with its availability tier + sticky-Unsupported verdict, and the per-endpoint capability snapshot + `chatMode` + the `streamingEndpoint`-preference resolution. The `@Synchronized` gateway-cache lock moved *with* the state (now the controller instance) — same mutual exclusion, different monitor. `rebuildApiClient` pushes the probed snapshot via `setCapabilitiesAndMode`.
- **`ProfileController`** (313 lines). Owns the merged `agentProfiles` list (relay `auth.ok` ∪ dashboard `/api/profiles`), the per-connection selected-profile state machine + its three persistence stores, `profileDisplayAlias`, `activeSessionTransport`, and the per-profile last-session restore. Because the state machine is co-driven by ViewModel-level lifecycle observers (connection switch / active-connection change / agent-profile arrival / gateway-availability settle), those observers stay in the ViewModel and call `profileController.*` lifecycle hooks **in their original order** — orchestration stays put; only state + logic moved, so the state machine is now unit-testable in isolation. The three stores are exposed as public vals so the connection-lifecycle orchestrators (`removeConnection`, duplicate-merge, `resetAppData`, `saveLastSessionId`) keep their clear/persist call sites byte-identical.
- **Deferred: `RelayTransportController`** (the plan's Step 2). Left in place per the plan's explicit "too entangled — don't force it" rule. `ConnectionManager` is referenced at ~38 ViewModel sites, including eager `StateFlow` initializers (`relayConnectionState`, `activeEndpoint`, `effectiveApiServerUrl`/`RelayUrl`/`DashboardUrl`, `relayReady`, `insecureMode`), the `ConnectionSwitchCoordinator`, and the `relayHttpClient`/`ScreenCapture`/`tailscaleDetector` URL-provider lambdas. The relay/route methods are inseparable from the central `_relayUrl`/`_apiServerUrl` state (also written by non-relay orchestrators + the switch coordinator) and from shared connection-store helpers (`persistActiveConnectionUrls`, `mergedRouteCandidates`); the route-probe public nested types (`RouteProbeStatus`, `RelayReachable`) are part of the frozen public API. A faithful extraction needs ~18 injected callbacks/refs — relocating coupling into lambdas rather than removing it, against the plan's "named seams, not emergent shared state" goal and at a regression risk the compile-plus-focused-slice verification can't catch. The other three controllers stand on their own; the `ChatTransportProvider` capstone (optional) is likewise not attempted.
- **Result.** `ConnectionViewModel` 5,531 → 5,089 lines (−442); cohesive transport/pairing/profile concerns now live in 811 lines of `viewmodel/connection/` collaborators with narrow, provider-injected interfaces. Public API unchanged.
- **Verification.** `:app:compileSideloadDebugKotlin` + `:app:compileSideloadDebugUnitTestKotlin` BUILD SUCCESSFUL after each extraction (every caller + test source compiles unchanged → public surface byte-identical). Focused slice `*ArchitectureBoundaryTest` (Konsist fence — `viewmodel/connection/` is outside `network.*` so it imports both worlds freely, fence stays green) + `*RelayUrlDeriverTest` + `*ConnectionSwitchTest` passes. `./gradlew lint` clean. On-device confirm pending (Bailey, via Studio).
## 2026-06-17 — Voice mode audit: relay bug fixes, spoken-output hint, Google enhanced voice
**Why.** A post-refactor audit of the standard and relay voice paths surfaced one reachable correctness bug plus several enhancement opportunities: leverage the desktop-style non-persisted per-turn context to instruct spoken-output formatting, and expose Google/Gemini enhanced voice (tone tags, voice/model/persona) that upstream added in recent PRs.
- **Relay realtime-agent loop drift (correctness).** The non-native ("render-after-Hermes") websocket loop in `plugin/relay/realtime_agent/broker.py` lacked a `playback.drained` branch, so the end-of-turn ack every client sends fell through to the unsupported-message error; the client treats `voice.error` as fatal and tore the session down on every turn whenever a non-native provider was configured. Extracted the client→server messages identical across the native and non-native loops (`session.start`, `session.resume`, `client.ack`, `playback.drained`, `hermes.confirm`) into a shared `_handle_common_client_message` dispatcher used by both loops so they can no longer drift, and added the missing `input_audio.clear` handling to the non-native path. The native loop's per-message `provider_task.done()` break check is preserved exactly. (`hermes.confirm` echo is by-design in both loops — the realtime model answers confirmations via its `hermes_confirm` tool, which returns `forwarded_to_hermes_ui` — so it was left unchanged.)
- **TTS file leak.** `plugin/relay/voice.py` synthesize streamed upstream-written `~/voice-memos/*.mp3` files and never deleted them. It now passes its own temp `output_path` into the TTS tool and deletes the artifact after streaming (reads bytes into memory and returns a `web.Response` so cleanup can run; audio is bounded by `MAX_TEXT_CHARS`).
- **Realtime "lab" session binding.** `plugin/relay/realtime_voice.py` `handle_ws` authenticated but never checked the websocket caller created the session. Added `_auth_matches_session` (mirrors `voice_output.py`) binding sessions to the creating principal's kind + device-id / token-hash, reusing `voice_auth`'s shared `AuthPrincipal` / `_bearer_from_request`.
- **Spoken-output formatting hint (standard + relay).** Enriched the per-turn voice interface context (`VoiceViewModel.STABLE_VOICE_INTERFACE_CONTEXT`) to tell the model its reply will be spoken — short conversational sentences, no markdown/emoji/URLs. This rides the existing non-persisted `system_message` slot (upstream `ephemeral_system_prompt`), so it never lands in history. Because the gateway `prompt.submit` RPC has no system-message slot, `ChatViewModel.startStream` now forces any turn carrying a per-turn interface context (voice) onto an SSE endpoint so the hint always reaches the model.
- **Provider-aware enhanced voice (Gemini + xAI) — relay path.** `/voice/synthesize` accepts optional per-request overrides (`voice`, `model`, `audio_tags`, `persona_prompt`, `language`) mapped onto the active provider. Upstream's `text_to_speech_tool` has no per-call override surface, so the relay merges overrides into a config copy and invokes the provider generator directly — `_generate_gemini_tts` (voice/model/`audio_tags` tone-tag rewrite/inline persona via a temp file) or `_generate_xai_tts` (`voice`→`voice_id`, `audio_tags`→`auto_speech_tags`, `language`). Gemini's audio-tag rewrite fails soft when the auxiliary LLM is unavailable. `/voice/config` advertises a provider-aware `tts.enhanced` capability block. Android plumbs `EnhancedVoiceOverrides` from new persisted prefs through `RelayVoiceClient.synthesize` + the relay adapter, with an "Enhanced Voice (<provider>)" Voice Settings card rendered from the capability flags. OpenAI is excluded — upstream exposes only voice/speed for it (no `instructions` tone steering).
- **Enhanced voice on the streaming renderer (`/voice/output`).** Because `voice_output_enabled` defaults true, the streaming renderer — not `/voice/synthesize` — is the normal relay playback path, so the per-request override was effectively fallback-only. Added xAI `auto_speech_tags` as a per-profile `voice_output:` setting threaded through every layer (config dataclass + env + YAML loader, `voice_output_settings`, profile override, `VoiceOutputSession`, `_provider_options`, `config_payload`, the `PATCH /voice/output/config` allow-list), mirroring `text_normalization`. The render applies `upstream_voice.apply_xai_speech_tags()` (fail-soft) to each chunk before the `voice_lab` `xai_tts` renderer — keeping `voice_lab` standalone and the upstream import in the patch-point. Android: `VoiceOutputConfig.auto_speech_tags` + `updateVoiceOutputConfig(autoSpeechTags=)` + an "Expressive speech tags" switch in the **Hermes Chat + Voice Output** card (xai_tts, persisted with the existing Save buttons). No Gemini streaming provider in `voice_lab`, so Gemini enhanced voice stays `/voice/synthesize`-only.
- **Render-path visibility (troubleshooting).** The streaming-vs-synthesize decision was logcat-only. Added a per-session `DiagnosticsLog` entry naming the active path and a persistent "Render path" row in the Voice Settings card (derived from `voiceOutputConfig`).
- **Docs.** `docs/upstream-surface-matrix.md` gained a "Voice Surfaces (standard vs. relay)" section with an explicit **route-ownership table** (every `/voice/*` route is relay-owned; only dashboard `/api/audio/*` is upstream — no upstream streaming/WS audio route) and an enhanced-voice matrix across both relay paths; `docs/spec.md` Phase V documents the override, the `tts.enhanced` block, and `voice_output` `auto_speech_tags`; `user-docs/features/voice.md` gained an "Enhanced Voice (Gemini & xAI)" section + the streaming speech-tags toggle and corrected the stale `~/voice-memos` note.
- **Standard voice polish.** Pre-flight 25 MB transcribe guard (matches upstream `_MAX_TRANSCRIPTION_UPLOAD_BYTES`) + friendly 413/400 copy in `StandardHermesVoiceClient`; hardened the dashboard audio HEAD probe to also try `/api/audio/speak`.
- **Verification.** Python: `py_compile` on all touched relay modules; relay voice suite green at 103 tests except one pre-existing xAI-OAuth env-dependent failure (identical on the unmodified tree). New tests: non-native `playback.drained` regression (red-on-bug/green-on-fix), Gemini + xAI synthesize-override integration, override-parser + capability-block units, `auto_speech_tags` PATCH round-trip, `apply_xai_speech_tags` call-through/fail-soft. Kotlin not built locally (Studio + `./gradlew lint` are the pre-push gate).
## 2026-06-17 — Upstream/Relay isolation: package fence + Konsist rule + vanilla contract test
**Why.** The load-bearing "standard path = vanilla upstream" invariant was enforced only by `CLAUDE.md` convention (network clients were cleanly named but co-located in one package, so nothing *stopped* a standard-path file importing a relay client), and the standard path had never been validated against *true* vanilla upstream (staging runs the fork with relay routes compiled in). ADR 34 records the decision; this lands all three parts. Net-additive, behavior-preserving.
- **Package fence.** Split `app/.../network/` into `network/{upstream,relay,shared}` — main + mirrored test sources (38 files moved, ~83 touched for import repointing). `VoiceAudioClient.kt` split three ways: the `VoiceAudioClient` interface + `AutoVoiceAudioClient` router → `shared`; `StandardHermesVoiceClient` → `upstream`; `RelayVoiceAudioClientAdapter` → `relay` (co-locating them would force one file to import both worlds). `ChatHandler` → `upstream` (per ADR 3 chat never flows through the relay multiplexer; the handler is fed only by upstream transports). `AndroidManifest` `GatewayKeepAliveService` FQCN and the `ci-android.yml` `RelayUrlDeriverTest` path updated for the move.
- **Hidden coupling surfaced.** The move exposed the one real upstream→relay dependency the import grep couldn't see (it was a same-package bare reference): `ChatHandler` renders phone-action bubbles from the bridge's `LocalDispatchResult` DTO. Resolved by moving that passive DTO to `network.shared` — both sides now depend only on shared to speak it.
- **Konsist boundary test.** `ArchitectureBoundaryTest` (`scopeFromProduction`) asserts `upstream` ⊥ `relay` and `shared` imports neither. Added to the `ci-android.yml` explicit `--tests` list (the broad aggregate hangs, issue #32, so a named test is the only way it runs in CI). Konsist 0.17.3 resolves clean on Kotlin 2.3.21.
- **Vanilla-upstream contract.** `scripts/check-upstream-route-contract.py` source-parses upstream's declared routes (aiohttp `add_*` + FastAPI decorators) — no server boot, no pip, no model keys. Two tiers: REQUIRED standard-path routes fail the build if missing; mode-dependent routes (auth-gate, `/api/pty`, `/v1/models`) only warn. A fork-marker guard refuses to pass against our own fork. `ci-contract.yml` checks out vanilla upstream with no relay bootstrap, asserts the checkout is vanilla, runs the contract; weekly schedule tracks upstream `main` as a drift siren, PR/push use a pinned ref. Notable: in the checked upstream commit the dashboard exposes no `/api/auth/ws-ticket` REST route (it uses the injected session token + a ws `ticket` query param), so the Desktop-style auth-gate routes are advisory, not required.
- **Deferred (tracked in ADR 34).** The `ConnectionViewModel` transport-strategy split — the one true god-object leak — is intentionally deferred as the riskiest change; the fence now contains the blast radius. Same-package redundant imports left behind by the move (e.g. a relay file importing its own `network.relay` sibling) are harmless and not cleaned. The contract job's PR-run `UPSTREAM_REF` defaults to `main` until pinned to a confirmed-public known-good SHA.
- **Verification.** `:app:compileSideloadDebugKotlin` + `:app:compileSideloadDebugUnitTestKotlin` BUILD SUCCESSFUL; `ArchitectureBoundaryTest` passes; contract script PASS against the local upstream clone (12/12 REQUIRED routes). `./gradlew lint`: BUILD SUCCESSFUL (clean, all four variants).
## 2026-06-17 — Gateway parity: live session.info sync + YOLO/Fast + stale-state refreshes
**Why.** An audit (full tui_gateway surface vs. what the official desktop uses vs. what we used) found we were dropping most `session.info` fields and fetching several server lists once. Goal: augment upstream, never show stale state. Verified every contract against the up-to-date upstream clone; a parallel review confirmed the new RPCs match `config.set` exactly and caught three race-window bugs (fixed).
- **More of `session.info` consumed live.** The interceptor now also surfaces `reasoning_effort`, `credential_warning`, `yolo`, and `fast` (added to `serverReasoningEffort`/`serverCredentialWarning`/`serverYolo`/`serverFast` flows); `startGatewayStateSync` gained one guarded collector each. A `/reasoning` change made on the desktop/TUI now reflects instantly instead of only on turn-complete.
- **Credential warnings no longer silent.** `session.info.credential_warning` (present only when the active provider key is missing/invalid) is surfaced once per distinct warning as a ⚠ system notice — dedup'd against the constant `session.info` echoes, cleared when the key is fixed. Previously such turns just failed silently.
- **YOLO + Fast mode.** New session-scoped toggles in the agent sheet (`config.set yolo` value `1`/`0` scope `session`; `config.set fast` value `fast`/`normal`) with optimistic set + rollback, live state from `session.info.yolo`/`fast`, and reset across every session/profile/connection switch. YOLO (approval bypass) renders loud — `error` caption + an `errorContainer` "Approvals are OFF" banner — and stays ephemeral so a backgrounded app can't leave global auto-approve armed.
- **No fetch-once staleness.** `refreshSkills()` + `refreshModels()` (SSE `/v1/models`) now fire on agent-sheet open alongside the personality/model refreshes, so server-side skill/model changes appear without an app reload.
- **Review fixes.** `activateGatewayProfile` now nulls YOLO/Fast (the missing 5th clear site); the `setYolo`/`setFast` optimistic rollback guards against a session switch landing during a slow `prewarm` (re-check client identity + only roll back if we still own the value).
- **Verification.** `:app:testSideloadDebugUnitTest` compiles clean; contract-fidelity review = all PASS. Touches only `GatewayChatClient`/`ChatViewModel`/`ConnectionInfoSheet`. The command-palette skills-refresh-on-open is the one optional follow-up (palette lives in the co-owned `ChatScreen.kt`). `./gradlew lint` + on-device confirm still pending.
## 2026-06-17 — Personality: server-owned on the gateway + picker-command handling
**Why.** Two reports against the personality flow. (1) Sending `/personality` (no arg) showed an agent reply bubble that appeared then vanished; (2) `/personality none` returned a confirmation but the app never reflected that the overlay was cleared. Verified the actual contract against the up-to-date upstream clone: `/personality` is a *picker command* (`hermes_cli/commands.py` `_PICKER_COMMANDS`) that the desktop/TUI never raw-forward — a bare command expands to an arg step, and a named/`none` value is applied via `config.set {key:"personality"}`, which persists `display.personality` + applies `ephemeral_system_prompt` live to the session and emits `session.info`. The app instead blindly forwarded every slash to `slash.exec`/`command.dispatch`, had no `none` concept, and never consumed `session.info` — so it kept injecting a stale per-turn personality prompt that fought the server.
- **Slash results stopped vanishing.** `ChatHandler.loadMessageHistory` did a wholesale reload preserving only `voice-intent-`/`steer-`/`ask-` ids; `system-notice-` (every `addSystemNotice` slash result) was wiped by the next turn's reconcile. Added `system-notice-` to the preserve allow-list — fixes the disappearing bubble for `/personality` and all other inline command output.
- **Gateway client owns personality.** `GatewayChatClient` gained `serverPersonality: StateFlow<String?>`, `getPersonality()` (`config.get`), and `setPersonality()` (`config.set {key:"personality", value, session_id}`), plus a connection-level `session.info` interceptor that captures the `personality` field even with no turn in flight. `"none"`/`"default"`/`"neutral"` all clear the overlay (upstream `_validate_personality` conflates them).
- **ViewModel mirrors server truth.** `selectPersonality` pushes via `config.set` on the gateway (optimistic, rolled back on a server reject with a now-durable notice) and only drives per-turn injection on the SSE fallbacks; `startStream` skips the persona-prompt injection entirely on the gateway so it can't double-apply. A `startPersonalitySync` collector + a ready-socket `config.get` seed keep `_selectedPersonality` reconciled to whatever the server/desktop/TUI set.
- **Picker-command UX.** `/personality` is intercepted client-side like the desktop: bare `/personality` opens the agent sheet's Personality section (new `openPersonalityPicker` one-shot), `/personality <name|none>` routes to `selectPersonality`. The synthetic client **"Default" row was removed** — the picker is now **None** + the server-provided personalities (the configured default, if any, shows tagged `(default)` and highlights when active); upstream's active value is just `none` or a name, so the client shouldn't invent a third state. Added a `/personality none` palette entry. `AgentDisplay` treats `none`/`neutral` as cleared-overlay aliases (base identity, not the literal word) via `isClearedPersonality`.
- **`/model` sibling fixed.** `/model` is the other picker command (`_PICKER_COMMANDS = {model, skin, personality}`; `skin` is `cli_only` and already excluded on mobile). A bare `/model` had the same raw-forward dead-end — now it opens the model picker (`openModelPicker` one-shot → `ModelPickerSheet`), while `/model <args>` stays a real gateway switch.
- **Live model/provider sync.** The `session.info` interceptor now also surfaces `model` + `provider` (`serverModel`/`serverProvider` flows); the VM's `startGatewayStateSync` (renamed from `startPersonalitySync`) drives the model pill from them, so a `/model` switch on the desktop/TUI reflects live. Format-safe — the pill normalizes through `AgentDisplay.displayModelName`, and `session.info` keeps model/provider separate like `model.options`.
- **No app reload for server-supplied data.** `refreshPersonalities()` (list + default + active `config.get`) now fires on agent-sheet open alongside `refreshModelOptions`, so a personality added/changed server-side appears without restarting the app; the active value also tracks live via `session.info`.
- **Profile SOUL double-inject fixed.** On the gateway the session is bound to the selected profile (SOUL applied server-side) AND the personality rides `config.set` — so `startStream` now sends NO persona/profile prompt on the gateway (only the phone-status block), where it previously re-injected the profile's `systemMessage` on top of the server's own SOUL. SSE fallbacks keep the client-side precedence rules.
- **Verification.** `:app:testSideloadDebugUnitTest` compiles the full module clean; new `AgentDisplayTest` cases for `isClearedPersonality` / `none`-as-cleared pass. `./gradlew lint` still the pre-push gate. On-device confirm of the live gateway round-trip pending.
## 2026-06-16 — Per-surface release notes (plugin + CLI parity with Android)
**Why.** Plugin and CLI GitHub Release bodies were static boilerplate baked into the workflow YAML (version-interpolated, but change-agnostic — a reader couldn't tell what a `plugin-v*`/`cli-v*` release actually changed). Only Android had real per-release notes (`RELEASE_NOTES.md` via `body_path`). Brought plugin and CLI up to the same Summary/Added/Changed/Fixed format.
+13
View File
@@ -0,0 +1,13 @@
# GEMINI.md
Agent instructions for **Hermes-Relay**. This file exists so Gemini CLI (which
does not read `AGENTS.md` natively) picks up the project's guidance.
**Read [AGENTS.md](AGENTS.md) — it is the single source of truth** for every
coding agent: the entry point, the non-negotiables (standard-path-is-vanilla-
upstream, verify-endpoints, Conventional Commits + `main`/`dev` branching, the
per-language stack rules), and the public-repo writing hygiene. It links on to
`CLAUDE.md` for the deep reference (architecture, upstream Hermes API, repo
layout, code style, the dev loop, and the Key Files map).
Do not restate rules here — keep them in `AGENTS.md` so they can't drift.
+13 -11
View File
@@ -1,23 +1,25 @@
# Hermes-Relay-Plugin v__VERSION__
**Release Date:** June 16, 2026
**Since the previous plugin release:** Easier setup and a fixed dashboard panel — plus mid-conversation `/relay` controls and a relay-status widget.
**Release Date:** June 20, 2026
**Since the previous plugin release:** A new, removable **enhancement layer** that lets the relay teach the agent things only the relay knows — starting with sensitive-media classification — plus provider-aware enhanced voice and an isolated, TUI-tuned tmux for relay terminals.
This release makes the relay plugin easier to install and live with. Setup now prompts for the optional voice-provider keys instead of asking you to hand-edit `.env`, tools-only hosts can install through the native `hermes plugins install` path, and the installer no longer breaks on `uv`-managed Hermes cores. The dashboard panel — which previously rendered as blank boxes on the host's design system — now displays correctly, and a header widget plus `/relay` slash commands surface relay state from anywhere. The standard no-plugin path needs none of this.
This release adds a clean way for the relay to extend the agent without forking or touching the user's soul/memory. The first use is **sensitive-media classification**: the relay appends a small, auditable system-prompt block teaching the agent to mark private/NSFW media so the paired phone can blur it — with sensitivity staying model-emitted. It's on by default for relay installs (installing the relay is the opt-in), reversible from the dashboard or an env flag, fully visible over a new audit route, and a complete no-op on vanilla upstream. Voice gains provider-aware controls for Gemini and xAI, and relay terminals now run on a dedicated, correctly-configured tmux.
## What's changed
### Added
- **Guided env-key setup.** The plugin declares its optional voice-provider keys (`XAI_API_KEY`, `OPENAI_API_KEY`, `ELEVENLABS_API_KEY`) in its manifest, so `hermes plugins install` prompts for them (masked, with a "get yours" link) instead of requiring a hand-edited `.env`. The standard no-plugin path needs none.
- **Native install path.** Tools-only setups can install via `hermes plugins install Codename-11/hermes-relay/plugin`; the full relay still uses the curl `install.sh`.
- **`/relay` slash commands.** `relay status · devices · pair` are usable mid-conversation from any platform (CLI / Discord / TUI).
- **Dashboard relay-status widget.** A `Relay · connected / offline / unpaired` badge in the dashboard header, visible on every page.
- **Session-start relay health check.** A minimal, fully-guarded `on_session_start` hook records relay reachability without slowing the gateway.
- **Relay enhancement layer + agent-context injection.** A reusable, removable layer that injects auditable, fenced blocks into the agent's system prompt at plugin-load. Fail-open at every step (seam absent / block build throws ⇒ base prompt unchanged), config-gated, and a byte-for-byte no-op on vanilla upstream. Built to be retired per-surface as upstream adds a context hook — the same pattern as the bootstrap route shims. See `docs/plans/2026-06-20-relay-enhancement-layer.md`.
- **Sensitive-media classification (first block).** Teaches the agent to mark private/NSFW media with the client's spoiler convention so the phone blurs it per the user's setting. **On by default for relay installs**; opt out with `RELAY_AGENT_CONTEXT_ENABLED=0` or the dashboard toggle. Sensitivity stays model-emitted — no relay-side or on-device classifier. No soul/memory is touched.
- **`GET /context/injected` audit route.** The relay exposes exactly what it would inject (loopback-open, bearer-gated remotely), so the injection is never hidden — surfaced in the Android chat "What the agent sees" sheet as "Relay context (server-side)".
- **Dashboard Agent-context controls.** The Relay management tab gained a master toggle and per-block toggles (labeled experimental / server-side / removable), shown on-by-default for relay installs.
- **Provider-aware enhanced voice (Gemini + xAI).** `/voice/synthesize` accepts per-request overrides so a paired client can steer a Gemini voice/model with expressive tone tags, or an xAI voice with expressive speech tags, without changing the server's global voice config.
### Changed
- **Relay terminals run on an isolated, TUI-tuned tmux.** Sessions spawn on a dedicated tmux server/socket with a generated config — `escape-time 0`, truecolor `tmux-256color`, `mouse`/`focus-events` on, `status off` — so editors and full-screen tools behave correctly without touching the user's personal tmux.
### Fixed
- **Installer failed on uv-managed Hermes hosts.** `install.sh` assumed `pip` lived in the hermes-agent virtualenv, but environments created by `uv` (the upstream default) ship no `pip` module, so the editable install aborted at step 2. The installer now bootstraps `pip` via `ensurepip`, or falls back to `uv pip`, so the plugin installs cleanly on uv-managed cores.
- **Dashboard buttons rendered as blank boxes.** The host dashboard's Nous design-system `Button` / `Badge` use boolean variant flags (`outlined` / `ghost` / `invert`) and a `tone` prop — not the shadcn-style `variant` prop the plugin passed — so every button collapsed to a solid near-white fill with an invisible label. The plugin now translates its props to the design-system contract via an adapter and drops a label-hiding CSS reset.
- **Unreadable button labels.** Solid buttons in the relay dashboard panel inherited the container text colour, which matched their background; solid button variants now keep their proper contrast colour.
- **Relay voice synthesis no longer leaves temporary audio files behind** on the server.
- **Clearer voice errors.** Standard voice rejects an over-long recording before uploading and returns a helpful message for audio the server can't read, instead of a generic HTTP error.
## Install
+14 -10
View File
@@ -38,6 +38,10 @@ Hermes-Relay puts your [Hermes agent](https://github.com/NousResearch/hermes-age
A vanilla [hermes-agent](https://github.com/NousResearch/hermes-agent) install is enough — chat, management, and voice need **no plugin**. Add the optional relay only when you want terminal, phone control, or the CLI's tools. **Pair once from either surface; both work.**
<p align="center">
<img src="docs/diagrams/architecture-homepage.png" alt="How Hermes-Relay connects — Vanilla Hermes (Chat, Manage, Voice) runs with no plugin; the optional Relay plugin adds Terminal, Bridge, relay voice and desktop tools to the app and CLI; Device Control needs the sideload build." width="900">
</p>
## Quick Start (Android)
Install → connect → talk, in about two minutes.
@@ -51,7 +55,7 @@ Sideload builds check GitHub for updates and show a one-tap banner when you're b
### 2 · Have Hermes running
The app needs your Hermes **API server enabled and reachable from your phone**, plus an **API key** — the token the app sends to authenticate Chat (pick any value you like). Installing Hermes and choosing a provider is standard Hermes setup; the [full walkthrough](https://codename-11.github.io/hermes-relay/guide/getting-started) covers Windows, the dashboard for **Manage**, LAN scan, and QR setup.
The app needs your Hermes **API server enabled and reachable from your phone**, plus an **API key** — the token the app sends to authenticate Chat (pick any value you like). Installing Hermes and choosing a provider is vanilla Hermes setup; the [full walkthrough](https://codename-11.github.io/hermes-relay/guide/getting-started) covers Windows, the dashboard for **Manage**, LAN scan, and QR setup.
```bash
hermes setup --portal # install / log in / pick a provider — skip if already done
@@ -78,8 +82,8 @@ hermes gateway
Open the app and pick how to connect — any of:
- **Standard Hermes** → tap **Scan for Hermes on LAN** to auto-find the server, then enter your key.
- **Standard Hermes** → type the address (`http://<host>:8642`) and key by hand.
- **Vanilla Hermes** → tap **Scan for Hermes on LAN** to auto-find the server, then enter your key.
- **Vanilla Hermes** → type the address (`http://<host>:8642`) and key by hand.
- **Scan setup QR** → ask your Hermes agent to generate a QR with your URL + key (e.g. `{"api_url":"http://<host>:8642","api_key":"<key>","dashboard_url":"http://<host>:9119"}`) and scan it. `dashboard_url` is optional when the dashboard uses the conventional same-host `:9119` URL.
The wizard probes everything and finishes with a capability card:
@@ -92,7 +96,7 @@ The wizard probes everything and finishes with a capability card:
| **Remote** | Fallback route configured — keeps working away from home |
| **Relay** | Optional power tools — fine to leave unpaired |
If your dashboard requires sign-in, do it once under the **Manage** tab — the same session unlocks voice. That's the whole standard setup.
If your dashboard requires sign-in, do it once under the **Manage** tab — the same session unlocks voice. That's the whole Vanilla Hermes setup.
> **Going places?** Put your server's Tailscale URL in the setup form's *Remote access* field (or add a route any time under **Settings → Connections → Routes**). The app uses LAN at home and switches routes automatically when you leave. See [Remote access](https://codename-11.github.io/hermes-relay/guide/remote-access).
@@ -140,10 +144,10 @@ Full server setup, TLS, and systemd details: [docs/relay-server.md](docs/relay-s
<td align="center" width="25%"><img src="assets/screenshots/04_sessions.png" alt="Session history" width="100%"><br><sub><b>Session history</b></sub></td>
</tr>
<tr>
<td align="center" width="25%"><img src="assets/screenshots/05_commands.png" alt="Command palette" width="100%"><br><sub><b>Command palette</b></sub></td>
<td align="center" width="25%"><img src="assets/screenshots/05_themes.png" alt="App themes" width="100%"><br><sub><b>App themes</b></sub></td>
<td align="center" width="25%"><img src="assets/screenshots/06_manage.png" alt="Manage your agent" width="100%"><br><sub><b>Manage your agent</b></sub></td>
<td align="center" width="25%"><img src="assets/screenshots/07_connections.png" alt="Connections and routes" width="100%"><br><sub><b>Connections &amp; routes</b></sub></td>
<td align="center" width="25%"><img src="assets/screenshots/08_settings.png" alt="Settings" width="100%"><br><sub><b>Settings</b></sub></td>
<td align="center" width="25%"><img src="assets/screenshots/08_appearance.png" alt="Agent avatar &amp; skins" width="100%"><br><sub><b>Avatars &amp; skins</b></sub></td>
</tr>
</table>
@@ -153,7 +157,7 @@ Full server setup, TLS, and systemd details: [docs/relay-server.md](docs/relay-s
### Android
- **Streaming chat** — rides standard Hermes, preferring the dashboard gateway (`/api/ws`, live thinking) when signed in to Manage and falling back to API-server SSE otherwise, with live markdown, tool-call cards, session history, a searchable command palette, file attachments, quote-in-reply, conversation share, and send-while-streaming queuing.
- **Streaming chat** — rides vanilla Hermes, preferring the dashboard gateway (`/api/ws`, live thinking) when signed in to Manage and falling back to API-server SSE otherwise, with live markdown, tool-call cards, session history, a searchable command palette, file attachments, quote-in-reply, conversation share, and send-while-streaming queuing.
- **Manage your agent** — the full Hermes dashboard, native: switch models from your provider catalog, manage keys (write-only, masked, rate-limited reveal), create and edit profiles including `SOUL.md`, and browse/install/update skills. One dashboard sign-in covers it all.
- **Hands-free voice** — talk on a vanilla install: speech rides your server's configured providers, unlocked by the same Manage sign-in. Relay-paired setups add per-profile voice and an opt-in provider-native Realtime Agent with background task handoff.
- **Works away from home** — add a Tailscale or public URL and the app roams automatically (LAN at home, fallback elsewhere). An unreachable server gets a diagnosis, not just a red dot.
@@ -189,14 +193,14 @@ It pairs against the **same relay and credential store** as the Android app —
## How It Works
```
Phone (HTTP/WSS) --> Hermes Dashboard (:9119) [chat gateway, manage, standard voice]
Phone (HTTP/WSS) --> Hermes Dashboard (:9119) [chat gateway, manage, vanilla voice]
Phone (HTTP/SSE) --> Hermes API Server (:8642) [chat fallback, sessions, runs]
Phone (WSS/HTTP) --> Relay (:8767) [terminal, bridge, media, relay voice, sessions]
CLI (WSS) --> Relay (:8767) [machine tools, tui, terminal]
```
Chat prefers the Hermes dashboard gateway when Manage auth is ready, then falls
back to the upstream API server SSE path with the API key. Manage and standard
back to the upstream API server SSE path with the API key. Manage and Vanilla Hermes
voice ride the Hermes dashboard with its own one-time sign-in, so a vanilla
install needs no plugin for either. The optional relay on `:8767` adds the power
surfaces: terminal, bridge phone control, media handoff, machine tools, and
@@ -232,7 +236,7 @@ Read the canonical setup recipe before acting:
Then guide me through:
- Verifying hermes-agent is already installed (it's a prerequisite — Hermes-Relay is a plugin, not standalone)
- Running the server-plugin install one-liner: `curl -fsSL https://raw.githubusercontent.com/Codename-11/hermes-relay/main/install.sh | bash`
- Connecting my phone by Standard Hermes API URL/key first, then optionally pairing Relay via `hermes pair` or `/hermes-relay-pair` for power tools; OR pairing my laptop via the Hermes-Relay CLI (`irm https://raw.githubusercontent.com/Codename-11/hermes-relay/main/desktop/scripts/install.ps1 | iex` on Windows, then `hermes-relay pair --remote ws://<host>:8767`)
- Connecting my phone by Vanilla Hermes API URL/key first, then optionally pairing Relay via `hermes pair` or `/hermes-relay-pair` for power tools; OR pairing my laptop via the Hermes-Relay CLI (`irm https://raw.githubusercontent.com/Codename-11/hermes-relay/main/desktop/scripts/install.ps1 | iex` on Windows, then `hermes-relay pair --remote ws://<host>:8767`)
- Verifying with `hermes-status` (server) or `hermes-relay doctor` (CLI)
Always confirm before running shell commands. Never restart hermes-gateway without asking. If any step fails, consult the Troubleshooting section in the SKILL.md and ask me for the exact error.
+13
View File
@@ -405,6 +405,13 @@ the new app version and a higher `appVersionCode`.
shown in the settings/about screen. Update with the version number
and a brief feature summary. Gets stale silently if forgotten
(v0.4.0 shipped with 0.1.0 content until caught post-release).
- `app/src/googlePlay/play/release-notes/en-US/default.txt` — the Play
Console **"What's new"** text, which gradle-play-publisher reads at
upload to fill the Production-draft release notes. This is **separate**
from `RELEASE_NOTES.md` (that one is only the GitHub Release body) — if
this file is missing or stale, the Play draft ships with empty/wrong
notes (shipped empty in v1.1.0 until caught post-release). Keep it
**≤500 chars per language**, user-facing, Android-only.
- `docs/play-store-listing.md` — Play Store listing copy. Update
the version reference and the "Release Notes" section that gets
pasted into the Play Console "What's new" field. Keep the Play
@@ -533,6 +540,12 @@ Release named `Hermes-Relay-Plugin v<version>` for the plugin package.
> Production **draft** — skip to the Play Console, confirm the draft, and click
> **Start rollout**. The manual path below is the fallback when the secret is
> unset (or for staging on a non-production track).
>
> This automated tag path is intentionally bundle-only. It uploads the
> `googlePlayRelease` AAB and release-scoped "What's new" notes, but it does
> not republish static listing assets such as screenshots, title, description,
> icon, or feature graphic. Use the Play Store Listing workflow when those
> assets change.
**Pick the track first.** The AAB is track-agnostic — the same
`-googlePlay-release.aab` goes to whichever track you publish on. Choose by intent,
+36 -25
View File
@@ -1,22 +1,22 @@
# Hermes-Relay-Android v1.1.0
# Hermes-Relay-Android v1.2.0
**Release Date:** June 16, 2026
**Since v1.0.0:** A settings + chat-UX overhaul — quieter status surfaces, a single state-aware plugin badge, and chat-settings polish — plus a force-close fix and release-pipeline upgrades.
**Release Date:** June 20, 2026
**Since v1.1.0:** A big personalization release — app themes, swappable sphere skins, and animated agent **pets** — paired with a transparency pass (see which transport you're on and exactly what the agent is told), a much faster cold start, in-app crash reporting, and a broad reliability sweep.
v1.1.0 is a refinement release on top of the 1.0 milestone. Settings is calmer and easier to read: status pills now appear only when a surface needs attention, the Power tools section shows one **Plugin active / required / offline** badge instead of an identical chip on every card, and the most-used controls sit where you reach for them. Chat settings render correctly, the system-prompt preview reflects your toggles, and a crash that could hit right after a successful pair is gone.
v1.2.0 is about making Hermes-Relay feel like *yours* and making it honest about what it's doing. Dress the app in one of eight themes, swap the agent orb for a hand-picked or AI-generated **pet** that reacts to what the agent is doing, and give each profile its own icon. At the same time, the chat status strip now names the actual streaming path (⚡ Gateway, 📡 Sessions, …), a "What the agent sees" sheet shows the exact extra context prepended to your next turn, and cold start is roughly three times faster. If something does go wrong, the app now catches the crash and offers a one-tap, pre-filled bug report.
---
## Download
v1.1.0 ships in two Android build flavors. APK and AAB filenames are version-tagged:
v1.2.0 ships in two Android build flavors. APK and AAB filenames are version-tagged:
| Flavor | File | Who it's for |
|---|---|---|
| Google Play | `hermes-relay-1.1.0-googlePlay-release.aab` | Upload this Android App Bundle to Play Console. It has no AccessibilityService, screen reading, screenshots, gestures, SMS/calls, contacts/location, overlays, or unattended phone control. |
| sideload | `hermes-relay-1.1.0-sideload-release.apk` | Direct-install APK for full Device Control. Installs as `com.axiomlabs.hermesrelay.sideload`. |
| googlePlay APK | `hermes-relay-1.1.0-googlePlay-release.apk` | Parity/testing artifact. |
| sideload AAB | `hermes-relay-1.1.0-sideload-release.aab` | Parity/testing artifact. |
| Google Play | `hermes-relay-1.2.0-googlePlay-release.aab` | Upload this Android App Bundle to Play Console. It has no AccessibilityService, screen reading, screenshots, gestures, SMS/calls, contacts/location, overlays, or unattended phone control. |
| sideload | `hermes-relay-1.2.0-sideload-release.apk` | Direct-install APK for full Device Control. Installs as `com.axiomlabs.hermesrelay.sideload`. |
| googlePlay APK | `hermes-relay-1.2.0-googlePlay-release.apk` | Parity/testing artifact. |
| sideload AAB | `hermes-relay-1.2.0-sideload-release.aab` | Parity/testing artifact. |
Verify integrity with `SHA256SUMS.txt` from the same release. See the [Sideload guide](https://codename-11.github.io/hermes-relay/guide/getting-started.html#sideload-apk) for APK install steps.
@@ -24,32 +24,43 @@ Verify integrity with `SHA256SUMS.txt` from the same release. See the [Sideload
## Highlights
### Settings screen overhaul
### Make it yours
Settings was reorganized around what you actually touch and quieted down everywhere else:
- **App themes.** A theme picker in Settings → Appearance ships eight looks — the signature Hermes Relay brand (full light/dark) plus ports of the Nous Hermes baselines: Hermes Teal, Nous Blue, Midnight, Ember, Mono, Cyberpunk, and Rosé. The whole app follows your choice; Light/Dark/Auto applies to themes that ship both modes.
- **Agent pets — a living avatar.** Replace the orb with an animated pet that reacts to the agent: idle / thinking / writing / speaking / listening, a distinct **working** pose during tool calls, one-shot **greet** and **celebrate** reactions, and a loop that speeds up as output streams. Add or remove pets right in Appearance (no `adb`), preview each state, tune playback speed, and toggle frame auto-stabilization. Pets are pure data — an AI authoring kit and JSON schema let you generate one from sprite art.
- **Hot-swappable sphere skins + per-profile icons.** Keep the orb but reskin it (Adaptive, Classic, Aurora, Solar, Mono, or your own JSON skin), and give each agent profile its own small icon beside its name — all client-side, never sent to Hermes.
- **Exception-only status pills.** Status pills now appear only when a surface needs attention and stay quiet when everything is healthy — no more a wall of green chips to read past.
- **One state-aware plugin badge.** The Power tools section shows a single **Plugin active / required / offline** badge instead of an identical "Relay paired" chip repeated on every card.
- **Layout that follows your reach.** Connections moved to the top (above the Hermes section), and Diagnostics + Developer options moved into the App section.
- **Restyled to match the app.** The status chips now use the app's translucent-bordered language, and the brand blue was deepened.
### See what's actually happening
### Chat settings polish
- **Transport path is visible.** The chat status strip now shows which streaming path is in use — ⚡ Gateway (live thinking), 📡 Sessions, Completions, or Runs — and Chat Settings adds a basic→best tier ladder explaining the active path and its fallback.
- **"What the agent sees" sheet.** Tap the context meter to see the exact extra context prepended to your next turn — persona/profile, phone status, any per-turn voice hint, and (when paired) the relay's own server-side context. The audit is honest about what the phone sends versus what's applied on the server.
- **Spoken-turn badges + voice render-path visibility.** Voice and Realtime Agent replies carry a chip in the scrollback, and Voice Settings shows whether speech is rendering over the streaming or basic path.
- **Streaming-endpoint picker fixed.** The picker no longer wraps "Gateway" / "Sessions" onto a second line.
- **Live system-prompt preview.** The system-prompt preview now reflects the context toggles you've enabled (foreground app, battery, safety rails) with representative placeholder values, instead of looking inert.
### Privacy
### Force-close fix
- **Sensitive-media blur.** When paired to the relay, the agent can mark private/NSFW media and the phone blurs it per your setting — sensitivity stays model-emitted (no on-device or relay-side classifier), and the exact instruction is visible in the "What the agent sees" sheet. Vanilla Hermes (no plugin) is unaffected.
A corrupt encrypted token store — which can happen after an app upgrade or a device restore — used to throw during construction and crash the app right after a successful pair, on both standard and relay connections. The token store now heals a corrupt keyset in place, and credential storage degrades to a re-pair instead of crashing if the device keystore is unusable.
### Faster, calmer, more honest
### Release pipeline
- **~3× faster cold start.** The app was building several hardware-keystore-encrypted stores at launch, serializing on a process-global lock and stalling the chat header for seconds. It now builds a single keyset shared with the dashboard cookies, cutting measured time-to-connected from ~2.9 s to ~1 s. Existing sign-ins migrate automatically.
- **Honest loading, never stale.** Model, personality, and approvals show a brief "checking…" state and fade in once the server confirms them; standard controls (Model, YOLO, Fast, reasoning effort) always appear — live when ready, "checking…" while loading, or cleanly disabled with the reason — instead of being hidden or showing a maybe-wrong value.
- **Automated Play Console upload.** When a `PLAY_SERVICE_ACCOUNT_JSON` secret is configured, pushing a stable `android-v*` tag uploads the `googlePlay` App Bundle to the Production track as a draft (a human still starts the rollout). Prereleases are skipped, and the `sideload` flavor is structurally blocked from ever publishing to Play. Without the secret, releases publish to GitHub Releases exactly as before.
- **Desktop UI preview harness (`:ui-preview`).** A non-shipped Compose for Desktop module renders presentational composables in a window on the PC with Compose Hot Reload, for fast UI iteration without a device build/install loop. It reuses the shared sphere algorithm as its single source of truth.
### More reliable
- **In-app crash reporting.** A force-close now surfaces a clean dialog on next launch with the stack trace — Copy it, or **Report** to open a pre-filled GitHub issue from the bug template. The handler re-raises so Play vitals still record the crash.
- **QR pairing hardened for foldables.** On devices where the camera can't initialize, the scanner shows a "camera unavailable — pair manually" card instead of force-closing.
- **Crash fixes.** No more crash opening a chat with a server-local image, and the PDF viewer no longer crashes when a document closes mid-render.
- **Chat correctness.** In-chat model picks now actually apply — both on a new chat and mid-conversation — server-side turn errors stay on screen as an error bubble, per-reply token counts and provenance badges survive the post-turn reload, and server steering markers (`[System: …]`) no longer appear as chat bubbles.
### Voice & terminal polish
- **Enhanced voice control (Gemini & xAI).** When the relay uses a Gemini or xAI voice provider, Voice Settings can pick a voice/model and turn on expressive tone/speech tags. Vanilla Hermes voice stays configured server-side.
- **Leaner terminal.** A scrollable, fully-legible key bar, TUI-correct arrows and bracketed paste, a compact single-row header, and relay sessions on an isolated, TUI-tuned tmux so editors and full-screen tools behave.
---
## Upgrade notes
- The force-close fix means devices that previously crashed on connect after an upgrade or restore will heal their token store automatically on first launch of this build — no manual re-pair required in most cases.
- `appVersionCode` is **13**.
- Cold-start speedup migrates the encrypted credential and dashboard-cookie stores automatically on first launch; in rare cases Manage/voice may ask for a one-time re-login (cookies are re-obtainable).
- App themes, sphere skins, and pets are available on **both** flavors — they're client-side and need no Device Control.
- `appVersionCode` is **14**.
+156 -40
View File
@@ -6,60 +6,118 @@ For shipped work, see `DEVLOG.md`. For architectural decisions, see `docs/decisi
---
## User-Added:
- [ ] Enhance the 'clean chat' view mode to allow more a little more vertical visible text area and scrolling within.
- [ ] Look into the voice-settings profile specific capabilities - confirm approach is sound - verify as I noticed that in 'auto' mode it didn't work, it still used the system default despite config despite override voice chosen being displayed to user in voice config in voice setting in app UI. Only switching to 'Relay' specifically allowed the user-override to work/apply.
- [ ] - analytics and diagnostics pages need cleaned up, improved, enhancements for UI/UX/layout. Diagnostics should have timeline vertical status checks with failure reason etc
- [x] **Per-profile agent icon + static-image avatar (shipped 2026-06-20 — `d827e46`, see DEVLOG).** Per-profile icon: client-side `ProfileIconStore` (per `(connection, profile)`, never sent to Hermes; stores a copied-file path) → small Coil image beside the agent name in `MessageBubble` via `LocalAgentIconPath`; picker is `AgentIconRow` under the local-name row in `ConnectionInfoSheet`. Static image: "Add a pet" accepts a single image (magic-byte detect → one-frame static pet). Scope shipped: small name-adjacent icon only; big avatar stays global. Follow-ups: on-device smoke (import an image as a pet; set a profile icon, confirm it shows by the name + persists across restart); optionally also show the icon in the profile picker.
## Hands-free agentic voice backlog
Goal: make Hermes usable for hands-free work without leaving the operator blind
to tool state, safety prompts, or the current task.
- **Waveform output-start sync** — current input waveform timing feels good, but
the agent-output waveform can unfold and begin movement before audible speech
starts. Split "preparing audio" from "speaking audio" in the visual layer, or
gate the unfolded Speaking waveform on the first real playback frame/audio
amplitude. Processing can stay as the folded circular spinner until output is
actually audible.
the agent-output waveform can unfold and begin movement before audible speech
starts. Split "preparing audio" from "speaking audio" in the visual layer, or
gate the unfolded Speaking waveform on the first real playback frame/audio
amplitude. Processing can stay as the folded circular spinner until output is
actually audible.
- **Voice command layer** — reserve local commands that bypass normal agent
routing: "pause", "resume", "stop talking", "cancel", "repeat that", "open
overlay", "return to Hermes", and "new chat". These should work while the
agent is thinking, speaking, or using tools.
routing: "pause", "resume", "stop talking", "cancel", "repeat that", "open
overlay", "return to Hermes", and "new chat". These should work while the
agent is thinking, speaking, or using tools.
- **Spoken tool progress** — when Hermes uses tools, voice mode should speak
short status updates such as "I'm checking the relay logs" or "I found an
error" without waiting for final assistant text. Long tool calls should emit
periodic, low-noise progress updates.
short status updates such as "I'm checking the relay logs" or "I found an
error" without waiting for final assistant text. Long tool calls should emit
periodic, low-noise progress updates.
- **Realtime tool timeline parity** — the voice overlay should render the same
live thinking blocks, streaming assistant text, and tool call progress as the
normal chat surface without requiring exit/reload.
live thinking blocks, streaming assistant text, and tool call progress as the
normal chat surface without requiring exit/reload.
- **Hands-free confirmation flow** — risky actions need first-class spoken and
visual confirmation: "yes", "no", "cancel", "confirm", plus a visible and
audible countdown for destructive actions.
visual confirmation: "yes", "no", "cancel", "confirm", plus a visible and
audible countdown for destructive actions.
- **Voice session memory/status** — add a compact "where are we?" summary for
the current voice task: active objective, last tool result, pending next step,
and whether the agent is waiting on the user.
the current voice task: active objective, last tool result, pending next step,
and whether the agent is waiting on the user.
- **Mode presets** — add presets such as Hands-free, Low latency, Careful tool
mode, and Quiet/visual-only. Hands-free should favor Continuous listening,
spoken tool progress, confirmations, and overlay availability.
mode, and Quiet/visual-only. Hands-free should favor Continuous listening,
spoken tool progress, confirmations, and overlay availability.
- **Barge-in hardening** — keep barge-in experimental until echo/self-recording
is solved. The target path is proper AEC, playback-ducking, and a rule that
output audio can never become a user turn.
is solved. The target path is proper AEC, playback-ducking, and a rule that
output audio can never become a user turn.
- **Audio quality guardrails** — normalize output volume across realtime and
fallback TTS providers, keep pronunciation hints/profile voice tuning, and
measure provider-specific delay, chunk gaps, and tail clipping.
fallback TTS providers, keep pronunciation hints/profile voice tuning, and
measure provider-specific delay, chunk gaps, and tail clipping.
- **Pluggable Realtime Agent media transports** — add an OpenAI-first WebRTC
transport option for Realtime Agent so mobile audio can use provider-native
jitter buffering, interruption, and media handling instead of only relay
WebSocket PCM. Design this as a provider transport interface
(`websocket`, `webrtc`, future `livekit`/SIP-style bridges) so other
realtime providers can opt in without forking the Hermes broker/tool
contract. Hermes must still own tools, memory, confirmations, current data,
and durable transcript state.
transport option for Realtime Agent so mobile audio can use provider-native
jitter buffering, interruption, and media handling instead of only relay
WebSocket PCM. Design this as a provider transport interface
(`websocket`, `webrtc`, future `livekit`/SIP-style bridges) so other
realtime providers can opt in without forking the Hermes broker/tool
contract. Hermes must still own tools, memory, confirmations, current data,
and durable transcript state.
- **Voice engine selector** — implemented as an opt-in experimental Realtime
Agent engine in `docs/plans/2026-05-19-realtime-hermes-voice-agent.md`.
Follow-up work is provider-native turn-taking, richer confirmation handling,
and quality/latency evaluation before promotion beyond Experimental.
Agent engine in `docs/plans/2026-05-19-realtime-hermes-voice-agent.md`.
Follow-up work is provider-native turn-taking, richer confirmation handling,
and quality/latency evaluation before promotion beyond Experimental.
- **Realtime-native Hermes bridge prototype** — first relay-brokered slice
implemented in `docs/plans/2026-05-19-realtime-hermes-voice-agent.md`.
Remaining work: let OpenAI/xAI realtime sessions own more of the live speech
turn while still proxying every tool, confirmation, memory, and Android bridge
action through Hermes/relay safety.
implemented in `docs/plans/2026-05-19-realtime-hermes-voice-agent.md`.
Remaining work: let OpenAI/xAI realtime sessions own more of the live speech
turn while still proxying every tool, confirmation, memory, and Android bridge
action through Hermes/relay safety.
---
@@ -77,7 +135,7 @@ Things to look into:
- **Skill distribution as separate from plugin distribution** — right now skills ride along with the plugin install via `external_dirs`. Should skills be installable independently (e.g. `hermes skill install <git-url>`)? Would that fragment maintenance or improve reuse?
- **Tool registration discoverability** — `android_*` tools register at gateway import time. There's no canonical "list installed plugin tools" API. Would adding one to upstream make sense, or is `gateway tool list` already enough?
- **Versioning + compatibility ranges** — `pip install -e` doesn't enforce version pins between hermes-agent and our plugin. A breaking change in upstream's plugin loader could silently break us. Do we need a `hermes_compat: ">=0.8.0,<1.0.0"` field somewhere?
- **`hermes-relay-self-setup` SKILL.md as a precedent** — we just shipped a self-installing skill that an LLM can fetch from a raw GitHub URL and execute. Does this pattern generalize? Could it become a recommended way for any third-party Hermes project to ship setup automation?
- `**hermes-relay-self-setup` SKILL.md as a precedent** — we just shipped a self-installing skill that an LLM can fetch from a raw GitHub URL and execute. Does this pattern generalize? Could it become a recommended way for any third-party Hermes project to ship setup automation?
- **Bootstrap injection** — `hermes_relay_bootstrap/` monkey-patches `aiohttp.web.Application` to inject endpoints into vanilla upstream. This is intentional but feels like a hack. Upstream PR #8556 (`feat/session-api`) will eventually let us delete it — verified 2026-04-15 that its scope covers the full bootstrap surface (sessions, memory, skills, config, available-models). Track that PR's status periodically.
- **Gateway slash-command preprocessor — upstream Stage 1 PR.** Sibling follow-up to #8556. Intercepts known gateway commands on `/v1/runs` + `/v1/chat/completions`, dispatches the stateless ones (`/help`, `/commands`) via `gateway_help_lines()`, returns a deterministic "use a channel with session state" notice for the stateful majority. Currently being prepared in `C:/Users/Bailey/Desktop/Open-Projects/hermes-agent-pr-prep/` on branch `feat/api-server-gateway-commands`; awaiting subagent's code + draft PR body before pushing. See `docs/upstream-contributions.md` §5.
- **Gateway slash-command preprocessor — bootstrap middleware (Stage 1 equivalent).** Sibling shim in `hermes_relay_bootstrap/_command_middleware.py` that mirrors the upstream Stage 1 PR as an aiohttp middleware injected at bootstrap time. Ships the hallucination fix to vanilla-upstream installs before the upstream PR lands. Planned for v0.4.1, after the current bridge feature branch wraps. See `ROADMAP.md` v0.4.1 entry.
@@ -94,6 +152,64 @@ When the answer becomes clearer, this section becomes either an ADR in `docs/dec
- **Wave 3 voice-bridge multi-turn confirmation** — currently a 5s TTS countdown with cancel; conversational confirmation is the follow-up
- **LLM client wiring for `android_navigate`** — `_default_vision_model` is stubbed; production swap to a real Anthropic/OpenAI vision client
- **Real screenshots of each flavor's a11y permission dialog** — for `user-docs/guide/release-tracks.md`
- **`llms.txt` standard** — explicitly skipped in favor of the `hermes-relay-self-setup` SKILL.md path; revisit if the standard gains traction in the agent ecosystem
- **`markdown-renderer` 0.40.x API update** — pinned at `0.30.0` in `gradle/libs.versions.toml` because 0.40.2 introduced breaking API changes that `app/src/main/kotlin/com/hermesandroid/relay/ui/components/MarkdownContent.kt` hasn't been updated for. Specifically: `markdownColor()` drops `codeText`/`linkText`, `MarkdownCodeBlock`/`MarkdownCodeFence` inner lambdas now take a 3rd `TextStyle` arg, and `MarkdownHighlightedCode`'s 3rd param is now `TextStyle` instead of `Highlights.Builder`. Dependabot auto-merged the bump on 2026-04-13 which silently broke CI; reverted for the v0.3.0 release. Update requires reading the new library API docs and testing in Studio — not a blind fix. Consider adding a dependabot ignore rule for `markdown-renderer` major bumps until this is handled.
- `**llms.txt` standard** — explicitly skipped in favor of the `hermes-relay-self-setup` SKILL.md path; revisit if the standard gains traction in the agent ecosystem
- `**markdown-renderer` 0.40.x API update** — pinned at `0.30.0` in `gradle/libs.versions.toml` because 0.40.2 introduced breaking API changes that `app/src/main/kotlin/com/hermesandroid/relay/ui/components/MarkdownContent.kt` hasn't been updated for. Specifically: `markdownColor()` drops `codeText`/`linkText`, `MarkdownCodeBlock`/`MarkdownCodeFence` inner lambdas now take a 3rd `TextStyle` arg, and `MarkdownHighlightedCode`'s 3rd param is now `TextStyle` instead of `Highlights.Builder`. Dependabot auto-merged the bump on 2026-04-13 which silently broke CI; reverted for the v0.3.0 release. Update requires reading the new library API docs and testing in Studio — not a blind fix. Consider adding a dependabot ignore rule for `markdown-renderer` major bumps until this is handled.
- **Dependabot auto-merge guardrails** — Dependabot merged breaking bumps despite CI failing. Investigate why `.github/workflows/dependabot-auto-merge.yml` isn't gating on CI status, and consider adding an ignore rule for packages we know need manual attention on major bumps (`markdown-renderer`, compose BOM, activity-compose).
---
## Crash reporting + foldable hardening (shipped 2026-06-20)
Triggered by a Play Store review: app "keeps crashing" during setup on a Samsung Galaxy Z Fold7 (Android 16 / SDK 36, version code 13). Shipped: in-app crash capture (`util/CrashReporter.kt` — uncaught handler that persists a report then re-raises so Play vitals still collects; `ui/components/CrashReportDialog.kt` — show-once dialog with Copy + pre-filled GitHub-issue "Report"); QR camera-init hardening (`QrPairingScanner.kt` — try/catch around `ProcessCameraProvider.get()` and `InputImage.fromMediaImage()`, graceful `CameraUnavailableCard` → manual pairing instead of force-close).
Follow-ups:
- **Confirm the actual crash from Play vitals.** Pull the top crash cluster for Galaxy Z Fold7 / version code 13 (Quality → Android vitals → Crashes & ANRs) to verify the camera path is the real cause vs. another setup-path throw. The hardening is correct regardless, but the trace closes the loop.
- **Portrait lock is moot on large screens under SDK 36.** `android:screenOrientation="portrait"` is largely ignored by Android 16's mandatory large-screen orientation override on foldables/tablets. Decide whether to keep the lock (it still applies on phones) or make it conditional; either way it does not *cause* the crash.
- **Foldable camera lifecycle races (from the 2026-06-20 audit, not yet fixed).** `QrPairingScanner` can still hit bind/unbind races on rapid fold/unfold recomposition (the `DisposableEffect` `unbindAll()` vs. an in-flight `addListener` bind), and `mapBoxToViewport` runs on possibly-stale `viewportSizePx` during a fold transition. Not crash-fatal after the try/catch hardening (logged + skipped), but worth a fold-aware guard if foldable adoption grows.
- **Optional: surface crash history in Settings.** The reporter keeps only the most recent crash (`files/crash/last-crash.json`, consumed on view). If repeat-crash diagnosis becomes common, keep a small ring of recent reports + a Settings entry to view/copy them.
---
## Relay enhancement layer + agent-context injection (shipped 2026-06-20 — `docs/plans/2026-06-20-relay-enhancement-layer.md`)
Shipped: `plugin/enhancements/` (registry + fail-open `context_injection` wrap of `AIAgent._build_system_prompt`), the `media-sensitivity` block, `GET /context/injected` audit route, dashboard toggles, client sensitivity re-thread + "Relay context (server-side)" audit section, and the transport-path UI (`ChatTransportStatusBadge` / `RelayStatusStrip` + tier ladder). OFF by default, removable, vanilla-safe.
Follow-ups:
- **Confirm the `AIAgent` seam on the live host before relying on it.** `context_injection._resolve_ai_agent_class()` tries `agent.system_prompt` / `run_agent`. When you flip `RELAY_AGENT_CONTEXT_ENABLED=1`, verify `GET /context/injected` shows the block AND that it actually lands in the prompt (the wrap is fail-open, so a wrong module = inert, not broken). If the class lives elsewhere, widen the module list.
- **Retire the monkey-patch when upstream adds a plugin context hook.** Drop `context_injection` (and migrate to the native hook) the moment hermes-agent ships a first-class system-prompt contributor — same as we retire bootstrap routes for native upstream routes.
- **Incremental bootstrap migration.** Fold the existing `hermes_relay_bootstrap` route-patches into `plugin/enhancements/` per-surface (startup phase) so patching is one surface; don't big-bang the working compat.
- **Structured media channel** — `docs/plans/2026-06-20-structured-media-channel.md` (design only). Replace fragile `MEDIA:`/markdown text markers with a structured channel carrying `sensitive` natively; lead with a relay `relay_send_media(path, sensitive, …)` tool.
- **Gateway voice-ephemeral via the same slot.** The enhancement layer's server-side injection can carry per-turn voice instructions on the gateway (which has no ephemeral `system_message`), letting voice stay on the gateway instead of being forced to SSE. Wire when the voice path is revisited.
---
## Attachments (shipped 2026-06-18 — `docs/plans/2026-06-18-attachment-experience.md`)
- **B3 — download progress + cancel.** Inbound fetch is un-cancelable; the previews work scaffolded an indeterminate bar + nullable `onCancel`. Live wiring needs the fetch-path owner (`ChatViewModel`/`Attachment`) to expose determinate progress (Content-Length) + a cancel hook.
- **A6 — multi-image gallery.** N images in one message → grid + swipe-across viewer (Telegram media-group parity).
- **C5 — agent-side sensitivity config gate.** `RELAY_MEDIA_SENSITIVITY_HINTS` (env or per-profile) instructing the agent to annotate sensitive media via the prompt-builder. Transport (relay `X-Media-Sensitive` header + client blur) already ships; the agent isn't asked to set the bit yet.
- **Relay thumbnails (D6).** Server-side thumbnail generation to avoid full-size download for cards/galleries. Needs an image lib (Pillow not currently a dep) — evaluate before adding.
- **D5 — outbound upload progress.** No per-attachment progress during the 60s gateway PDF-render window.
## Voice overhaul (shipped 2026-06-18 — `docs/plans/2026-06-18-voice-overhaul.md`)
- **Per-profile voice on Standard (upstream PR).** Upstream `/api/profiles/*` has no voice field and `/api/audio/*` is host-global. Long-term: PR a voice section to the profile config + make `/api/audio/*` honor the active/`?profile=` profile. The relay path already carries per-profile voice; ship that first.
- **Wire connectionId for per-profile voice namespacing.** `VoicePreferencesRepository` is scope-aware (`base_connId_profile`), but `RelayApp` passes only the profile *name* to `onProfileChanged`, so `connectionId` is null and keys namespace by profile-only. Wire `setVoicePrefsConnection` to `ConnectionViewModel.activeConnectionId` (in `RelayApp`) so two connections with same-named profiles don't share voice settings.
- **Realtime-PCM waveform output gating.** The basic-TTS output waveform is now Visualizer-accurate (gated on real playback amplitude), but the realtime path gates `outputAudioActive` on `audioSeen` (first decoded PCM bytes) in `VoiceViewModel.handleRealtimeVoiceEvent`, which can still lead audible output by the `RealtimePcmPlayer` start prebuffer. Gate realtime on actual playback-start (head moved) to match the basic-TTS path.
## Chat clean-mode + pets (shipped 2026-06-18 — `docs/plans/2026-06-18-chat-clean-mode-and-pets.md`)
- **Part-A chat polish (optional bundle).** Per-code-block copy + horizontal scroll, visible copy affordance, mid-stream stall feedback, profile/skill-aware empty-state chips, the ~40-flow recomposition hotspot at the top of `ChatScreen`. (Sphere `contentDescription`/reduced-motion was handled by the clean-mode a11y work.)
- **Pet hot-load + in-app add/remove (shipped 2026-06-20).** Pets now live-refresh: an `avatarsRefreshTick` keys the avatar `produceState` in `RelayApp`, and Appearance re-scans `pets/` on open and after in-app import/delete — no app restart. Appearance gained "Add a pet" (SAF `.zip` import via `PetImporter`, zip-slip/zip-bomb guarded + validated through `toAvatar`) and an "Installed pets" list with per-pet remove (`PetLoader.deletePet`, confirm dialog, Sphere fallback). Remaining:
- **Sphere-skin parity.** Skins are still process-scoped + `adb push` only — the live tick and the importer cover pets, not skins. Extend the tick to `loadUserSkins` and add a `.json` skin import if hot-loading/adding skins in-app is wanted.
- **`adb push` into `Android/data` hangs on Samsung scoped storage.** Confirmed: pushing a pet pack to `/sdcard/Android/data/<pkg>/files/pets/` stalls (no bytes written) although `adb shell ls` of the dir works. In-app `.zip` import is the supported path; `/sdcard/Download` pushes fine. Consider softening `docs/pet-spec.md` + user-docs to lead with in-app import over adb.
- **On-device import/delete smoke.** Import `/sdcard/Download/lucy.zip` via Add a pet → confirm Lucy appears, selects, and animates all states; then remove it and confirm the avatar falls back to the Sphere.
- **Pet state-change re-decode can flash one blank frame.** When the agent state switches clips, the first frame of the new clip may briefly be blank during decode; prewarm/hold-last-frame to smooth it. Root cause is the same as the next item: `PetAvatar.Render` re-decodes from disk on every clip change.
- **Pet frame-sequence memory: no cap or downsample (audit 2026-06-19).** `decodeClip` decodes every frame of the selected clip into `List<ImageBitmap>` at full resolution with no `inSampleSize` downscale to the display size and no frame-count/dimension ceiling — a long sequence of large PNGs can use a lot of RAM and a single very large image can OOM `BitmapFactory`. Add `inSampleSize` downsampling to the avatar's draw size and/or a documented hard cap. Spec now warns authors (prefer sprite sheets), but the renderer doesn't enforce it.
- **Pet decoded-clip cache (audit 2026-06-19).** `PetAvatar.Render` keys `produceState` on `clip`, so idle→thinking→speaking→idle within one turn re-runs `BitmapFactory.decodeFile` from disk each transition (repeated I/O + GC churn, and the blank-frame flash above). Add a small per-avatar `Map<SphereState, PetFrames>` decode cache.
- **Pet behavior model — richer state association (spec'd 2026-06-19, `docs/pet-spec.md` "Agent states & pet behavior").** Shipped: the honesty clamp (declared reactivity ∩ `PET_RENDERER_CAPABILITIES`), the friendly `writing` alias, the `**working`/tool-use overlay** (pet-local sub-state from `toolCallBurst`; opt-in `working` clip drives both the swap and the Tools badge), the **one-shot reaction layer** (`greet`/`wake` on appear, `done`/`celebrate` on turn-finish — opt-in, play-once-then-revert, transition-derived; `ONE_SHOT_MAX_MS` backstop), and `**intensity` modulation** (opt-in `reactive.intensity` → live playback speedup ≤1.6× via `rememberUpdatedState`; un-clamps the Activity badge). Voice · Tools · Activity reactivity is now complete. Remaining:
- `**attention` one-shot (only deferred behavior).** A reaction on notification arrival — needs a host event the avatar doesn't yet receive (unlike `greet`/`done`, which ride state transitions). Would plumb a notification edge into `AvatarRenderState` (or a side channel) + a `PetOneShot.Attention`. Low priority: the avatar is rarely on-screen when notifications land (backgrounded) — see the value analysis; revisit only if the avatar becomes an always-on surface (persistent overlay / Quest port).
- **On-device verification (working + one-shots + intensity).** Best seen in clean mode (`AgentTextFlow` feeds `toolCallBurst` + `streamingIntensity` + state transitions). Confirm: a `working` clip swaps in during a tool run and releases ~600ms after (`WORKING_BURST_THRESHOLD` 0.5); a `done` clip plays once on reply completion then returns to idle; a `greet` clip plays once when the avatar appears; with `intensity:true`, a writing/working loop visibly quickens while streaming. Watch for the known clip re-decode flash on each swap (separate TODO — decoded-clip cache).
- **Undecodable-but-present image appears valid (audit 2026-06-19).** A file that exists but isn't a decodable image passes the loader's `isFile` check, so the pet shows in the picker but renders blank. Documented as a caveat; consider a cheap header sniff at load time if false-valid pets become a support issue.
+26
View File
@@ -179,6 +179,10 @@ android {
// Robolectric (VoicePlayerTest) needs merged Android resources +
// manifest on the unit-test classpath to bootstrap its sandbox.
unitTests.isIncludeAndroidResources = true
// [POC] Roborazzi runs without its Gradle plugin (the plugin needs AGP's
// removed TestedExtension). Force record mode via the test-JVM system
// property the plugin would otherwise inject, so captureRoboImage writes.
unitTests.all { it.systemProperty("roborazzi.test.record", "true") }
}
}
@@ -198,6 +202,17 @@ kotlin {
jvmToolchain(17)
}
// [screenshots] Host-side screenshot tests render MessageBubble -> MarkdownContent,
// whose code-highlighter (dev.snipme.highlights) ships Java-21 bytecode. The build
// toolchain pins test execution to JDK 17, which can't load class-file v65, so run
// unit tests on a 21 JVM. Compile target stays 17; on-device (dexed) is unaffected.
// foojay (settings.gradle.kts) auto-provisions the 21 JDK if absent.
tasks.withType<Test>().configureEach {
javaLauncher.set(
javaToolchains.launcherFor { languageVersion.set(JavaLanguageVersion.of(21)) }
)
}
dependencies {
// Compose BOM
val composeBom = platform(libs.compose.bom)
@@ -283,8 +298,19 @@ dependencies {
// across priority groups against real local sockets so the behavior we
// validate matches on-device.
testImplementation(libs.okhttp.mockwebserver)
// Konsist — enforces the ADR 34 upstream/relay/shared package fence as a JUnit test
testImplementation(libs.konsist)
androidTestImplementation(libs.compose.ui.test.junit4)
debugImplementation(libs.compose.ui.tooling)
debugImplementation(libs.compose.ui.test.manifest)
// [POC] Roborazzi host-side screenshot rendering (src/test, Robolectric).
// Renders real composables on the JVM at an exact canvas — no device, no
// status bar, no clipping. See StoreScreenshotTest.
testImplementation("io.github.takahirom.roborazzi:roborazzi:1.43.1")
testImplementation("io.github.takahirom.roborazzi:roborazzi-compose:1.43.1")
testImplementation(libs.compose.ui.test.junit4)
testImplementation(libs.compose.ui.test.manifest)
testImplementation("androidx.test.ext:junit:1.2.1")
}
@@ -126,7 +126,7 @@ class OnboardingFlowTest {
navigateToPage(4)
composeTestRule
.onNodeWithText("Standard Hermes")
.onNodeWithText("Vanilla Hermes")
.assertIsDisplayed()
}
@@ -135,7 +135,7 @@ class OnboardingFlowTest {
setOnboardingContent()
navigateToPage(4)
composeTestRule.onNodeWithText("Standard Hermes").performClick()
composeTestRule.onNodeWithText("Vanilla Hermes").performClick()
composeTestRule.waitForIdle()
composeTestRule
@@ -151,7 +151,7 @@ class OnboardingFlowTest {
setOnboardingContent()
navigateToPage(4)
composeTestRule.onNodeWithText("Standard Hermes").performClick()
composeTestRule.onNodeWithText("Vanilla Hermes").performClick()
composeTestRule.waitForIdle()
composeTestRule
@@ -172,6 +172,17 @@ class OnboardingFlowTest {
.assertIsDisplayed()
}
@Test
fun powerPage_linksToPermissionReview() {
setOnboardingContent()
navigateToPage(3)
composeTestRule
.onNodeWithText("Review permissions")
.assertIsDisplayed()
.assertIsEnabled()
}
@Test
fun skipButton_visibleOnIntroPages_andWizardSkipOnConnectPage() {
setOnboardingContent()
@@ -0,0 +1,45 @@
package com.hermesandroid.relay.ui.screens
import androidx.compose.ui.test.assertIsDisplayed
import androidx.compose.ui.test.junit4.createComposeRule
import androidx.compose.ui.test.onNodeWithText
import androidx.compose.ui.test.performScrollTo
import com.hermesandroid.relay.ui.theme.HermesRelayTheme
import org.junit.Rule
import org.junit.Test
class PermissionsStatusScreenTest {
@get:Rule
val composeTestRule = createComposeRule()
@Test
fun permissionsScreen_showsStandardAndOnDemandRows() {
composeTestRule.setContent {
HermesRelayTheme {
PermissionsStatusScreen(
onBack = {},
onOpenBridge = {},
)
}
}
composeTestRule
.onNodeWithText("Permissions and capabilities")
.assertIsDisplayed()
composeTestRule
.onNodeWithText("Chat and Manage")
.assertIsDisplayed()
composeTestRule
.onNodeWithText("No Android runtime permission needed. API/dashboard auth is configured separately.")
.assertIsDisplayed()
composeTestRule
.onNodeWithText("Camera")
.performScrollTo()
.assertIsDisplayed()
composeTestRule
.onNodeWithText("Microphone")
.performScrollTo()
.assertIsDisplayed()
}
}
@@ -1,8 +1,8 @@
package com.hermesandroid.relay.voice
import com.hermesandroid.relay.network.ChannelMultiplexer
import com.hermesandroid.relay.network.handlers.LocalDispatchResult
import com.hermesandroid.relay.network.models.Envelope
import com.hermesandroid.relay.network.relay.ChannelMultiplexer
import com.hermesandroid.relay.network.shared.LocalDispatchResult
import com.hermesandroid.relay.network.relay.models.Envelope
/**
* Local in-process bridge dispatcher type. The Play flavor never invokes
@@ -0,0 +1 @@
en-US
@@ -0,0 +1,62 @@
Hermes-Relay is the native Android client for the Hermes agent platform. Point it at your own Hermes instance and chat with your agent, talk to it hands-free, and manage models, keys, skills, and profiles from anywhere.
It is not a hosted AI service. It is a companion app for the Hermes agent you run, and it talks only to the instances you configure.
QUICK START
1. Run hermes-agent with its API server and dashboard enabled on your computer or home server.
2. Install Hermes-Relay and enter your server address, for example http://192.168.1.100:8642.
3. The setup wizard checks what your server supports and shows a readiness card, then you are ready to chat.
A plain Hermes install is enough. Chat, management, and voice work with no plugin or extra service.
HOW IT WORKS
Chat streams directly from your Hermes API Server or dashboard gateway in real time. Manage and voice use your Hermes dashboard with one sign-in. Run the optional relay service and the app can pair by QR code to add power tools: remote terminal, notification companion, media handoff, relay-session management, and additional voice engines.
GOOGLE PLAY BUILD
The Google Play build ships Hermes Bridge Core only. It has no AccessibilityService Device Control: it cannot read your screen, tap, type, swipe, screenshot, send SMS, place calls, or access contacts or location. Device Control is reserved for sideload builds distributed outside Google Play.
FEATURES
- Streaming Chat: real-time responses with reasoning, markdown, tool-call visibility, attachments, mid-turn steering, edit-and-resend, and a searchable command palette.
- Manage Your Agent: use your Hermes dashboard from your phone to switch models, manage provider keys, edit profiles, and browse, install, and update skills.
- Voice Mode: talk hands-free using your server's speech providers. Relay-paired setups add per-profile voices and an experimental realtime engine.
- Works Away From Home: add LAN, Tailscale, or public routes and the app chooses the best available path on connect.
- Sessions: create, switch, rename, and delete chats. Message history loads on demand.
- Multiple Servers and Profiles: connect to more than one server and switch in a tap; overlay an agent profile or personality per conversation.
- Relay Power Tools: optional QR pairing for remote terminal, relay-session management, media handoff, and per-feature grants.
- Notification Companion: optionally forward notification metadata to your paired relay so your assistant can summarize it. Toggle it anytime in system settings.
- Stats for Nerds: local-only counters for response timing, token usage, cost, and stream health.
- Material You: Material 3 dynamic color, light/dark/system themes, and haptics.
SECURITY AND PRIVACY
- API keys and relay tokens are stored in encrypted Android storage.
- HTTPS is enforced for remote connections; cleartext is limited to localhost or LAN setups.
- No telemetry, ads, tracking, or third-party analytics SDKs.
- Notification access and the microphone are optional and user-controlled.
- All app traffic goes only to servers you configure.
REQUIREMENTS
- Android 8.0 or later.
- A running Hermes agent for chat, management, and voice.
- Optional Hermes relay service for power tools such as terminal, notifications, and media.
- Network access to your server by local network, VPN, or internet.
OPEN SOURCE
Hermes-Relay is MIT licensed. Source, docs, and issue tracking are on GitHub.
This app is a community project and is not affiliated with or endorsed by NousResearch.
Binary file not shown.

After

Width:  |  Height:  |  Size: 44 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 37 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 152 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 182 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 112 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 131 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 129 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 246 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 140 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 165 KiB

@@ -0,0 +1 @@
Your Hermes AI agent, in your pocket - chat, voice, and control.
@@ -0,0 +1 @@
Hermes-Relay
@@ -0,0 +1,7 @@
v1.2.0 — Make it yours.
• Eight app themes, swappable sphere skins, and animated agent "pets" that react to what your agent is doing.
• See which streaming path you're on, plus a "What the agent sees" sheet showing the agent's exact context.
• ~3× faster cold start and honest loading states.
• In-app crash reporting with one-tap bug reports.
• Fixes: QR pairing on foldables, server-image & PDF crashes, in-chat model picks now apply.
+5 -2
View File
@@ -1,5 +1,6 @@
<?xml version="1.0" encoding="utf-8"?>
<manifest xmlns:android="http://schemas.android.com/apk/res/android">
<manifest xmlns:android="http://schemas.android.com/apk/res/android"
xmlns:tools="http://schemas.android.com/tools">
<uses-permission android:name="android.permission.INTERNET" />
<uses-permission android:name="android.permission.ACCESS_NETWORK_STATE" />
@@ -36,6 +37,8 @@
android:name=".MainActivity"
android:exported="true"
android:launchMode="singleTask"
android:screenOrientation="portrait"
tools:ignore="LockedOrientationActivity"
android:configChanges="uiMode|fontScale|locale|density|orientation|screenSize|screenLayout|keyboardHidden"
android:windowSoftInputMode="adjustResize"
android:theme="@style/Theme.HermesRelay.Splash">
@@ -73,7 +76,7 @@
runs while the user has explicitly enabled the toggle. specialUse
needs a Play Console foreground-service declaration at submission. -->
<service
android:name=".network.GatewayKeepAliveService"
android:name=".network.upstream.GatewayKeepAliveService"
android:exported="false"
android:foregroundServiceType="specialUse">
<property
+52 -1
View File
@@ -26,7 +26,9 @@
left: 0;
right: 0;
bottom: 0;
padding: 8px 6px 0 8px;
/* Bottom gap so xterm's last row clears the extra-keys footer
instead of butting flush against it (read as an overlap). */
padding: 8px 6px 8px 8px;
box-sizing: border-box;
}
.xterm .xterm-viewport {
@@ -149,6 +151,18 @@
}
});
// Report scroll position so the host can show a "jump to latest" pill
// while the user is scrolled up into scrollback. atBottom is true when
// the viewport is pinned to the live tail.
const reportScroll = function () {
if (!(window.AndroidBridge && window.AndroidBridge.onScrollPosition)) return;
try {
const buf = term.buffer.active;
window.AndroidBridge.onScrollPosition(buf.viewportY >= buf.baseY);
} catch (_) {}
};
term.onScroll(function () { reportScroll(); });
// ── Inbound: Android → terminal ───────────────────────────────────
// Base64-encoded payloads avoid JS string-escaping headaches when the
// stream contains control characters, raw escape sequences, or bytes
@@ -223,6 +237,13 @@
try { term.focus(); } catch (_) {}
};
// Current xterm selection as plain text ('' when nothing selected).
// Read back via WebView.evaluateJavascript for the toolbar Copy key,
// since long-press copy is unreliable inside an Android WebView.
window.getSelectionText = function () {
try { return term.getSelection() || ''; } catch (_) { return ''; }
};
window.clearTerminal = function () {
try { term.clear(); } catch (_) {}
};
@@ -239,6 +260,36 @@
}
};
// Mode-aware encoder for the on-screen toolbar's special keys
// (arrows / Home / End / Page). Arrows must follow xterm's current
// DECCKM (application cursor keys) mode: when an app like vim, less,
// or readline has requested it, an arrow is SS3-encoded (\eOA) rather
// than CSI (\e[A). The old path always sent CSI from Kotlin, which the
// running TUI could misread. We read term.modes here (where the mode
// actually lives) and route bytes back through onInput so sticky
// modifiers still apply. Page keys are mode-independent.
window.termSendKey = function (name) {
var appCursor = false;
try {
appCursor = !!(term.modes && term.modes.applicationCursorKeysMode);
} catch (_) {}
var p = appCursor ? 'O' : '[';
var map = {
ArrowUp: p + 'A',
ArrowDown: p + 'B',
ArrowRight: p + 'C',
ArrowLeft: p + 'D',
Home: p + 'H',
End: p + 'F',
PageUp: '[5~',
PageDown: '[6~',
};
var seq = map[name];
if (seq && window.AndroidBridge && window.AndroidBridge.onInput) {
window.AndroidBridge.onInput(seq);
}
};
// ── Scroll shims + gesture ────────────────────────────────────────
// xterm.js ships a scrollback buffer (scrollback: 10000 above) but
// has no built-in mobile touch-to-scroll — its input handlers are
+27 -31
View File
@@ -1,36 +1,32 @@
v1.0.0 - The 1.0 release
v1.2.0 - Make it yours
Standard path
* Chat, Manage, and voice now work on a plain Hermes agent — no relay
plugin required. The plugin is optional and only adds power tools.
Personalize
* Eight app themes in Settings → Appearance — the Hermes Relay brand plus
ports of the Nous Hermes looks (Teal, Nous Blue, Midnight, Ember, Mono,
Cyberpunk, Rosé), with light/dark.
* Swap the agent orb for an animated pet that reacts to what the agent is
doing — add, preview, and tune pets right in the app, or generate one
from sprite art with the AI authoring kit.
* Reskin the sphere, and give each agent profile its own icon.
Chat
* New gateway transport streams the agent's reasoning live, so the
Thinking block fills in during generation instead of after.
* Warm-start + opt-in "Keep connected in background" make returning to a
conversation fast.
* Attachments at desktop parity: images, PDFs, and files upload over the
gateway. If a connection can't carry a file, you'll see a notice
instead of a silent drop.
* Steer a running turn, edit & resend your messages, watch subagent
lanes, and a context-window meter — plus turn-complete notifications.
* Tap an image to open it full-screen (pinch to zoom); save or share
images and other attachments.
* Redesigned input bar: pill field, one morphing Send/Voice/Stop button.
See what's happening
* The chat status strip names the actual streaming path (Gateway, Sessions,
Completions, Runs), with a basic→best tier ladder in Chat Settings.
* Tap the context meter for a "What the agent sees" sheet — the exact extra
context prepended to your next turn.
* Voice and Realtime turns are badged in the scrollback.
Profiles
* Switch the whole agent — model, persona, and skills — per conversation.
The drawer scopes to the active profile, and switching is ephemeral: it
never changes your server's default agent.
Privacy
* When paired to the relay, the agent can mark private media and the phone
blurs it per your setting — sensitivity stays model-emitted.
Manage
* Models, provider keys, profiles + SOUL.md, and a skills hub — parity
with the desktop dashboard. Cached for instant cold-launch.
Faster & more reliable
* Cold start is about 3× faster, and model/personality/approvals load
honestly instead of showing a maybe-wrong value.
* In-app crash reporting offers a one-tap, pre-filled bug report.
* QR pairing no longer force-closes on unusual cameras (foldables); fixed
crashes opening server images and PDFs; in-chat model picks now apply.
Voice
* Realtime Agent keeps one session across turns; long runs continue in
the background and are spoken when ready.
Polish
* Seamless LAN/Tailscale handoffs (no chat reload), slide-down status
toasts, and a broad round of fixes.
Voice & terminal
* Enhanced voice control for Gemini and xAI providers.
* Leaner terminal with TUI-correct input and an isolated, tuned tmux.
@@ -10,6 +10,7 @@ import com.hermesandroid.relay.bridge.UnattendedAccessManager
import com.hermesandroid.relay.data.AppAnalytics
import com.hermesandroid.relay.power.WakeLockManager
import com.hermesandroid.relay.util.AppForegroundTracker
import com.hermesandroid.relay.util.CrashReporter
class HermesRelayApp : Application(), SingletonImageLoader.Factory {
@@ -28,6 +29,9 @@ class HermesRelayApp : Application(), SingletonImageLoader.Factory {
override fun onCreate() {
super.onCreate()
instance = this
// Install the crash handler FIRST so any failure in the rest of app
// init (or anywhere later) is captured and surfaced on next launch.
CrashReporter.install(this)
AppAnalytics.initialize(this)
// A8 — wire the bridge-gesture wake-lock wrapper so
// ActionExecutor.tap/tapText/typeText/swipe/scroll can hold
@@ -21,7 +21,6 @@ import com.hermesandroid.relay.bridge.UnattendedAccessManager
import com.hermesandroid.relay.data.BuildFlavor
import com.hermesandroid.relay.notifications.TurnCompleteNotifier
import com.hermesandroid.relay.ui.RelayApp
import com.hermesandroid.relay.util.ComposeArrWorkaround
import com.hermesandroid.relay.util.NavRouteRequest
import com.hermesandroid.relay.viewmodel.ConnectionViewModel
@@ -118,9 +117,6 @@ class MainActivity : ComponentActivity() {
setContent {
RelayApp()
}
window.decorView.post {
ComposeArrWorkaround.disableForViewTree(window.decorView)
}
}
override fun onNewIntent(intent: Intent) {
@@ -1551,7 +1551,7 @@ class ActionExecutor(private val service: HermesAccessibilityService) {
* googlePlay as a dialer-opener" per the plan.
*
* The destructive-verb confirmation modal is fired in
* [com.hermesandroid.relay.network.handlers.BridgeCommandHandler]
* [com.hermesandroid.relay.network.relay.BridgeCommandHandler]
* before we even get here — by the time this method runs, the user
* has explicitly approved the call.
*/
@@ -12,8 +12,8 @@ import android.util.Log
import com.hermesandroid.relay.bridge.BridgeSafetyManager
import com.hermesandroid.relay.bridge.UnattendedAccessManager
import com.hermesandroid.relay.data.BuildFlavor
import com.hermesandroid.relay.network.ChannelMultiplexer
import com.hermesandroid.relay.network.models.Envelope
import com.hermesandroid.relay.network.relay.ChannelMultiplexer
import com.hermesandroid.relay.network.relay.models.Envelope
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Job
import kotlinx.coroutines.delay
@@ -68,7 +68,7 @@ class HermesAccessibilityService : AccessibilityService() {
* service is not running. Written on [onServiceConnected],
* cleared on [onUnbind] / [onDestroy].
*
* Read by [com.hermesandroid.relay.network.handlers.BridgeCommandHandler]
* Read by [com.hermesandroid.relay.network.relay.BridgeCommandHandler]
* and by the Bridge UI screen (bridge-ui) to check live status.
*/
@Volatile
@@ -38,7 +38,13 @@ import kotlin.math.sqrt
* The Visualizer is attached exactly once against the ExoPlayer's
* [ExoPlayer.getAudioSessionId]. There is a known gotcha where re-attaching
* the Visualizer on every track transition invalidates the session id — the
* single-attach lifecycle here sidesteps it entirely.
* single-attach lifecycle here sidesteps it entirely. The single attach is
* triggered by whichever of {playback became live, a real session id landed}
* arrives last, so a late AudioTrack allocation (deep-buffer cold-start) can't
* leave amplitude pinned at 0 for the turn — see [attachVisualizerIfPlaying].
* That promptness matters because the voice overlay gates its output waveform
* on the first real playback-amplitude frame, so the visual follows audible
* speech instead of leading it.
*
* @param context used for [ExoPlayer.Builder]. Application context is fine;
* the player holds no view references.
@@ -109,6 +115,21 @@ class VoicePlayer(
audioSessionId: Int,
) {
cachedAudioSessionId = audioSessionId
// Deep-buffer cold-start guard. On some OEM pipelines the
// AudioTrack — and therefore a real (non-zero) session id —
// isn't allocated until *after* onIsPlayingChanged(true) has
// already fired. In that race the isPlaying-driven attach
// below ran with id == 0, no-oped, and isPlaying will not
// toggle again for the rest of a continuous TTS turn, so the
// Visualizer would never attach and [amplitude] would stay
// pinned at 0 for the whole turn. The output waveform gates
// its unfold on the first real playback-amplitude frame, so a
// never-firing amplitude leaves it stuck in the folded
// processing/spinner shape even though audio is audible.
// Attaching here — the moment a real session id lands while
// playback is already live — makes the first-audible-frame
// signal reliable regardless of when the track allocates.
attachVisualizerIfPlaying()
}
})
exoPlayer.addListener(object : Player.Listener {
@@ -124,11 +145,11 @@ class VoicePlayer(
// runs on the main thread too, so reading the getter here
// is safe and guarantees the cache is warm by the time
// playback is audible (and thus by the time barge-in
// starts its IO reader).
// starts its IO reader). If the id isn't ready yet, the
// analytics callback above re-tries the attach the instant
// it lands (see attachVisualizerIfPlaying).
cachedAudioSessionId = exoPlayer.audioSessionId
if (!visualizerAttached) {
attachVisualizer(cachedAudioSessionId)
}
attachVisualizerIfPlaying()
}
}
@@ -308,6 +329,24 @@ class VoicePlayer(
exoPlayer.release()
}
/**
* Attach the [Visualizer] iff playback is live and we haven't attached for
* this session yet. Idempotent and main-thread-only: both call sites
* ([Player.Listener.onIsPlayingChanged] and the [AnalyticsListener]'s
* `onAudioSessionIdChanged`) are delivered on the player's application
* thread, so the [visualizerAttached] check needs no extra synchronization.
*
* The delegate [attachVisualizer] still no-ops (without latching
* [visualizerAttached]) when the cached session id is 0, which preserves
* the retry: whichever of {isPlaying, valid session id} arrives last drives
* the single attach. This is the cold-start race fix — see the
* `onAudioSessionIdChanged` comment in `init`.
*/
private fun attachVisualizerIfPlaying() {
if (visualizerAttached || !_isPlaying.value) return
attachVisualizer(cachedAudioSessionId)
}
private fun attachVisualizer(audioSessionId: Int) {
if (audioSessionId == 0) {
// ExoPlayer returns 0 before the audio track is allocated; retry
@@ -6,8 +6,8 @@ import com.hermesandroid.relay.data.Connection
import com.hermesandroid.relay.data.EndpointCandidate
import com.hermesandroid.relay.data.PairingPreferences
import com.hermesandroid.relay.data.Profile
import com.hermesandroid.relay.network.ChannelMultiplexer
import com.hermesandroid.relay.network.models.Envelope
import com.hermesandroid.relay.network.relay.ChannelMultiplexer
import com.hermesandroid.relay.network.relay.models.Envelope
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.flow.MutableStateFlow
@@ -92,6 +92,19 @@ class AuthManager(
* legacy connection intentionally keeps [Connection.LEGACY_TOKEN_STORE_KEY].
*/
private val tokenStoreKey: String? = null,
/**
* When false, [init] skips the eager session-token hydration (and the
* keyset decrypt it forces). Used for the throwaway LEGACY SENTINEL manager
* that `ConnectionViewModel` builds at field-init and replaces as soon as
* the active connection hydrates — decrypting its keyset only to discard it
* is a measured ~600 ms of wasted startup keystore work, and on a device
* whose active connection isn't connection 0 the sentinel's file has no
* token anyway. The real per-connection manager (created via the active
* connection, [eagerHydrate] = true) hydrates normally; the
* `restorePersistedActiveConnectionContext` path even awaits its
* Paired/Failed state. Channel handlers are still registered either way.
*/
private val eagerHydrate: Boolean = true,
) : ChannelMultiplexer.ChannelHandler {
companion object {
@@ -102,6 +115,10 @@ class AuthManager(
private const val KEY_API_KEY = "api_server_key"
private const val HINT_API_KEY_PRESENT = "api_key_present"
private const val KEY_PAIRED_META = "paired_session_meta_json"
// Marker (in the connection-0 token store) recording that the one-shot
// pre-StrongBox `hermes_companion_auth` → `hermes_companion_auth_hw`
// migration has run, so we never rebuild the legacy keyset to re-check.
private const val KEY_LEGACY_MIGRATED = "legacy_migrated"
private const val PAIRING_CODE_LENGTH = 6
private val PAIRING_CODE_CHARS = ('A'..'Z') + ('0'..'9')
@@ -329,31 +346,29 @@ class AuthManager(
_store?.let { return it }
return storeMutex.withLock {
_store?.let { return it }
withContext(Dispatchers.IO) {
// Multi-connection: [tokenPrefsName] picks the
// EncryptedSharedPreferences filename for the bound
// connection. The legacy sentinel keeps the pre-multi-
// connection install on its original file so the existing
// paired device keeps working with no migration.
// Both encrypted backends decrypt their Tink keyset eagerly on
// construction, so a corrupt file can throw AEADBadTagException
// here. KeystoreTokenStore.tryCreate already degrades to null;
// the legacy store self-heals its file in its constructor. If
// even that rebuild fails (a fundamentally broken keystore),
// fall back to a non-persistent store rather than force-close —
// the user re-pairs, but the app stays up.
val picked: SessionTokenStore =
KeystoreTokenStore.tryCreate(context, tokenPrefsName)
?: runCatching {
LegacyEncryptedPrefsTokenStore(context, tokenPrefsName)
}.getOrElse { e ->
Log.w(TAG, "Legacy token store unavailable (${e.message}) — using in-memory fallback; re-pair required")
InMemoryTokenStore()
}
migrateFromLegacyIfNeeded(picked)
_store = picked
picked
val picked = withContext(Dispatchers.IO) {
// One keyset build per file, process-wide (see [SecureStoreCache]).
// The legacy sentinel is deferred (eagerHydrate=false) and the
// dashboard cookie store now shares this same file, so the active
// connection's token keyset is the ONLY one built on the cold-
// start critical path. [tokenPrefsName] picks the file.
//
// The build decrypts its Tink keyset eagerly, so a corrupt file
// can throw AEADBadTagException — KeystoreTokenStore.tryCreate
// degrades to null, the legacy store self-heals in its ctor, and
// a fundamentally broken keystore falls back to InMemory (the app
// stays up; the user re-pairs). See [buildRawTokenStore].
val s = SecureStoreCache.getOrBuild(tokenPrefsName) {
buildRawTokenStore(context, tokenPrefsName)
}
// Migration runs AFTER the (shared) build so the cookie store can
// trigger the build without needing token-migration logic; a
// marker makes it read the legacy file at most once ever.
migrateFromLegacyIfNeeded(s)
s
}
_store = picked
picked
}
}
@@ -365,14 +380,33 @@ class AuthManager(
*/
private fun migrateFromLegacyIfNeeded(picked: SessionTokenStore) {
if (picked is LegacyEncryptedPrefsTokenStore) return
// Multi-connection: only the legacy connection inherits from the pre-
// multi-connection `hermes_companion_auth` file. A freshly-minted
// per-connection store must NOT be seeded from the legacy file or
// Gate on the FILE, not the connection id. Only the legacy connection-0
// file (`hermes_companion_auth_hw`) inherits from the pre-multi-
// connection `hermes_companion_auth` file; a freshly-minted per-
// connection store (`hermes_auth_<id>`) must NOT be seeded from it or
// we'd copy connection 0's token into every new connection.
if (connectionId != CONNECTION_ID_LEGACY) return
//
// Why file-gated rather than `connectionId == CONNECTION_ID_LEGACY`:
// the store build is now cached/deduped across the legacy sentinel and
// the real connection-0 manager, so whichever one builds the file first
// runs this migration. Both share `tokenPrefsName == LEGACY_TOKEN_STORE_KEY`
// but only the sentinel had `connectionId == CONNECTION_ID_LEGACY`, so
// the old id-based gate would skip migration whenever the real manager
// won the race — dropping a pre-StrongBox user's token. The file name is
// the same for both, so gating on it is race-proof.
if (tokenPrefsName != Connection.LEGACY_TOKEN_STORE_KEY) return
// Read the legacy file at most ONCE ever. The build is now cache-shared
// (and the cookie store can trigger it without migrating), so without
// this marker every freshly-rebuilt connection-0 AuthManager would
// re-build the legacy `hermes_companion_auth` keyset just to find it
// already drained — re-introducing the startup cost we just removed.
if (picked.contains(KEY_LEGACY_MIGRATED)) return
val legacy = try {
LegacyEncryptedPrefsTokenStore(context)
} catch (_: Exception) {
// Legacy file unreadable/corrupt — nothing to inherit. Still mark
// done so its keyset isn't rebuilt on every launch.
picked.putString(KEY_LEGACY_MIGRATED, "1")
return
}
@@ -396,6 +430,7 @@ class AuthManager(
// backup copies of the session token lying around.
legacy.clearAll()
}
picked.putString(KEY_LEGACY_MIGRATED, "1")
}
/** Cert pin store — shared across all relay connections. */
@@ -524,24 +559,28 @@ class AuthManager(
// one-line change in [onMessage].
multiplexer.registerHandler("pairing", this)
// Check for existing session token off main thread
scope.launch {
val s = store()
val existingToken = s.getString(KEY_SESSION_TOKEN)
if (existingToken != null) {
_authState.value = AuthState.Paired(existingToken)
_currentPairedSession.value = loadStoredMetadata(existingToken)
Log.i(
TAG,
"init: hydrated existing session_token=${existingToken.take(8)}… " +
"→ authState=Paired (stale-at-startup unless this is a real continuous session)"
)
} else {
Log.i(TAG, "init: no stored session_token → authState stays Unpaired")
// Check for existing session token off main thread. Skipped for the
// throwaway sentinel (eagerHydrate=false) so it never pays the keyset
// decrypt for a store that's about to be replaced (see [eagerHydrate]).
if (eagerHydrate) {
scope.launch {
val s = store()
val existingToken = s.getString(KEY_SESSION_TOKEN)
if (existingToken != null) {
_authState.value = AuthState.Paired(existingToken)
_currentPairedSession.value = loadStoredMetadata(existingToken)
Log.i(
TAG,
"init: hydrated existing session_token=${existingToken.take(8)}… " +
"→ authState=Paired (stale-at-startup unless this is a real continuous session)"
)
} else {
Log.i(TAG, "init: no stored session_token → authState stays Unpaired")
}
// Converge the plain api-key-present hint with the decrypted
// truth (also repairs a hint that predates legacy migration).
recordApiKeyHint(!s.getString(KEY_API_KEY).isNullOrBlank())
}
// Converge the plain api-key-present hint with the decrypted
// truth (also repairs a hint that predates legacy migration).
recordApiKeyHint(!s.getString(KEY_API_KEY).isNullOrBlank())
}
}
@@ -6,6 +6,44 @@ import android.os.Build
import android.util.Log
import androidx.security.crypto.EncryptedSharedPreferences
import androidx.security.crypto.MasterKey
import java.util.concurrent.ConcurrentHashMap
/**
* Process-global cache for encrypted stores, keyed by prefs-file name.
*
* `EncryptedSharedPreferences.create()` unwraps a Tink keyset via a KeyStore op
* (~0.6–1 s on StrongBox), and Tink serializes those process-globally — so a
* second build of the SAME file is pure waste (the measured cold-start
* `Long monitor contention … AndroidKeysetManager.build()` with `waiters=1..4`).
*
* Caching by file name means each file's keyset builds ONCE process-wide. The
* cache is **synchronous** ([ConcurrentHashMap.computeIfAbsent], which holds a
* per-key lock so the build runs at most once per file) precisely so the SAME
* instance serves both the suspend token path (callers wrap this in
* [kotlinx.coroutines.Dispatchers.IO]) AND the synchronous OkHttp cookie-jar
* path — which is how the dashboard cookies now ride the connection's
* already-built token keyset instead of building a second one.
*
* The build is ~1 s on StrongBox: call only from IO / OkHttp threads, never the
* main thread.
*/
internal object SecureStoreCache {
private val instances = ConcurrentHashMap<String, SessionTokenStore>()
fun getOrBuild(prefsName: String, build: () -> SessionTokenStore): SessionTokenStore =
instances.computeIfAbsent(prefsName) { build() }
}
/**
* Build the raw encrypted store for [prefsName] — Keystore-backed when possible,
* self-healing legacy fallback, in-memory last resort. No migration. Shared by
* the token store and the dashboard cookie store so a given file always yields
* the SAME backend, via [SecureStoreCache].
*/
internal fun buildRawTokenStore(context: Context, prefsName: String): SessionTokenStore =
KeystoreTokenStore.tryCreate(context, prefsName)
?: runCatching { LegacyEncryptedPrefsTokenStore(context, prefsName) }
.getOrElse { InMemoryTokenStore() }
/**
* Abstraction over the storage backend for the relay session token + API key
@@ -25,7 +25,6 @@ import androidx.savedstate.SavedStateRegistryOwner
import androidx.savedstate.setViewTreeSavedStateRegistryOwner
import com.hermesandroid.relay.ui.components.BridgeStatusOverlayChip
import com.hermesandroid.relay.ui.components.DestructiveVerbConfirmDialog
import com.hermesandroid.relay.util.ComposeArrWorkaround
import java.util.concurrent.ConcurrentHashMap
/**
@@ -158,7 +157,6 @@ class BridgeStatusOverlay(context: Context) : ConfirmationOverlayHost {
Log.w(TAG, "addView(chip) failed", it)
return
}
compose.post { ComposeArrWorkaround.disableForViewTree(compose) }
chipView = compose
chipUnattended = unattended
}
@@ -227,7 +225,6 @@ class BridgeStatusOverlay(context: Context) : ConfirmationOverlayHost {
onResult(false)
return
}
compose.post { ComposeArrWorkaround.disableForViewTree(compose) }
activeConfirmations[request.id] = compose
}
@@ -233,7 +233,7 @@ object UnattendedAccessManager {
* Acquire the screen-bright wake lock + opportunistically request
* keyguard dismiss. Synchronous — does not suspend. The caller (
* [com.hermesandroid.relay.accessibility.ActionExecutor] wrapper, or
* [com.hermesandroid.relay.network.handlers.BridgeCommandHandler]
* [com.hermesandroid.relay.network.relay.BridgeCommandHandler]
* pre-dispatch hook) holds onto the result and decides whether to
* proceed with the action.
*
@@ -66,11 +66,17 @@ object AgentDisplay {
localDisplayAlias(localDisplayAlias)?.let { return it }
profileDisplayName(profile)?.let { return it }
// "none"/"neutral" are the upstream "cleared overlay" aliases — treat
// them like "default" for identity: fall through to the server default
// (or the base connection identity) rather than rendering the literal
// word as an agent name.
val personalityName = if (
selectedPersonality == "default" &&
isClearedPersonality(selectedPersonality) &&
defaultPersonality.isNotBlank()
) {
defaultPersonality
} else if (isClearedPersonality(selectedPersonality)) {
""
} else {
selectedPersonality
}
@@ -83,10 +89,18 @@ object AgentDisplay {
}
}
/** True for the upstream "clear the overlay" aliases (default == none == neutral). */
fun isClearedPersonality(value: String): Boolean =
value.trim().lowercase() in setOf("default", "none", "neutral", "")
fun personalityLabel(
selectedPersonality: String,
defaultPersonality: String,
): String = when {
// Explicit "none" — show "None" (or the configured default name, if any)
// so the cleared-overlay state is legible in the picker header.
selectedPersonality.trim().lowercase() in setOf("none", "neutral") ->
if (defaultPersonality.isNotBlank()) titleCase(defaultPersonality.trim()) else "None"
selectedPersonality != "default" && selectedPersonality.isNotBlank() ->
titleCase(selectedPersonality.trim())
defaultPersonality.isNotBlank() -> titleCase(defaultPersonality.trim())
@@ -99,6 +113,19 @@ object AgentDisplay {
?.takeIf { it.isNotEmpty() }
?.takeUnless { it.lowercase() in GENERIC_MODEL_ALIASES }
/**
* A model string safe to SEND to the server as a model override or
* `config.set model=…`. Returns null for the generic agent placeholders
* ("hermes-agent", …) which are NOT real models — the server rejects them
* (HTTP 400) and falls back. Null means "send no model; use the server's
* configured default."
*/
fun requestModelName(model: String?): String? =
model
?.trim()
?.takeIf { it.isNotEmpty() }
?.takeUnless { it.lowercase() in GENERIC_MODEL_ALIASES }
fun isServerDefaultAlias(profileName: String?): Boolean =
profileName?.trim()?.equals("default", ignoreCase = true) == true
@@ -46,7 +46,7 @@ data class ChatMessage(
/**
* Rich content cards emitted by the agent via `CARD:{json}` line
* markers in the text stream. Parsed in
* [com.hermesandroid.relay.network.handlers.ChatHandler.scanForCardMarkers]
* [com.hermesandroid.relay.network.upstream.ChatHandler.scanForCardMarkers]
* and rendered inline by
* [com.hermesandroid.relay.ui.components.HermesCardBubble]. Mirrors
* [attachments]' lifecycle — the marker line is stripped from
@@ -76,7 +76,7 @@ data class ChatMessage(
* The sync builder treats messages with [voiceIntent] != null and
* [VoiceIntentTrace.syncedToServer] == false as the inputs to its
* synthesis pass; on a successful send we flip [VoiceIntentTrace.syncedToServer]
* to true via [com.hermesandroid.relay.network.handlers.ChatHandler.markVoiceIntentsSynced]
* to true via [com.hermesandroid.relay.network.upstream.ChatHandler.markVoiceIntentsSynced]
* so they're not re-sent on the next turn.
*/
val voiceIntent: VoiceIntentTrace? = null,
@@ -89,12 +89,33 @@ data class ChatMessage(
* the durable session turn; the provider's spoken summary is UI/runtime
* provenance, not another canonical assistant message.
*/
val realtimeTurn: RealtimeTurnTrace? = null
val realtimeTurn: RealtimeTurnTrace? = null,
/**
* True for bubbles that exist ONLY on the client and have no server-side
* row — slash-command notices, voice-intent traces, the steer echo, gateway
* ask cards, an errored turn the server never persisted, and a provider-only
* (non-Hermes-backed) realtime turn. The post-turn history reload
* ([com.hermesandroid.relay.network.upstream.ChatHandler.loadMessageHistory])
* preserves any client-only message whose id is absent from the reloaded
* server transcript; without the flag those orphans would be silently
* wiped by the reconcile.
*
* Replaces the old id-prefix whitelist (`voice-intent-`/`steer-`/`ask-`/
* `system-notice-`) + "Error"-badge sniffing: each creator now declares its
* own provenance instead of the reconcile having to know every id
* convention. Defaults false so every server-backed message and existing
* call site stays correct.
*
* NOTE: an "Error" badge alone does NOT make a message preservable — a turn
* can error *after* persisting server-side, and that message must still
* reconcile normally. Only [clientOnly] gates orphan preservation.
*/
val clientOnly: Boolean = false,
)
/**
* Structured details about a phone-local voice intent that was dispatched
* in-process via [com.hermesandroid.relay.network.handlers.BridgeCommandHandler.handleLocalCommand].
* in-process via [com.hermesandroid.relay.network.relay.BridgeCommandHandler.handleLocalCommand].
*
* Captured on a [ChatMessage] (id prefix `voice-intent-`) so the next chat
* payload can include synthetic OpenAI-format `assistant` + `tool` message
@@ -123,12 +144,12 @@ data class ChatMessage(
* includes an `error` field.
* @property resultJson Compact JSON object describing the dispatch outcome.
* On success, typically `{"ok":true,...}` with any tool-specific fields
* from [com.hermesandroid.relay.network.handlers.LocalDispatchResult.resultJson].
* from [com.hermesandroid.relay.network.shared.LocalDispatchResult.resultJson].
* On failure, an error envelope including `ok:false`, `error`, optionally
* `error_code`. Stored as a string and rendered verbatim into the
* synthetic `tool`-role message's `content` field.
* @property syncedToServer Idempotency guard. Flipped to true by
* [com.hermesandroid.relay.network.handlers.ChatHandler.markVoiceIntentsSynced]
* [com.hermesandroid.relay.network.upstream.ChatHandler.markVoiceIntentsSynced]
* the moment we hand the request payload to the API client. Once true,
* the trace is excluded from future sync passes — the server-side
* session has already absorbed it.
@@ -186,7 +207,23 @@ data class Attachment(
/** Opaque token from `MEDIA:hermes-relay://<token>` — identifies the file on the relay. */
val relayToken: String? = null,
/** content:// URI from the FileProvider once bytes are cached to disk. */
val cachedUri: String? = null
val cachedUri: String? = null,
/**
* Whether this attachment was flagged sensitive (NSFW / spoiler) and should
* render blurred until the user taps to reveal — honored per the user's
* `MediaSettings.blurMode`.
*
* The flag is **model-emitted metadata, never an on-device or relay-side
* classifier** (see `docs/plans/2026-06-18-attachment-experience.md` §C): the
* agent annotates media it surfaces, the relay transports the bit
* authoritatively via the `X-Media-Sensitive` response header, and the
* client merely renders the blur. Populated for inbound attachments from
* [com.hermesandroid.relay.network.relay.RelayHttpClient.FetchedMedia.sensitive]
* when the bytes flip to [AttachmentState.LOADED]. Defaults false so every
* existing outbound/inbound call site stays valid and unflagged media
* renders exactly as before.
*/
val sensitive: Boolean = false
) {
val isImage: Boolean get() = contentType.startsWith("image/")
@@ -30,7 +30,7 @@ data class DashboardConnectionStatus(
* open the token store.
*
* Switching connection is a HEAVY context swap — caller is expected to tear down
* the current [com.hermesandroid.relay.network.ConnectionManager],
* the current [com.hermesandroid.relay.network.relay.ConnectionManager],
* [com.hermesandroid.relay.auth.AuthManager], and API client, then construct
* fresh ones pointed at the new connection's `tokenStoreKey`.
*
@@ -290,6 +290,8 @@ data class Connection(
role = role.ifBlank { inferRouteRole(apiServerUrl) },
priority = priority,
api = ApiEndpoint(host = host, port = port, tls = tls),
dashboard = deriveDefaultDashboardUrl(apiServerUrl)
?.let { DashboardEndpoint(url = it) },
relay = RelayEndpoint(url = resolvedRelayUrl, transportHint = transportHint),
)
}
@@ -8,8 +8,8 @@ import androidx.datastore.preferences.core.edit
import androidx.datastore.preferences.core.stringPreferencesKey
import com.hermesandroid.relay.auth.AuthManager
import com.hermesandroid.relay.auth.ConnectionAuthSecrets
import com.hermesandroid.relay.network.EncryptedDashboardCookieStore
import com.hermesandroid.relay.network.StoredDashboardCookie
import com.hermesandroid.relay.network.upstream.EncryptedDashboardCookieStore
import com.hermesandroid.relay.network.upstream.StoredDashboardCookie
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.flow.first
import kotlinx.coroutines.flow.map
@@ -147,6 +147,7 @@ class DataManager(
dashboardCookies = EncryptedDashboardCookieStore(
context = context,
connectionId = connection.id,
tokenStoreKey = connection.tokenStoreKey,
).load().map { it.toBackup() },
)
}
@@ -185,6 +186,7 @@ class DataManager(
EncryptedDashboardCookieStore(
context = context,
connectionId = connection.id,
tokenStoreKey = connection.tokenStoreKey,
).save(secret.dashboardCookies.map { it.toStoredCookie() })
}
}
@@ -40,6 +40,10 @@ data class EndpointCandidate(
val priority: Int = 0,
val api: ApiEndpoint,
val relay: RelayEndpoint,
val dashboard: DashboardEndpoint? = null,
val proxy: ProxyEndpoint? = null,
val security: String? = null,
val recommended: Boolean = false,
)
/**
@@ -61,6 +65,17 @@ data class ApiEndpoint(
get() = "${if (tls) "https" else "http"}://$host:$port"
}
/**
* Dashboard/admin surface for an [EndpointCandidate]. This is optional so
* older v3 payloads that only carried API + Relay endpoints keep
* deserializing; when absent, Android derives the conventional same-host
* `:9119` dashboard URL from [ApiEndpoint].
*/
@Serializable
data class DashboardEndpoint(
val url: String,
)
/**
* The relay-server half of an [EndpointCandidate] — the WSS URL the phone
* opens for the bridge + terminal channels.
@@ -78,6 +93,22 @@ data class RelayEndpoint(
val transportHint: String? = null,
)
/**
* Optional plugin-owned secure proxy route. Unlike [api], [dashboard], and
* [relay], this is one app-facing base that can cover all Hermes-Relay
* supported traffic after pairing. It is deliberately optional so plugin
* proxy support can be advertised by newer payloads without changing the
* standard upstream connection model.
*/
@Serializable
data class ProxyEndpoint(
val url: String,
@SerialName("transport_hint")
val transportHint: String? = null,
@SerialName("pin_sha256")
val pinSha256: String? = null,
)
/**
* Returns true when [EndpointCandidate.role] is one of the built-in, styled
* roles: `lan`, `tailscale`, or `public`. Case-insensitive match — but the
@@ -89,7 +120,7 @@ data class RelayEndpoint(
*/
fun EndpointCandidate.isKnownRole(): Boolean {
return when (role.lowercase()) {
"lan", "tailscale", "public" -> true
"lan", "tailscale", "public", "plugin_proxy", "plugin-proxy", "https" -> true
else -> false
}
}
@@ -106,7 +137,15 @@ fun EndpointCandidate.displayLabel(): String {
return when (role.lowercase()) {
"lan" -> "LAN"
"tailscale" -> "Tailscale"
"public" -> "Public"
"public" -> if (api.tls) "HTTPS" else "Public"
"https" -> "HTTPS"
"plugin_proxy", "plugin-proxy" -> "Plugin proxy"
else -> "Custom VPN ($role)"
}
}
fun EndpointCandidate.hasSecureProxy(): Boolean =
proxy?.url?.startsWith("https://", ignoreCase = true) == true ||
proxy?.url?.startsWith("wss://", ignoreCase = true) == true ||
role.equals("plugin_proxy", ignoreCase = true) ||
role.equals("plugin-proxy", ignoreCase = true)
@@ -11,7 +11,7 @@ import androidx.datastore.preferences.core.edit
* Shared by [com.hermesandroid.relay.viewmodel.ConnectionViewModel] (the
* StateFlow + setter that drive the foreground service and the client's
* no-background-close flag) and
* [com.hermesandroid.relay.network.GatewayKeepAliveService]'s Stop notification
* [com.hermesandroid.relay.network.upstream.GatewayKeepAliveService]'s Stop notification
* action, so both read/write the same key.
*/
val KEY_GATEWAY_KEEP_ALIVE = booleanPreferencesKey("gateway_keep_alive_background")
@@ -6,7 +6,7 @@ import kotlinx.serialization.Serializable
/**
* A rich content card emitted inline in an assistant message via the
* `CARD:{json}` line marker. Parsed by
* [com.hermesandroid.relay.network.handlers.ChatHandler] and rendered by
* [com.hermesandroid.relay.network.upstream.ChatHandler] and rendered by
* [com.hermesandroid.relay.ui.components.HermesCardBubble].
*
* The marker lives in the text stream alongside `MEDIA:...` for the same
@@ -226,7 +226,7 @@ data class HermesCardAction(
* (with structured `tool_calls`) + `tool` message pairs under a synthetic
* `hermes_card_action` tool name, splicing them into the session history
* the LLM sees. After the API client takes ownership of the request,
* [com.hermesandroid.relay.network.handlers.ChatHandler.markCardDispatchesSynced]
* [com.hermesandroid.relay.network.upstream.ChatHandler.markCardDispatchesSynced]
* flips [syncedToServer] so subsequent turns don't re-send the same
* trace.
*/
@@ -238,7 +238,7 @@ data class HermesCardDispatch(
/**
* Idempotency guard for the server-side session sync path.
* Flipped to true by
* [com.hermesandroid.relay.network.handlers.ChatHandler.markCardDispatchesSynced]
* [com.hermesandroid.relay.network.upstream.ChatHandler.markCardDispatchesSynced]
* once the API client has accepted the request that carried this
* dispatch's synthetic message pair. Once true, the dispatch is
* excluded from future
@@ -4,9 +4,26 @@ import android.content.Context
import androidx.datastore.preferences.core.booleanPreferencesKey
import androidx.datastore.preferences.core.edit
import androidx.datastore.preferences.core.intPreferencesKey
import androidx.datastore.preferences.core.stringPreferencesKey
import kotlinx.coroutines.flow.Flow
import kotlinx.coroutines.flow.map
/**
* How aggressively inbound media is blurred behind a "tap to reveal" gate.
*
* - [OFF] never blur — show everything immediately.
* - [FLAGGED] blur only media the agent flagged sensitive (the model-emitted
* `X-Media-Sensitive` bit; see
* `docs/plans/2026-06-18-attachment-experience.md` §C). This is
* the product default: zero blur when nothing is flagged.
* - [ALL_IMAGES] blur every inbound image regardless of source. Works on the
* pure standard path with no server support at all.
*
* Persisted by [Enum.name] so adding cases later is forward-safe; an unknown
* stored value decodes back to the default rather than throwing.
*/
enum class BlurMode { OFF, FLAGGED, ALL_IMAGES }
/**
* User-tunable limits for inbound media attachments fetched from the relay.
*
@@ -20,12 +37,17 @@ import kotlinx.coroutines.flow.map
* - [autoFetchOnCellular] master switch: when false, the cellular-network
* case always inserts a manual-download placeholder.
* - [cachedMediaCapMb] LRU cap on the `hermes-media/` cache directory.
* - [blurSensitive] whether (and which) inbound images render behind a
* tap-to-reveal blur — see [BlurMode]. Unlike the four knobs above this one
* also applies on the standard (no-Relay) path, since [BlurMode.ALL_IMAGES]
* needs no server cooperation.
*/
data class MediaSettings(
val maxInboundSizeMb: Int = 25,
val autoFetchThresholdMb: Int = 2,
val autoFetchOnCellular: Boolean = false,
val cachedMediaCapMb: Int = 200
val cachedMediaCapMb: Int = 200,
val blurSensitive: BlurMode = BlurMode.FLAGGED
)
/**
@@ -39,11 +61,18 @@ class MediaSettingsRepository(private val context: Context) {
private val KEY_AUTO_FETCH_THRESHOLD_MB = intPreferencesKey("media_auto_fetch_threshold_mb")
private val KEY_AUTO_FETCH_ON_CELLULAR = booleanPreferencesKey("media_auto_fetch_on_cellular")
private val KEY_CACHED_MEDIA_CAP_MB = intPreferencesKey("media_cached_cap_mb")
private val KEY_BLUR_SENSITIVE = stringPreferencesKey("media_blur_sensitive")
const val DEFAULT_MAX_INBOUND_MB = 25
const val DEFAULT_AUTO_FETCH_THRESHOLD_MB = 2
const val DEFAULT_AUTO_FETCH_ON_CELLULAR = false
const val DEFAULT_CACHED_MEDIA_CAP_MB = 200
val DEFAULT_BLUR_SENSITIVE = BlurMode.FLAGGED
/** Decode a persisted [BlurMode] name, falling back to the default. */
private fun parseBlurMode(raw: String?): BlurMode =
raw?.let { name -> BlurMode.entries.firstOrNull { it.name == name } }
?: DEFAULT_BLUR_SENSITIVE
}
val settings: Flow<MediaSettings> = context.relayDataStore.data.map { prefs ->
@@ -51,10 +80,21 @@ class MediaSettingsRepository(private val context: Context) {
maxInboundSizeMb = prefs[KEY_MAX_INBOUND_MB] ?: DEFAULT_MAX_INBOUND_MB,
autoFetchThresholdMb = prefs[KEY_AUTO_FETCH_THRESHOLD_MB] ?: DEFAULT_AUTO_FETCH_THRESHOLD_MB,
autoFetchOnCellular = prefs[KEY_AUTO_FETCH_ON_CELLULAR] ?: DEFAULT_AUTO_FETCH_ON_CELLULAR,
cachedMediaCapMb = prefs[KEY_CACHED_MEDIA_CAP_MB] ?: DEFAULT_CACHED_MEDIA_CAP_MB
cachedMediaCapMb = prefs[KEY_CACHED_MEDIA_CAP_MB] ?: DEFAULT_CACHED_MEDIA_CAP_MB,
blurSensitive = parseBlurMode(prefs[KEY_BLUR_SENSITIVE])
)
}
/**
* Just the blur knob — a standalone flow so per-bubble UI can observe it
* without collecting (and recomposing on) the whole [MediaSettings].
* Built here (outside composition) on purpose so callers can
* `collectAsState()` it without tripping `FlowOperatorInvokedInComposition`.
*/
val blurMode: Flow<BlurMode> = context.relayDataStore.data.map { prefs ->
parseBlurMode(prefs[KEY_BLUR_SENSITIVE])
}
suspend fun setMaxInboundSize(mb: Int) {
context.relayDataStore.edit { it[KEY_MAX_INBOUND_MB] = mb.coerceAtLeast(1) }
}
@@ -70,4 +110,8 @@ class MediaSettingsRepository(private val context: Context) {
suspend fun setCachedMediaCap(mb: Int) {
context.relayDataStore.edit { it[KEY_CACHED_MEDIA_CAP_MB] = mb.coerceAtLeast(10) }
}
suspend fun setBlurSensitive(mode: BlurMode) {
context.relayDataStore.edit { it[KEY_BLUR_SENSITIVE] = mode.name }
}
}
@@ -0,0 +1,70 @@
package com.hermesandroid.relay.data
import android.content.Context
import androidx.datastore.core.DataStore
import androidx.datastore.preferences.core.Preferences
import androidx.datastore.preferences.core.edit
import androidx.datastore.preferences.core.stringPreferencesKey
import androidx.datastore.preferences.preferencesDataStore
import kotlinx.coroutines.flow.Flow
import kotlinx.coroutines.flow.map
/**
* Local-only per-profile agent icons — the visual twin of [ProfileDisplayAliasStore].
*
* Stores a **file path** to an image that was copied into app storage (not a SAF
* content URI, so it survives without a persistable-permission grant). Like the
* name alias, these are phone-UI labels only: never sent to Hermes, and keyed by
* connection + profile context so the same server-default agent can wear a
* different face on each configured host.
*/
class ProfileIconStore(
private val dataStore: DataStore<Preferences>,
) {
constructor(context: Context) : this(context.profileIconsDataStore)
companion object {
private const val PREFIX = "profile_icon__"
private fun keyName(connectionId: String, profileName: String?): String =
"$PREFIX${connectionId}__${AgentDisplay.profileSessionKey(profileName)}"
private fun keyFor(connectionId: String, profileName: String?) =
stringPreferencesKey(keyName(connectionId, profileName))
private fun connectionPrefix(connectionId: String): String =
"$PREFIX${connectionId}__"
}
suspend fun setIcon(connectionId: String, profileName: String?, path: String?) {
dataStore.edit { prefs ->
val key = keyFor(connectionId, profileName)
if (path.isNullOrBlank()) {
prefs.remove(key)
} else {
prefs[key] = path
}
}
}
fun iconFlow(connectionId: String, profileName: String?): Flow<String?> {
val key = keyFor(connectionId, profileName)
return dataStore.data.map { prefs -> prefs[key] }
}
suspend fun clearConnection(connectionId: String) {
val prefix = connectionPrefix(connectionId)
dataStore.edit { prefs ->
prefs.asMap().keys
.filter { it.name.startsWith(prefix) }
.forEach { prefs.remove(it) }
}
}
suspend fun clearAll() {
dataStore.edit { prefs -> prefs.clear() }
}
}
internal val Context.profileIconsDataStore: DataStore<Preferences>
by preferencesDataStore(name = "profile_icons")
@@ -8,8 +8,11 @@ import androidx.datastore.preferences.core.longPreferencesKey
import androidx.datastore.preferences.core.Preferences
import androidx.datastore.preferences.core.stringPreferencesKey
import kotlinx.coroutines.flow.Flow
import kotlinx.coroutines.flow.MutableStateFlow
import kotlinx.coroutines.flow.StateFlow
import kotlinx.coroutines.flow.asStateFlow
import kotlinx.coroutines.flow.combine
import kotlinx.coroutines.flow.distinctUntilChanged
import kotlinx.coroutines.flow.map
/**
* User-tunable voice mode preferences.
@@ -37,8 +40,60 @@ data class VoiceSettings(
* docs/plans/2026-05-24-realtime-persistent-session.md.
*/
val realtimePersistentSession: Boolean = true,
/**
* Enhanced-voice overrides for the relay TTS path, mapped onto the active
* provider (Gemini / xAI). Empty string / false means "use the server's
* saved config" — the relay only applies a field when it is set. Surfaced
* in Voice Settings only when the relay advertises an enhanced provider
* (`/voice/config` `tts.enhanced.supported`). Field meaning is generic:
* `enhancedVoice` → Gemini voice / xAI voice_id; `enhancedAudioTags` →
* Gemini audio_tags / xAI auto_speech_tags; `enhancedPersona` is Gemini-only
* and `enhancedLanguage` is xAI-only.
*/
val enhancedVoice: String = "",
val enhancedModel: String = "",
val enhancedAudioTags: Boolean = false,
val enhancedPersona: String = "",
val enhancedLanguage: String = "",
)
/**
* Per-request enhanced-voice overrides forwarded to the relay's
* `/voice/synthesize`. Mirrors the generic fields recognized by
* `plugin/relay/voice.py:_extract_voice_overrides`; the relay maps them onto
* the active provider's config.
*/
data class EnhancedVoiceOverrides(
val voice: String? = null,
val model: String? = null,
val audioTags: Boolean? = null,
val personaPrompt: String? = null,
val language: String? = null,
) {
val isEmpty: Boolean
get() = voice == null && model == null && audioTags == null &&
personaPrompt == null && language == null
companion object {
/**
* Build overrides from persisted settings, or null when nothing is set
* (so the relay falls back to the server's saved config). The audio-tags
* toggle only sends `true` — leaving it off defers to the server default
* rather than forcing it off.
*/
fun fromSettings(s: VoiceSettings): EnhancedVoiceOverrides? {
val overrides = EnhancedVoiceOverrides(
voice = s.enhancedVoice.takeIf { it.isNotBlank() },
model = s.enhancedModel.takeIf { it.isNotBlank() },
audioTags = true.takeIf { s.enhancedAudioTags },
personaPrompt = s.enhancedPersona.takeIf { it.isNotBlank() },
language = s.enhancedLanguage.takeIf { it.isNotBlank() },
)
return overrides.takeUnless { it.isEmpty }
}
}
}
enum class VoiceEngineMode(val storageValue: String) {
HermesVoiceOutput("hermes_voice_output"),
RealtimeAgent("realtime_agent");
@@ -60,13 +115,63 @@ enum class VoiceAudioRoute(val storageValue: String) {
}
}
/**
* Active scope for per-profile voice prefs.
*
* Mirrors [ProfileSelectionStore]'s `_<connectionId>` keying and extends it to
* `_<connectionId>_<profile>` so per-profile voice picks don't leak across
* profiles (or across connections that expose a same-named profile).
*
* A null/blank [profileName] is the "default / launch profile" and resolves to
* the un-namespaced global keys — i.e. the default profile *is* the base layer
* that named profiles override. A null/blank [connectionId] degrades to
* profile-only namespacing, which still isolates profiles within one
* connection; it just can't disambiguate two connections with a same-named
* profile. See [VoicePreferencesRepository.setActiveScope].
*/
data class VoiceProfileScope(
val connectionId: String? = null,
val profileName: String? = null,
) {
companion object {
val Global = VoiceProfileScope()
}
}
class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>) {
constructor(context: Context) : this(context.relayDataStore)
companion object {
private val KEY_ENGINE_MODE = stringPreferencesKey("voice_engine_mode")
private val KEY_AUDIO_ROUTE = stringPreferencesKey("voice_audio_route")
// --- Per-profile keys (override map; namespaced by active scope) -----
// These are stored as base NAME strings (not typed Key<>s) so the
// scoped key can be built per (connectionId, profile) at read/write
// time. Resolution layers a per-profile value over the global value
// over the hard default — see [scopedName] / [resolveString].
//
// Why these are per-profile: engine mode, audio route, and the
// enhanced-voice overrides describe *which voice the agent speaks
// with*, which is a property of the profile (the relay already
// persists `voice_output:`/`realtime_voice:` per profile and
// `RelayVoiceClient` already sends `?profile=`). Keeping them global
// leaked one profile's voice onto every other profile.
private const val KEY_ENGINE_MODE = "voice_engine_mode"
private const val KEY_AUDIO_ROUTE = "voice_audio_route"
private const val KEY_ENH_VOICE = "voice_enh_voice"
private const val KEY_ENH_MODEL = "voice_enh_model"
private const val KEY_ENH_AUDIO_TAGS = "voice_enh_audio_tags"
private const val KEY_ENH_PERSONA = "voice_enh_persona"
private const val KEY_ENH_LANGUAGE = "voice_enh_language"
// --- Global keys (shared across profiles; never namespaced) ----------
// Why these stay global: interaction-mode and silence-threshold are
// ergonomic input preferences about *how the user drives the mic*, not
// about the agent's voice — a user wants the same tap/hold/continuous
// habit regardless of which profile is active. auto-tts and the STT
// language hint are dead/experimental controls today, and the two
// realtime diagnostic toggles (trace details, persistent session) are
// engine-behaviour switches that aren't profile-specific. Keeping them
// un-namespaced means switching profiles never churns these.
private val KEY_INTERACTION_MODE = stringPreferencesKey("voice_interaction_mode")
private val KEY_SILENCE_THRESHOLD_MS = longPreferencesKey("voice_silence_threshold_ms")
private val KEY_AUTO_TTS = booleanPreferencesKey("voice_auto_tts")
@@ -83,37 +188,149 @@ class VoicePreferencesRepository(private val dataStore: DataStore<Preferences>)
const val DEFAULT_LANGUAGE = ""
const val DEFAULT_REALTIME_TRACE_DETAILS = false
const val DEFAULT_REALTIME_PERSISTENT_SESSION = true
/**
* Build the storage name for a per-profile [base] key under [scope].
*
* - null/blank profile → returns [base] verbatim (the global base
* layer; the default profile reads/writes the un-namespaced key).
* - profile set, no connection → `<base>_<profile>`.
* - profile + connection set → `<base>_<connectionId>_<profile>`,
* matching [ProfileSelectionStore]'s connection-first ordering.
*/
internal fun scopedName(base: String, scope: VoiceProfileScope): String {
val profile = scope.profileName?.trim()?.takeIf { it.isNotEmpty() } ?: return base
val conn = scope.connectionId?.trim()?.takeIf { it.isNotEmpty() }
return if (conn != null) "${base}_${conn}_$profile" else "${base}_$profile"
}
}
val settings: Flow<VoiceSettings> = dataStore.data
.map { prefs ->
VoiceSettings(
engineMode = VoiceEngineMode.fromStorage(
prefs[KEY_ENGINE_MODE] ?: DEFAULT_ENGINE_MODE,
).storageValue,
audioRoute = VoiceAudioRoute.fromStorage(
prefs[KEY_AUDIO_ROUTE] ?: DEFAULT_AUDIO_ROUTE,
).storageValue,
interactionMode = prefs[KEY_INTERACTION_MODE] ?: DEFAULT_INTERACTION_MODE,
silenceThresholdMs = prefs[KEY_SILENCE_THRESHOLD_MS] ?: DEFAULT_SILENCE_THRESHOLD_MS,
autoTts = prefs[KEY_AUTO_TTS] ?: DEFAULT_AUTO_TTS,
language = prefs[KEY_LANGUAGE] ?: DEFAULT_LANGUAGE,
realtimeTraceDetails = prefs[KEY_REALTIME_TRACE_DETAILS]
?: DEFAULT_REALTIME_TRACE_DETAILS,
realtimePersistentSession = prefs[KEY_REALTIME_PERSISTENT_SESSION]
?: DEFAULT_REALTIME_PERSISTENT_SESSION,
)
// In-memory active scope. Defaults to global so un-scoped consumers (and
// every existing call site) behave exactly as before until a scope is set.
private val _scope = MutableStateFlow(VoiceProfileScope.Global)
/** The active per-profile scope. Set via [setActiveScope]. */
val activeScope: StateFlow<VoiceProfileScope> = _scope.asStateFlow()
/**
* Point the repository at a (connection, profile) scope. Per-profile reads
* and writes (engine/route/enhanced) re-target the namespaced keys for that
* profile; global prefs are unaffected. Passing a null/blank profile name
* reverts per-profile reads/writes to the global base layer (the default
* profile). Idempotent — a no-op when the normalized scope is unchanged.
*/
fun setActiveScope(connectionId: String?, profileName: String?) {
val next = VoiceProfileScope(
connectionId = connectionId?.trim()?.takeIf { it.isNotEmpty() },
profileName = profileName?.trim()?.takeIf { it.isNotEmpty() },
)
if (_scope.value != next) {
_scope.value = next
}
.distinctUntilChanged()
}
/**
* Emits the resolved [VoiceSettings] for the [activeScope]. Re-emits when
* either the underlying DataStore or the active scope changes. Per-profile
* fields are resolved as: per-profile key → global key → hard default.
*/
val settings: Flow<VoiceSettings> = combine(_scope, dataStore.data) { scope, prefs ->
VoiceSettings(
// --- per-profile (override map) ---
engineMode = VoiceEngineMode.fromStorage(
resolveString(prefs, KEY_ENGINE_MODE, scope, DEFAULT_ENGINE_MODE),
).storageValue,
audioRoute = VoiceAudioRoute.fromStorage(
resolveString(prefs, KEY_AUDIO_ROUTE, scope, DEFAULT_AUDIO_ROUTE),
).storageValue,
enhancedVoice = resolveString(prefs, KEY_ENH_VOICE, scope, ""),
enhancedModel = resolveString(prefs, KEY_ENH_MODEL, scope, ""),
enhancedAudioTags = resolveBoolean(prefs, KEY_ENH_AUDIO_TAGS, scope, false),
enhancedPersona = resolveString(prefs, KEY_ENH_PERSONA, scope, ""),
enhancedLanguage = resolveString(prefs, KEY_ENH_LANGUAGE, scope, ""),
// --- global (shared across profiles) ---
interactionMode = prefs[KEY_INTERACTION_MODE] ?: DEFAULT_INTERACTION_MODE,
silenceThresholdMs = prefs[KEY_SILENCE_THRESHOLD_MS] ?: DEFAULT_SILENCE_THRESHOLD_MS,
autoTts = prefs[KEY_AUTO_TTS] ?: DEFAULT_AUTO_TTS,
language = prefs[KEY_LANGUAGE] ?: DEFAULT_LANGUAGE,
realtimeTraceDetails = prefs[KEY_REALTIME_TRACE_DETAILS]
?: DEFAULT_REALTIME_TRACE_DETAILS,
realtimePersistentSession = prefs[KEY_REALTIME_PERSISTENT_SESSION]
?: DEFAULT_REALTIME_PERSISTENT_SESSION,
)
}.distinctUntilChanged()
// --- per-profile resolution (per-profile key → global key → default) -----
private fun resolveString(
prefs: Preferences,
base: String,
scope: VoiceProfileScope,
default: String,
): String {
val scopedName = scopedName(base, scope)
if (scopedName != base) {
prefs[stringPreferencesKey(scopedName)]?.let { return it }
}
return prefs[stringPreferencesKey(base)] ?: default
}
private fun resolveBoolean(
prefs: Preferences,
base: String,
scope: VoiceProfileScope,
default: Boolean,
): Boolean {
val scopedName = scopedName(base, scope)
if (scopedName != base) {
prefs[booleanPreferencesKey(scopedName)]?.let { return it }
}
return prefs[booleanPreferencesKey(base)] ?: default
}
// --- per-profile setters (write the namespaced key for the active scope) -
suspend fun setEngineMode(mode: VoiceEngineMode) {
dataStore.edit { it[KEY_ENGINE_MODE] = mode.storageValue }
val key = stringPreferencesKey(scopedName(KEY_ENGINE_MODE, _scope.value))
dataStore.edit { it[key] = mode.storageValue }
}
suspend fun setAudioRoute(route: VoiceAudioRoute) {
dataStore.edit { it[KEY_AUDIO_ROUTE] = route.storageValue }
val key = stringPreferencesKey(scopedName(KEY_AUDIO_ROUTE, _scope.value))
dataStore.edit { it[key] = route.storageValue }
}
/** "" clears the override (relay falls back to the server's saved voice). */
suspend fun setEnhancedVoice(voice: String) {
val key = stringPreferencesKey(scopedName(KEY_ENH_VOICE, _scope.value))
dataStore.edit { it[key] = voice.trim() }
}
/** "" clears the override (relay falls back to the server's saved model). */
suspend fun setEnhancedModel(model: String) {
val key = stringPreferencesKey(scopedName(KEY_ENH_MODEL, _scope.value))
dataStore.edit { it[key] = model.trim() }
}
suspend fun setEnhancedAudioTags(enabled: Boolean) {
val key = booleanPreferencesKey(scopedName(KEY_ENH_AUDIO_TAGS, _scope.value))
dataStore.edit { it[key] = enabled }
}
/** "" clears the inline persona/style direction (Gemini). */
suspend fun setEnhancedPersona(persona: String) {
val key = stringPreferencesKey(scopedName(KEY_ENH_PERSONA, _scope.value))
dataStore.edit { it[key] = persona }
}
/** "" clears the language override (xAI). */
suspend fun setEnhancedLanguage(language: String) {
val key = stringPreferencesKey(scopedName(KEY_ENH_LANGUAGE, _scope.value))
dataStore.edit { it[key] = language.trim() }
}
// --- global setters (always the un-namespaced key) -----------------------
suspend fun setInteractionMode(mode: String) {
dataStore.edit { it[KEY_INTERACTION_MODE] = mode }
}
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network.handlers
package com.hermesandroid.relay.network.relay
import android.content.ActivityNotFoundException
import android.content.ClipData
@@ -23,9 +23,10 @@ import kotlinx.serialization.json.booleanOrNull
// === PHASE3-tier-C: flavor gate for sideload-only tools ===
import com.hermesandroid.relay.data.BuildFlavor
// === END PHASE3-tier-C ===
import com.hermesandroid.relay.network.ChannelMultiplexer
import com.hermesandroid.relay.network.RelayHttpClient
import com.hermesandroid.relay.network.models.Envelope
import com.hermesandroid.relay.network.relay.ChannelMultiplexer
import com.hermesandroid.relay.network.relay.RelayHttpClient
import com.hermesandroid.relay.network.relay.models.Envelope
import com.hermesandroid.relay.network.shared.LocalDispatchResult
import com.hermesandroid.relay.util.MediaCacheWriter
import kotlin.coroutines.AbstractCoroutineContextElement
import kotlin.coroutines.CoroutineContext
@@ -2418,33 +2419,9 @@ class BridgeCommandHandler(
}
}
/**
* Captured outcome of a local bridge dispatch. Voice mode reads this to
* emit follow-up chat traces showing the real success/failure state of
* an action after the safety modal resolves and the underlying
* [ActionExecutor] method returns. The fields mirror what the LLM path
* would see on a `bridge.response` envelope:
*
* - [status] — HTTP-style status: 200 success, 400 client error,
* 403 user denial / bridge disabled, 500 executor error
* - [errorMessage] — free-text error from the response payload, or null
* on success. Safe to speak / display verbatim to the user.
* - [errorCode] — structured classification (e.g. `permission_denied`,
* `bridge_disabled`, `user_denied`) when `respondFromResult` or a
* direct respond call includes one. Null for errors we haven't
* classified yet.
* - [resultJson] — the raw result object, for callers that need
* action-specific fields (e.g. the resolved phone number from
* /search_contacts). Optional.
*/
data class LocalDispatchResult(
val status: Int,
val errorMessage: String?,
val errorCode: String?,
val resultJson: JsonObject?,
) {
val isSuccess: Boolean get() = status in 200..299
}
// LocalDispatchResult moved to network.shared (ADR 34 fence): it is a passive
// DTO shared with the upstream chat path (ChatHandler), so it cannot live in
// this relay-package file without creating an upstream -> relay import.
/**
* Coroutine context marker installed by [BridgeCommandHandler.handleLocalCommand]
@@ -1,6 +1,6 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.relay
import com.hermesandroid.relay.network.models.Envelope
import com.hermesandroid.relay.network.relay.models.Envelope
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.relay
import android.content.Context
import android.net.ConnectivityManager
@@ -12,7 +12,8 @@ import com.hermesandroid.relay.data.PairingPreferences
import com.hermesandroid.relay.diagnostics.DiagnosticCategory
import com.hermesandroid.relay.diagnostics.DiagnosticSeverity
import com.hermesandroid.relay.diagnostics.DiagnosticsLog
import com.hermesandroid.relay.network.models.Envelope
import com.hermesandroid.relay.network.relay.models.Envelope
import com.hermesandroid.relay.network.shared.EndpointResolver
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.SupervisorJob
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.relay
import android.util.Log
import com.hermesandroid.relay.auth.PairedDeviceInfo
@@ -7,6 +7,7 @@ import com.hermesandroid.relay.diagnostics.DiagnosticSeverity
import com.hermesandroid.relay.diagnostics.DiagnosticsLog
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.Serializable
import kotlinx.serialization.builtins.ListSerializer
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
@@ -22,7 +23,7 @@ import java.io.IOException
*
* The chat SSE stream can emit tool output containing a marker of the form
* `MEDIA:hermes-relay://<opaque-token>`
* [ChatHandler][com.hermesandroid.relay.network.handlers.ChatHandler] parses
* [ChatHandler][com.hermesandroid.relay.network.upstream.ChatHandler] parses
* the marker, and [ChatViewModel][com.hermesandroid.relay.viewmodel.ChatViewModel]
* calls [fetchMedia] to pull the actual bytes over plain HTTP(S). The relay
* base URL is the WSS relay URL with `ws`/`wss` swapped for `http`/`https`.
@@ -40,7 +41,11 @@ import java.io.IOException
class RelayHttpClient(
private val okHttpClient: OkHttpClient,
private val relayUrlProvider: () -> String?,
private val sessionTokenProvider: suspend () -> String?
private val sessionTokenProvider: suspend () -> String?,
/** Synchronous snapshot of the paired session token (null when not currently
* paired). Lets [mediaUrlConfigured] check fetch-readiness without
* suspending; mirrors what [sessionTokenProvider] resolves. */
private val pairedTokenSnapshot: () -> String? = { null },
) {
companion object {
@@ -53,6 +58,19 @@ class RelayHttpClient(
}
}
/**
* True when relay media is actually FETCHABLE right now: a non-blank relay
* URL AND a current paired session token. Synchronous. The token check
* matters because the relay's SessionManager is in-memory and wiped on
* restart, so a configured relay URL can outlive the pairing — gating on URL
* alone made the media-capability badge read "available" while every
* `/media/by-path` fetch failed for a missing token. Now the badge (and the
* SSE media hint) agree with what the fetch can do, and self-correct on
* re-pair.
*/
fun mediaUrlConfigured(): Boolean =
!relayUrlProvider().isNullOrBlank() && !pairedTokenSnapshot().isNullOrBlank()
/**
* The result of a successful [fetchMedia] call.
*
@@ -61,28 +79,53 @@ class RelayHttpClient(
* @property bytes raw response body.
* @property fileName best-effort filename parsed from
* `Content-Disposition: inline; filename="..."`, or null.
* @property sensitive model-emitted sensitivity hint, read from the
* relay's `X-Media-Sensitive` response header (`"1"`/`"true"`
* → true). The relay never classifies media — it transports
* whatever the producing tool/agent declared. Absent header →
* false. Consumed by `ChatViewModel` to blur per the user's
* setting.
*/
data class FetchedMedia(
val contentType: String,
val bytes: ByteArray,
val fileName: String?
val fileName: String?,
val sensitive: Boolean = false
) {
override fun equals(other: Any?): Boolean {
if (this === other) return true
if (other !is FetchedMedia) return false
return contentType == other.contentType &&
bytes.contentEquals(other.bytes) &&
fileName == other.fileName
fileName == other.fileName &&
sensitive == other.sensitive
}
override fun hashCode(): Int {
var result = contentType.hashCode()
result = 31 * result + bytes.contentHashCode()
result = 31 * result + (fileName?.hashCode() ?: 0)
result = 31 * result + sensitive.hashCode()
return result
}
}
/**
* Server-side relay context that would be injected into the next agent turn.
* Mirrors `GET /context/injected`; Android treats it as audit-only state.
*/
@Serializable
data class InjectedContextAudit(
val enabled: Boolean = false,
val blocks: List<InjectedContextBlock> = emptyList(),
)
@Serializable
data class InjectedContextBlock(
val name: String,
val text: String,
)
/**
* Fetch `GET /media/<token>` from the relay over HTTP(S). Returns a
* [Result] — success carries a [FetchedMedia], failure wraps the
@@ -141,12 +184,16 @@ class RelayHttpClient(
response.header("Content-Disposition")
)
val sensitive = parseSensitiveHeader(
response.header("X-Media-Sensitive")
)
val body = response.body
if (body == null) {
return@withContext Result.failure(IOException("Empty response body"))
}
val bytes = body.bytes()
Result.success(FetchedMedia(contentType, bytes, fileName))
Result.success(FetchedMedia(contentType, bytes, fileName, sensitive))
}
} catch (e: IOException) {
Log.w(TAG, "fetchMedia failed for $token: ${e.message}")
@@ -248,12 +295,16 @@ class RelayHttpClient(
response.header("Content-Disposition")
)
val sensitive = parseSensitiveHeader(
response.header("X-Media-Sensitive")
)
val body = response.body
if (body == null) {
return@withContext Result.failure(IOException("Empty response body"))
}
val bytes = body.bytes()
Result.success(FetchedMedia(contentType, bytes, fileName))
Result.success(FetchedMedia(contentType, bytes, fileName, sensitive))
}
} catch (e: IOException) {
Log.w(TAG, "fetchMediaByPath failed for $path: ${e.message}")
@@ -264,6 +315,85 @@ class RelayHttpClient(
}
}
/**
* Fetch the relay's server-side injected-context audit. This endpoint is
* optional and fail-open: old/plugin-absent relays return an empty disabled
* audit rather than breaking the client-side context sheet.
*/
suspend fun fetchInjectedContext(): Result<InjectedContextAudit> = withContext(Dispatchers.IO) {
val relayUrl = relayUrlProvider()?.trim().orEmpty()
if (relayUrl.isEmpty()) {
return@withContext Result.failure(
IllegalStateException("Relay URL not configured")
)
}
val sessionToken = sessionTokenProvider()
if (sessionToken.isNullOrBlank()) {
return@withContext Result.failure(
IllegalStateException("Relay not paired — session token missing")
)
}
val httpBase = relayUrl
.replace(Regex("^wss://", RegexOption.IGNORE_CASE), "https://")
.replace(Regex("^ws://", RegexOption.IGNORE_CASE), "http://")
.trimEnd('/')
val url = try {
"$httpBase/context/injected".toHttpUrl()
} catch (e: IllegalArgumentException) {
return@withContext Result.failure(
IOException("Invalid relay URL: ${e.message}")
)
}
val request = Request.Builder()
.url(url)
.get()
.header("Authorization", "Bearer $sessionToken")
.header("Accept", "application/json")
.build()
val auditClient = okHttpClient.newBuilder()
.callTimeout(3, java.util.concurrent.TimeUnit.SECONDS)
.build()
try {
auditClient.newCall(request).execute().use { response ->
if (response.code == 404) {
return@withContext Result.success(InjectedContextAudit())
}
if (!response.isSuccessful) {
val reason = when (response.code) {
401, 403 -> "Unauthorized — re-pair with the relay"
in 500..599 -> "Relay error (HTTP ${response.code})"
else -> "HTTP ${response.code}: ${response.message.ifBlank { "request failed" }}"
}
return@withContext Result.failure(IOException(reason))
}
val body = response.body?.string().orEmpty()
if (body.isBlank()) {
return@withContext Result.failure(IOException("Empty response body"))
}
Result.success(
sessionsJson.decodeFromString(
InjectedContextAudit.serializer(),
body,
)
)
}
} catch (e: IOException) {
Log.w(TAG, "fetchInjectedContext failed: ${e.message}")
Result.failure(e)
} catch (e: Exception) {
Log.w(TAG, "fetchInjectedContext parse error: ${e.message}")
Result.failure(e)
}
}
// ------------------------------------------------------------------
// Paired-device management (2026-04-11 security overhaul)
// ------------------------------------------------------------------
@@ -777,4 +907,16 @@ class RelayHttpClient(
val match = Regex("""filename\s*=\s*"?([^";]+)"?""", RegexOption.IGNORE_CASE).find(header)
return match?.groupValues?.get(1)?.trim()?.ifBlank { null }
}
/**
* Parse the relay's `X-Media-Sensitive` response header into a bool.
*
* The relay emits the header only when the media was flagged sensitive,
* with value `"1"` (and tolerates `"true"`). Any other value — or an
* absent header — means "not sensitive", so when in doubt we don't blur.
*/
private fun parseSensitiveHeader(header: String?): Boolean {
val value = header?.trim()?.lowercase() ?: return false
return value == "1" || value == "true"
}
}
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.relay
import android.util.Log
import com.hermesandroid.relay.data.ProfileConfigResponse
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.relay
import java.net.URI
@@ -0,0 +1,24 @@
package com.hermesandroid.relay.network.relay
import com.hermesandroid.relay.data.EnhancedVoiceOverrides
import com.hermesandroid.relay.data.VoiceAudioRoute
import com.hermesandroid.relay.network.shared.VoiceAudioClient
import java.io.File
/**
* Adapts the relay-only [RelayVoiceClient] (same package) to the neutral
* [VoiceAudioClient] routing seam in `network.shared`. Relay → shared is an
* allowed dependency direction under the ADR 34 package fence.
*/
class RelayVoiceAudioClientAdapter(
private val relayVoiceClient: RelayVoiceClient,
private val enhancedOverridesProvider: () -> EnhancedVoiceOverrides? = { null },
) : VoiceAudioClient {
override val route: VoiceAudioRoute = VoiceAudioRoute.Relay
override suspend fun transcribe(audioFile: File): Result<String> =
relayVoiceClient.transcribe(audioFile)
override suspend fun synthesize(text: String): Result<File> =
relayVoiceClient.synthesize(text, enhancedOverridesProvider())
}
@@ -1,7 +1,8 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.relay
import android.content.Context
import android.util.Log
import com.hermesandroid.relay.data.EnhancedVoiceOverrides
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.data.RealtimeConversationContextMessage
import kotlinx.coroutines.CompletableDeferred
@@ -208,7 +209,10 @@ class RelayVoiceClient(
* when done (typical pattern: keep the last N mp3s in the cache dir and
* let the OS reclaim on cache pressure).
*/
suspend fun synthesize(text: String): Result<File> = withContext(Dispatchers.IO) {
suspend fun synthesize(
text: String,
enhanced: EnhancedVoiceOverrides? = null,
): Result<File> = withContext(Dispatchers.IO) {
val httpBase = resolveHttpBase()
?: return@withContext Result.failure(IllegalStateException("Relay URL not configured"))
val token = resolveBearerToken()
@@ -223,6 +227,16 @@ class RelayVoiceClient(
val bodyJson = buildJsonObject {
put("text", JsonPrimitive(text))
// Per-request enhanced-voice overrides. The relay maps these generic
// fields onto the active provider (Gemini/xAI) and ignores them for
// others — see voice.py:_extract_voice_overrides.
enhanced?.let { ov ->
ov.voice?.let { put("voice", JsonPrimitive(it)) }
ov.model?.let { put("model", JsonPrimitive(it)) }
ov.audioTags?.let { put("audio_tags", JsonPrimitive(it)) }
ov.personaPrompt?.let { put("persona_prompt", JsonPrimitive(it)) }
ov.language?.let { put("language", JsonPrimitive(it)) }
}
putProfile()
}.toString()
@@ -724,6 +738,7 @@ class RelayVoiceClient(
codec: String? = null,
optimizeStreamingLatency: Int? = null,
textNormalization: Boolean? = null,
autoSpeechTags: Boolean? = null,
fallbackEnabled: Boolean? = null,
): Result<VoiceOutputConfig> = withContext(Dispatchers.IO) {
val httpBase = resolveHttpBase()
@@ -755,6 +770,7 @@ class RelayVoiceClient(
put("optimize_streaming_latency", JsonPrimitive(it))
}
textNormalization?.let { put("text_normalization", JsonPrimitive(it)) }
autoSpeechTags?.let { put("auto_speech_tags", JsonPrimitive(it)) }
fallbackEnabled?.let { put("fallback_enabled", JsonPrimitive(it)) }
}
@@ -2304,11 +2320,44 @@ data class VoiceProviderInfo(
val voiceId: String? = null,
val enabled: Boolean = false,
val available: Boolean = true,
/**
* Provider-specific enhanced-voice capability hint. Present (non-null) only
* for the TTS block when the relay's active provider supports per-request
* enhanced control (today: Gemini and xAI).
*/
val enhanced: EnhancedVoiceCapabilities? = null,
) {
val displayVoice: String? get() = voice ?: voiceId
val isEnabled: Boolean get() = enabled || (!provider.isNullOrBlank() && available)
}
/**
* Wire shape of the `tts.enhanced` block on `GET /voice/config` — the relay's
* provider-aware enhanced-voice capability advertisement. Mirrors
* `plugin/relay/voice.py:_enhanced_voice_block`. `voices`/`models` may be empty
* (e.g. xAI uses a free-text voice field); the UI renders from the flags.
*/
@Serializable
data class EnhancedVoiceCapabilities(
val provider: String? = null,
val supported: Boolean = false,
val voices: List<String> = emptyList(),
val models: List<String> = emptyList(),
@SerialName("audio_tag_models")
val audioTagModels: List<String> = emptyList(),
@SerialName("audio_tags_enabled")
val audioTagsEnabled: Boolean = false,
@SerialName("audio_tags_label")
val audioTagsLabel: String = "Expressive tone tags",
@SerialName("supports_persona")
val supportsPersona: Boolean = false,
@SerialName("supports_language")
val supportsLanguage: Boolean = false,
@SerialName("persona_prompt_file")
val personaPromptFile: String? = null,
val overrides: List<String> = emptyList(),
)
@Serializable
data class RealtimeVoiceConfig(
val success: Boolean = false,
@@ -2481,6 +2530,8 @@ data class VoiceOutputConfig(
val codec: String = "pcm",
val optimize_streaming_latency: Int = 1,
val text_normalization: Boolean = false,
/** xAI expressive speech tags on the streaming renderer (xai_tts only). */
val auto_speech_tags: Boolean = false,
val fallback_enabled: Boolean = true,
val fallback_provider: String? = null,
val providers: List<RealtimeProviderInfo> = emptyList(),
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network.models
package com.hermesandroid.relay.network.relay.models
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.JsonArray
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.shared
import android.content.Context
import android.net.ConnectivityManager
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.shared
import android.util.Log
import com.hermesandroid.relay.data.EndpointCandidate
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.shared
import android.content.Context
import android.net.ConnectivityManager
@@ -0,0 +1,31 @@
package com.hermesandroid.relay.network.shared
import kotlinx.serialization.json.JsonObject
/**
* Transport-neutral result of a phone-control dispatch.
*
* Shared vocabulary between the relay bridge path
* ([com.hermesandroid.relay.network.relay.BridgeCommandHandler], which produces
* it) and the upstream chat path
* ([com.hermesandroid.relay.network.upstream.ChatHandler], which renders a
* phone-action bubble from it). It is a passive DTO — not a client — so it
* lives in `network.shared` to keep the upstream↔relay package fence intact
* (ADR 34); neither side depends on the other to speak it.
*
* - [status] — HTTP-style status of the dispatch (200 = ok).
* - [errorMessage] — human-readable failure text, or null on success.
* - [errorCode] — machine error code when the dispatch failed and was
* classified, or null.
* - [resultJson] — the raw result object, for callers that need
* action-specific fields (e.g. the resolved phone number from
* /search_contacts). Optional.
*/
data class LocalDispatchResult(
val status: Int,
val errorMessage: String?,
val errorCode: String?,
val resultJson: JsonObject?,
) {
val isSuccess: Boolean get() = status in 200..299
}
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.shared
import java.net.URI
@@ -0,0 +1,117 @@
package com.hermesandroid.relay.network.shared
import com.hermesandroid.relay.data.VoiceAudioRoute
import java.io.File
/**
* Transport-neutral STT/TTS contract. The routing seam between the Standard
* (dashboard) and Relay voice clients — implementations live in `network.upstream`
* (`StandardHermesVoiceClient`) and `network.relay` (`RelayVoiceAudioClientAdapter`),
* while this interface and the [AutoVoiceAudioClient] router stay dependency-neutral
* so neither voice backend leaks across the upstream/relay package fence (ADR 34).
*/
interface VoiceAudioClient {
val route: VoiceAudioRoute
/**
* The route a call would ACTUALLY use right now. For a concrete backend this
* equals [route]; for the [AutoVoiceAudioClient] router it resolves `Auto`
* against live readiness (relay-first). Callers that need to reason about
* the backend's capabilities (e.g. "is standard global-TTS in play?") must
* use this, not the configured preference.
*/
val effectiveRoute: VoiceAudioRoute
get() = route
suspend fun transcribe(audioFile: File): Result<String>
suspend fun synthesize(text: String): Result<File>
}
/**
* Routes each STT/TTS call to the Standard (dashboard) or Relay voice client.
*
* Auto preference order is **Relay first, then Standard**: a paired Relay is
* the purpose-built mobile facade — profile-aware voice config, no dashboard
* sign-in dependency — so users who installed the plugin keep the richer
* path. Standard is the zero-plugin route for vanilla Hermes installs and is
* used whenever Relay isn't configured/paired (or fails mid-call). Power
* users can force either route in Voice Settings.
*
* Depends only on the [VoiceAudioClient] abstraction (both backends are passed
* in as the interface), so this router carries no upstream or relay imports.
*/
class AutoVoiceAudioClient(
private val standardClient: VoiceAudioClient,
private val relayClient: VoiceAudioClient,
private val routeProvider: () -> VoiceAudioRoute,
private val standardReadyProvider: () -> Boolean,
private val relayReadyProvider: () -> Boolean,
) : VoiceAudioClient {
override val route: VoiceAudioRoute
get() = routeProvider()
/**
* Resolve the configured preference to the backend a call would land on:
* `Standard`/`Relay` are honored verbatim; `Auto` prefers Relay when it's
* ready (matching [runAuto]) and falls back to Standard otherwise. Used to
* decide whether standard-only limitations (global TTS) currently apply.
*/
override val effectiveRoute: VoiceAudioRoute
get() = when (routeProvider()) {
VoiceAudioRoute.Standard -> VoiceAudioRoute.Standard
VoiceAudioRoute.Relay -> VoiceAudioRoute.Relay
VoiceAudioRoute.Auto ->
if (relayReadyProvider()) VoiceAudioRoute.Relay else VoiceAudioRoute.Standard
}
override suspend fun transcribe(audioFile: File): Result<String> =
runWithSelectedRoute { it.transcribe(audioFile) }
override suspend fun synthesize(text: String): Result<File> =
runWithSelectedRoute { it.synthesize(text) }
private suspend fun <T> runWithSelectedRoute(
block: suspend (VoiceAudioClient) -> Result<T>,
): Result<T> {
return when (routeProvider()) {
VoiceAudioRoute.Standard -> {
if (!standardReadyProvider()) {
Result.failure(
IllegalStateException(
"Vanilla Hermes voice is not available — check dashboard sign-in in Manage",
),
)
} else {
block(standardClient)
}
}
VoiceAudioRoute.Relay -> {
if (!relayReadyProvider()) {
Result.failure(IllegalStateException("Relay voice is not available"))
} else {
block(relayClient)
}
}
VoiceAudioRoute.Auto -> runAuto(block)
}
}
private suspend fun <T> runAuto(
block: suspend (VoiceAudioClient) -> Result<T>,
): Result<T> {
var relayFailure: Result<T>? = null
if (relayReadyProvider()) {
val result = block(relayClient)
if (result.isSuccess || !standardReadyProvider()) return result
relayFailure = result
}
if (standardReadyProvider()) {
val result = block(standardClient)
if (result.isSuccess) return result
return relayFailure ?: result
}
return relayFailure ?: Result.failure(
IllegalStateException("Voice needs a reachable Hermes dashboard or Relay voice route"),
)
}
}
@@ -1,6 +1,7 @@
package com.hermesandroid.relay.network.handlers
package com.hermesandroid.relay.network.upstream
import android.util.Log
import com.hermesandroid.relay.data.Attachment
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.ChatSession
import com.hermesandroid.relay.data.HermesCard
@@ -8,9 +9,10 @@ import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.data.RealtimeTurnTrace
import com.hermesandroid.relay.data.ToolCall
import com.hermesandroid.relay.data.VoiceIntentTrace
import com.hermesandroid.relay.network.GatewaySubagentEvent
import com.hermesandroid.relay.network.models.MessageItem
import com.hermesandroid.relay.network.models.SessionItem
import com.hermesandroid.relay.network.shared.LocalDispatchResult
import com.hermesandroid.relay.network.upstream.GatewaySubagentEvent
import com.hermesandroid.relay.network.upstream.models.MessageItem
import com.hermesandroid.relay.network.upstream.models.SessionItem
import kotlinx.coroutines.flow.MutableStateFlow
import kotlinx.coroutines.flow.StateFlow
import kotlinx.coroutines.flow.asStateFlow
@@ -80,7 +82,11 @@ class ChatHandler {
// reachable when the tool fired, so we render an "unavailable"
// placeholder instead of attempting a fetch.
private val mediaRelayRegex = Regex("""MEDIA:hermes-relay://([A-Za-z0-9_-]+)""")
private val mediaBarePathRegex = Regex("""^\s*MEDIA:(/\S+)\s*$""")
// `/.+?` (not `/\S+`) so absolute paths containing spaces — e.g.
// `MEDIA:/mnt/media/Coralee Adshade/undressher.jpg` — still match. The
// trailing `\s*$` trims any trailing whitespace; non-greedy keeps the
// capture to the path. OkHttp re-encodes the space for /media/by-path.
private val mediaBarePathRegex = Regex("""^\s*MEDIA:(/.+?)\s*$""")
// Rich card marker — single line, full JSON object payload.
//
// Agents emit:
@@ -105,6 +111,14 @@ class ChatHandler {
/** Whether to parse tool annotations from assistant text (for servers that don't emit tool events). */
var parseToolAnnotations: Boolean = false
/**
* When false (default — TUI/desktop parity), server-injected role:system
* STEERING markers ("[System: The active model … has changed …]") are
* hidden from the rendered transcript. A developer toggle flips this to
* surface them for debugging. They always remain in server-side history.
*/
var showSystemMarkers: Boolean = false
/** Active personality/agent name — set by ChatViewModel before each stream. Included on new assistant messages. */
var activeAgentName: String? = null
@@ -189,6 +203,18 @@ class ChatHandler {
private val _messages = MutableStateFlow<List<ChatMessage>>(emptyList())
val messages: StateFlow<List<ChatMessage>> = _messages.asStateFlow()
/**
* Latest gateway `status.update` lifecycle line for the in-flight turn
* (model fallback, retries, errors). Surfaced as a transient status line
* above the composer; cleared when the turn completes.
*/
private val _turnStatus = MutableStateFlow<String?>(null)
val turnStatus: StateFlow<String?> = _turnStatus.asStateFlow()
fun setTurnStatus(text: String) {
_turnStatus.value = text
}
private val _isStreaming = MutableStateFlow(false)
val isStreaming: StateFlow<Boolean> = _isStreaming.asStateFlow()
@@ -221,6 +247,7 @@ class ChatHandler {
role = MessageRole.SYSTEM,
content = text,
timestamp = System.currentTimeMillis(),
clientOnly = true,
)
(list + notice).let { if (it.size > MAX_MESSAGES) it.drop(it.size - MAX_MESSAGES) else it }
}
@@ -229,8 +256,9 @@ class ChatHandler {
/**
* Append an assistant message that carries ONLY a gateway ask card
* (clarify / approval / sudo / secret). Local-only — the server never
* stores the ask as a message, so [loadMessageHistory] preserves the
* `ask-` id prefix the same way it preserves voice-intent traces.
* stores the ask as a message, so the bubble is flagged
* [ChatMessage.clientOnly] and [loadMessageHistory] preserves it across the
* reload the same way it preserves voice-intent traces.
* Idempotent on [messageId] so a re-emitted ask never duplicates.
*/
fun appendAskCardMessage(messageId: String, card: HermesCard) {
@@ -243,6 +271,7 @@ class ChatHandler {
timestamp = System.currentTimeMillis(),
cards = listOf(card),
agentName = activeAgentName,
clientOnly = true,
)
(list + msg).let { if (it.size > MAX_MESSAGES) it.drop(it.size - MAX_MESSAGES) else it }
}
@@ -314,6 +343,7 @@ class ChatHandler {
role = MessageRole.USER,
content = userText,
timestamp = ts,
clientOnly = true,
)
val assistantMsg = ChatMessage(
id = "voice-intent-action-$ts",
@@ -331,6 +361,7 @@ class ChatHandler {
// session memory. Null for the pre-dispatch user bubble (the
// raw transcribed utterance carries no structure on its own).
voiceIntent = voiceIntent,
clientOnly = true,
)
_messages.update { list ->
(list + userMsg + assistantMsg).let {
@@ -379,6 +410,7 @@ class ChatHandler {
// pre-dispatch trace was appended) leave this null and rely
// on the pre-dispatch trace's voiceIntent field.
voiceIntent = voiceIntent,
clientOnly = true,
)
_messages.update { list ->
(list + resultMsg).let {
@@ -451,7 +483,13 @@ class ChatHandler {
val mapped = messages.map { msg ->
if (msg.id == messageId && msg.role == MessageRole.ASSISTANT) {
changed = true
msg.copy(realtimeTurn = trace)
// A realtimeTurn trace is attached ONLY for provider-only
// (non-Hermes-backed) turns — Hermes-backed ones leave it
// null because the server owns that turn. So a trace ⟺ a
// purely local bubble: mark it clientOnly so the post-turn
// reload preserves it instead of wiping the only record of
// the turn (and the trace the next chat payload still needs).
msg.copy(realtimeTurn = trace, clientOnly = true)
} else {
msg
}
@@ -720,8 +758,19 @@ class ChatHandler {
}
/**
* Load message history from API response into the messages list.
* Replaces current messages with the loaded history.
* Reconcile the in-memory transcript against the server message history.
*
* This is a surgical DELTA-MERGE keyed by message id, not a wholesale
* replace:
* - a server message whose id matches a local row UPDATES that row's
* server-authoritative fields (content, tool calls, cards, reasoning)
* in place while keeping every client-only field;
* - a server message with no local row is INSERTED in timestamp order;
* - a client-only orphan ([ChatMessage.clientOnly], no server row) is KEPT;
* - a local row that WAS server-backed (not clientOnly) but is no longer in
* the transcript is DROPPED (genuinely deleted / forked / truncated
* server-side).
*
* Reconstructs tool calls from assistant messages' tool_calls field.
*
* Server-persisted message content still contains raw `MEDIA:...` markers
@@ -746,11 +795,69 @@ class ChatHandler {
// mutateMessage lookups find the newly-loaded messages.
val pendingMediaHits = mutableListOf<Pair<String, MediaMarkerHit>>()
// Reconcile optimistic (client-UUID) live ids to their server ids BEFORE
// building the carry map, so the id-keyed delta-merge updates rows in
// place instead of dropping-and-reinserting them. SSE assistant rows are
// already reconciled mid-turn (replaceMessageId ← message.started); but
// GATEWAY assistant rows keep a local UUID (the gateway exposes no
// per-message server id during the turn) and USER rows of every transport
// keep a local UUID. Without this, the id-keyed carry-forward below
// silently misses those rows, so a gateway turn's tokens/badges survived
// only if a content match happened to cover them. See
// [reconcileLiveIdsToServer].
val serverItemIds = items.mapNotNullTo(HashSet()) { it.id }
val idRemap = reconcileLiveIdsToServer(items, serverItemIds)
// Carry CLIENT-ONLY enrichment forward across the reload, keyed by the
// RECONCILED message id. The server transcript (MessageItem) rebuilds
// content, tool calls, and reasoning — but it does NOT persist per-message
// token usage/cost, provenance badges, tapped-card confirmations, or the
// voice/realtime sync traces. Carrying these forward is preserve-by-default
// (not a per-field whitelist the next new field forgets). With the remap
// above, an id-matched (SSE) row keeps its server id and a positionally
// reconciled (gateway/user) row adopts its server id, so both match the
// reloaded item id and update in place.
val priorById = _messages.value.associateBy { idRemap[it.id] ?: it.id }
// Outbound (user-authored) attachments are the safety net for any USER row
// that did NOT reconcile to a server id at merge time (they live on USER
// messages, are not echoed in server content, and are not re-dispatched —
// only inbound `MEDIA:` markers are). Reconciled rows already carry their
// attachment by id via priorById, so we queue ONLY the still-unreconciled
// rows here — otherwise a later same-content row could pull a duplicate
// from the queue. Match by content, consume-once so repeated identical
// sends don't cross-assign; inbound (relayToken != null) attachments are
// excluded since the marker re-dispatch re-fetches them.
val priorOutboundByContent = HashMap<String, ArrayDeque<List<Attachment>>>()
for (msg in _messages.value) {
if (msg.role != MessageRole.USER) continue
if ((idRemap[msg.id] ?: msg.id) in serverItemIds) continue // carried by id already
val outbound = msg.attachments.filter { it.relayToken == null }
if (outbound.isNotEmpty()) {
priorOutboundByContent.getOrPut(msg.content) { ArrayDeque() }.addLast(outbound)
}
}
val loaded = items.mapNotNull { item ->
val role = when (item.role) {
"user" -> MessageRole.USER
"assistant" -> MessageRole.ASSISTANT
"system" -> MessageRole.SYSTEM
"system" ->
// Upstream injects role:system STEERING markers into the
// session history on model/personality change — e.g.
// "[System: The active model for this chat has changed to …]"
// (tui_gateway/server.py) — so the LLM picks up the new
// runtime/persona. They are NOT user-facing; the TUI/desktop
// keep them invisible. Hide them for parity UNLESS the
// developer "Show system messages" toggle is on (debugging).
// They remain in server-side history for the model regardless.
if (!showSystemMarkers &&
item.contentText?.trimStart()?.startsWith("[System:") == true
) {
return@mapNotNull null
} else {
MessageRole.SYSTEM
}
"tool" -> return@mapNotNull null // Merged into assistant tool calls above
else -> return@mapNotNull null
}
@@ -788,24 +895,77 @@ class ChatHandler {
afterMedia to emptyList()
}
ChatMessage(
id = messageId,
role = role,
content = cleanedContent,
timestamp = timestampMs,
isStreaming = false,
toolCalls = toolCalls,
cards = extractedCards,
agentName = if (role == MessageRole.ASSISTANT) activeAgentName else null,
// Server persists per-message reasoning — restore it so the
// Thought-process block survives returning to the chat
// instead of existing only for the live turn.
thinkingContent = if (role == MessageRole.ASSISTANT) {
item.resolvedReasoning?.trim() ?: ""
} else {
""
},
)
val prior = priorById[messageId]
// Outbound attachments: prefer an id-match (covers any future
// user-message id reconciliation), else fall back to the
// content-keyed queue. Inbound attachments are intentionally
// excluded — they come back via the marker re-dispatch.
val carriedAttachments = run {
val byId = prior?.attachments.orEmpty().filter { it.relayToken == null }
when {
byId.isNotEmpty() -> byId
role == MessageRole.USER ->
priorOutboundByContent[cleanedContent]?.removeFirstOrNull().orEmpty()
else -> emptyList()
}
}
// Server reasoning is authoritative when present; absent, keep the
// live-streamed thinking rather than blanking it on reload.
val serverThinking =
if (role == MessageRole.ASSISTANT) item.resolvedReasoning?.trim() else null
if (prior != null) {
// DELTA-MERGE UPDATE — a local row with this id already exists.
// Refresh ONLY the server-authoritative fields (content, tool
// calls, cards, reasoning, role, timestamp) and keep every
// client-only field (tokens, cost, badges, tapped-card
// confirmations, voice/realtime traces, the clientOnly flag, …)
// by copying the existing message. This is strictly broader than
// the old curated carry list — a new client-only field is
// preserved automatically — and produces an object equal to the
// prior one when nothing server-side changed, so unchanged rows
// don't churn.
//
// `id = messageId` adopts the server id: for an id-matched (SSE)
// row it's a no-op, but for a positionally reconciled (gateway /
// user) row whose `prior` still carries a client UUID it swaps in
// the server id so EVERY future reload matches by id.
prior.copy(
id = messageId,
role = role,
content = cleanedContent,
attachments = carriedAttachments,
timestamp = timestampMs,
isStreaming = false,
isThinkingStreaming = false,
toolCalls = toolCalls,
cards = extractedCards,
agentName = if (role == MessageRole.ASSISTANT) activeAgentName else null,
thinkingContent = if (role == MessageRole.ASSISTANT) {
serverThinking ?: prior.thinkingContent
} else {
""
},
)
} else {
// INSERT — a server message with no local row yet. Built from
// server data; client-only enrichment defaults empty (there is
// nothing local to carry).
ChatMessage(
id = messageId,
role = role,
content = cleanedContent,
attachments = carriedAttachments,
timestamp = timestampMs,
isStreaming = false,
toolCalls = toolCalls,
cards = extractedCards,
agentName = if (role == MessageRole.ASSISTANT) activeAgentName else null,
// Server persists per-message reasoning — restore it so the
// Thought-process block survives returning to the chat.
thinkingContent = if (role == MessageRole.ASSISTANT) serverThinking ?: "" else "",
)
}
}
// Reload swaps the entire message list — any stale dedupe entries keyed
@@ -813,41 +973,37 @@ class ChatHandler {
// just collected against the reloaded IDs are guaranteed to fire.
dispatchedMediaMarkers.clear()
// === PHASE3-voice-intents-chathistory ===
// Preserve local-only voice-intent trace messages across a reload.
// These messages are injected by [appendLocalVoiceIntentTrace] with
// IDs prefixed "voice-intent-" and never reach the server-side
// session, so a wholesale `_messages.value = loaded` assignment
// would wipe them. Bailey hit this 2026-04-15: voice fall-through
// ("proceed" → not a recognized intent → chat.sendMessage) triggered
// a history reload on stream complete and the previous voice trace
// vanished, making it look like "the chat cleared". Server-side
// sync (so these traces reach the LLM's session memory too) is
// still a v0.4.1 follow-up, but preserving them client-side is
// enough to fix the disappearing-scrollback bug today.
// Gateway-local bubbles ride the same preservation: steered text
// (id "steer-…") lives inside a server-side tool result, never as a
// user message, and ask cards (id "ask-…") are built from gateway
// events that have no server-side message at all — a wholesale
// reload would silently erase both.
val preservedVoiceTraces = _messages.value.filter {
it.id.startsWith("voice-intent-") ||
it.id.startsWith("steer-") ||
it.id.startsWith("ask-")
}
val merged = if (preservedVoiceTraces.isEmpty()) {
// Preserve client-only orphans across the reload. A bubble flagged
// [ChatMessage.clientOnly] has no server-side row — slash-command
// notices (addSystemNotice), voice-intent traces
// (appendLocalVoiceIntent*), the steer echo, gateway ask cards
// (appendAskCardMessage), an errored turn the server never persisted
// (markError), and provider-only realtime turns (attachRealtimeTurnTrace).
// A wholesale `_messages.value = loaded` assignment would silently wipe
// any whose id is absent from the reloaded transcript: the
// disappearing-scrollback / "reply appears then vanishes" class of bug.
//
// This replaces the old id-prefix whitelist (voice-intent-/steer-/ask-/
// system-notice-) plus "Error"-badge sniffing — each creator now declares
// its own provenance, so a new client-only bubble type is preserved the
// moment it sets the flag, with no reconcile-side change. Note the badge
// subtlety: a turn that errors *after* persisting keeps an "Error" badge
// but IS in the transcript, so it reconciles normally; only clientOnly +
// absent-from-transcript marks a preservable orphan.
val loadedIds = loaded.mapTo(HashSet()) { it.id }
val preservedLocal = _messages.value.filter { it.clientOnly && it.id !in loadedIds }
val merged = if (preservedLocal.isEmpty()) {
loaded
} else {
// Merge by timestamp so voice traces interleave with the
// reloaded server messages in chronological order. The voice
// trace IDs carry `System.currentTimeMillis()` in their suffix
// (see appendLocalVoiceIntentTrace), so ChatMessage.timestamp
// is the source of truth here.
(loaded + preservedVoiceTraces).sortedBy { it.timestamp }
// Merge by timestamp so preserved orphans interleave with the
// reloaded server messages in chronological order. Voice trace IDs
// carry `System.currentTimeMillis()` in their suffix (see
// appendLocalVoiceIntentTrace); other orphans keep their live
// ChatMessage.timestamp — the source of truth either way.
(loaded + preservedLocal).sortedBy { it.timestamp }
}
_messages.value = if (merged.size > MAX_MESSAGES) merged.takeLast(MAX_MESSAGES) else merged
// === END PHASE3-voice-intents-chathistory ===
// Now that the reloaded messages are in state, fire callbacks so the
// ViewModel can insert LOADING/FAILED attachments via mutateMessage.
@@ -871,6 +1027,105 @@ class ChatHandler {
}
}
/** One adoptable server row during id reconciliation. `taken` enforces consume-once. */
private class ReconcileSlot(
val serverId: String,
val role: MessageRole,
val key: String,
) {
var taken: Boolean = false
}
/**
* Map optimistic (client-UUID) live message ids → their server ids so the
* id-keyed delta-merge in [loadMessageHistory] updates rows in place.
*
* Strategy: match each still-unreconciled, non-[ChatMessage.clientOnly] live
* row to an unclaimed server row by (role, marker-stripped content),
* consume-once in document order, and adopt the server id. Rows whose id is
* already a server id (SSE assistant, reconciled mid-turn) are skipped so we
* never double-swap; clientOnly orphans have no server row and are never
* mapped; a live row that matches no slot is left alone (graceful fallback —
* the content-keyed attachment fallback and drop-and-reinsert still apply, so
* a count divergence / truncation / compaction never forces a wrong map).
*
* Returns oldLiveId → serverId for the rows that reconciled (empty when there
* is nothing to adopt).
*/
private fun reconcileLiveIdsToServer(
items: List<MessageItem>,
serverItemIds: Set<String>,
): Map<String, String> {
val live = _messages.value
if (live.isEmpty()) return emptyMap()
val liveIds = live.mapTo(HashSet()) { it.id }
// Adoptable slots: rendered server rows (not tool, not a hidden steering
// marker) whose id no live row already carries.
val slots = items.mapNotNull { item ->
val serverId = item.id ?: return@mapNotNull null
if (serverId in liveIds) return@mapNotNull null
val role = renderedRoleOf(item) ?: return@mapNotNull null
ReconcileSlot(serverId, role, reconcileKey(item.contentText))
}
if (slots.isEmpty()) return emptyMap()
val remap = HashMap<String, String>()
for (msg in live) {
if (msg.clientOnly) continue // no server row to adopt
if (msg.id in serverItemIds) continue // already a server id
val key = reconcileKey(msg.content)
val slot = slots.firstOrNull { !it.taken && it.role == msg.role && it.key == key }
?: continue
slot.taken = true
remap[msg.id] = slot.serverId
}
return remap
}
/**
* The rendered role for a server message item, or null for rows that never
* become a visible bubble (tool results, unknown roles, and hidden
* `[System:` steering markers). Mirrors the role/skip logic in
* [loadMessageHistory] so reconciliation only adopts ids onto rows that
* actually render.
*/
private fun renderedRoleOf(item: MessageItem): MessageRole? = when (item.role) {
"user" -> MessageRole.USER
"assistant" -> MessageRole.ASSISTANT
"system" ->
if (!showSystemMarkers &&
item.contentText?.trimStart()?.startsWith("[System:") == true
) {
null
} else {
MessageRole.SYSTEM
}
else -> null
}
/**
* Normalize content for reconciliation matching: strip `MEDIA:`/`CARD:` marker
* lines (raw server content still carries them; live content already had them
* stripped during streaming) and collapse to trimmed, non-blank lines joined
* by newlines. Lets a marker-bearing assistant turn match its live row while
* staying byte-stable for plain user/assistant text.
*/
private fun reconcileKey(content: String?): String {
val text = content.orEmpty()
if (text.isEmpty()) return ""
val sb = StringBuilder()
for (line in text.lines()) {
val t = line.trim()
if (t.isEmpty()) continue
if (mediaRelayRegex.containsMatchIn(t) || mediaBarePathRegex.containsMatchIn(t)) continue
if (cardMarkerRegex.containsMatchIn(t)) continue
if (sb.isNotEmpty()) sb.append('\n')
sb.append(t)
}
return sb.toString()
}
/**
* Marker hit collected during [loadMessageHistory] for post-assignment dispatch.
*/
@@ -2083,12 +2338,57 @@ class ChatHandler {
// Note: do NOT set _isStreaming to false — the run is still active
}
/**
* Stamp a "Stopped" badge on a message whose turn the user cancelled, so
* the bubble carries a persistent status (not just a transient toast).
* No-op if already present. Call before [onStreamComplete] on cancel.
*/
fun markStopped(messageId: String) {
_messages.update { messages ->
messages.map { msg ->
if (msg.id == messageId && "Stopped" !in msg.badges) {
msg.copy(badges = msg.badges + "Stopped")
} else {
msg
}
}
}
}
/**
* Stamp an "Error" badge on a message whose turn ended in a server error
* (e.g. a gateway ❌ lifecycle status), so a failed turn doesn't read as a
* normal answer. No-op if already present.
*
* Also marks the bubble [ChatMessage.clientOnly] = true: the only caller is
* the gateway ❌ terminal-error path, which fires on an in-flight turn the
* server never persists. That makes this assistant bubble a client-only
* orphan, so the reload must preserve it (the "reply appears then vanishes"
* regression). If the same id later turns up in the server transcript (a
* turn that errored *after* persisting), the reload reconciles it as a
* normal server-backed message and the Error badge rides along via the
* priorById carry — the clientOnly flag only gates the not-in-transcript
* orphan case.
*/
fun markError(messageId: String) {
_messages.update { messages ->
messages.map { msg ->
if (msg.id == messageId && "Error" !in msg.badges) {
msg.copy(badges = msg.badges + "Error", clientOnly = true)
} else {
msg
}
}
}
}
/**
* The entire agent run is complete (run.completed / done).
* Marks the stream as finished and finalizes all messages.
*/
fun onStreamComplete(messageId: String) {
_isStreaming.value = false
_turnStatus.value = null
insideThinkingBlock = false
// Flush any remaining annotation text that didn't end with a newline
@@ -2249,7 +2549,7 @@ class ChatHandler {
*
* The label parameter is the short human-readable action name
* ("Send SMS", "Open App", "Call", etc). Error-code branches mirror the
* `error_code` strings [com.hermesandroid.relay.network.handlers.BridgeCommandHandler]
* `error_code` strings [com.hermesandroid.relay.network.relay.BridgeCommandHandler]
* emits on destructive-verb rejections.
*/
internal fun formatPhoneActionResult(
@@ -1,14 +1,14 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import android.content.Context
import com.hermesandroid.relay.data.Profile
import com.hermesandroid.relay.network.models.MessageItem
import com.hermesandroid.relay.network.models.MessageListResponse
import com.hermesandroid.relay.network.models.SessionItem
import com.hermesandroid.relay.network.models.SessionListResponse
import com.hermesandroid.relay.auth.KeystoreTokenStore
import com.hermesandroid.relay.auth.LegacyEncryptedPrefsTokenStore
import com.hermesandroid.relay.network.upstream.models.MessageItem
import com.hermesandroid.relay.network.upstream.models.MessageListResponse
import com.hermesandroid.relay.network.upstream.models.SessionItem
import com.hermesandroid.relay.network.upstream.models.SessionListResponse
import com.hermesandroid.relay.auth.SecureStoreCache
import com.hermesandroid.relay.auth.SessionTokenStore
import com.hermesandroid.relay.auth.buildRawTokenStore
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.Serializable
@@ -520,15 +520,23 @@ class DashboardApiClient(
* an auth-gated 401/403 also proves the route is registered.
*/
suspend fun audioRoutesPresent(): Boolean = withContext(Dispatchers.IO) {
val request = Request.Builder()
.url("$baseUrl/api/audio/transcribe")
.head()
.build()
try {
okHttpClient.newCall(request).execute().use { it.code != 404 }
} catch (_: Exception) {
false
// Route exists if HEAD returns anything but a clean 404:
// - 405 Method Not Allowed: path registered, POST-only (FastAPI/Starlette)
// - 401/403: registered but auth-gated
// - 2xx: handled
// A reverse proxy fronting the dashboard can rewrite a 405 into a 404,
// which would read as absent. To cut that false-negative, probe BOTH
// audio routes and treat the surface as present if EITHER answers
// non-404 (they ship together upstream, so one reachable implies both).
fun probe(path: String): Boolean {
val request = Request.Builder().url("$baseUrl$path").head().build()
return try {
okHttpClient.newCall(request).execute().use { it.code != 404 }
} catch (_: Exception) {
false
}
}
probe("/api/audio/transcribe") || probe("/api/audio/speak")
}
suspend fun requestWsTicket(): Result<DashboardWsTicket> = withContext(Dispatchers.IO) {
@@ -808,22 +816,56 @@ class InMemoryDashboardCookieStore : DashboardCookieStore {
class EncryptedDashboardCookieStore(
context: Context,
connectionId: String,
/**
* The connection's TOKEN-store file key. When non-null the dashboard cookies
* ride that already-built keyset (so there is NO second keyset build on cold
* start), and any cookies in this connection's old stand-alone
* `hermes_dashboard_<id>` file are migrated across once. Null preserves the
* original stand-alone-file behavior for callers that can't resolve the key.
*/
tokenStoreKey: String? = null,
private val json: Json = Json { ignoreUnknownKeys = true },
) : DashboardCookieStore {
private val serializer = ListSerializer(StoredDashboardCookie.serializer())
private val appContext = context.applicationContext
private val prefsName = prefsName(connectionId)
private val standaloneCookiePrefsName = prefsName(connectionId)
// Unify onto the connection's token file when we know it; else stand alone.
// (Explicit type + distinct name avoids a type-inference cycle with the
// companion `prefsName(connectionId)` function above.)
private val storePrefsName: String = tokenStoreKey ?: standaloneCookiePrefsName
private val unified = tokenStoreKey != null && tokenStoreKey != standaloneCookiePrefsName
// DEFERRED on purpose. Building the Keystore-backed prefs takes 1-4s
// on StrongBox devices and serializes through a process-GLOBAL Tink
// lock (AndroidKeysetManager.Builder.build) — eager construction here
// froze the main thread for ~11s at app start when several stores were
// built concurrently (frozen-sphere incident, 2026-06-11). Construction
// is now free on any thread; the expensive build happens on the first
// actual cookie access, which is always an OkHttp/IO thread.
// DEFERRED on purpose. Building the Keystore-backed prefs takes 1-4s on
// StrongBox devices and serializes through a process-GLOBAL Tink lock
// (AndroidKeysetManager.Builder.build) — eager construction here froze the
// main thread for ~11s at app start (frozen-sphere incident, 2026-06-11).
// Construction is free on any thread; the expensive build happens on the
// first actual cookie access, always an OkHttp/IO thread. Going through
// SecureStoreCache means that build is SHARED with the connection's token
// store — so when unified there is NO second keyset build at all.
private val store: SessionTokenStore by lazy {
KeystoreTokenStore.tryCreate(appContext, prefsName)
?: LegacyEncryptedPrefsTokenStore(appContext, prefsName)
val s = SecureStoreCache.getOrBuild(storePrefsName) {
buildRawTokenStore(appContext, storePrefsName)
}
if (unified) migrateCookiesFromStandaloneFile(s)
s
}
/**
* One-shot copy of this connection's cookies from the old stand-alone
* `hermes_dashboard_<id>` file into the unified token file, marker-gated so
* the old file's keyset is built at most once ever. On failure (corrupt old
* file) the user simply re-signs-in to Manage — cookies are re-obtainable,
* unlike the relay session token.
*/
private fun migrateCookiesFromStandaloneFile(target: SessionTokenStore) {
if (target.contains(KEY_COOKIES_MIGRATED)) return
runCatching {
val old = buildRawTokenStore(appContext, standaloneCookiePrefsName)
old.getString(KEY_COOKIES)?.let { target.putString(KEY_COOKIES, it) }
old.clearAll()
}
target.putString(KEY_COOKIES_MIGRATED, "1")
}
override fun load(): List<StoredDashboardCookie> {
@@ -842,6 +884,7 @@ class EncryptedDashboardCookieStore(
companion object {
private const val KEY_COOKIES = "dashboard_cookies_json"
private const val KEY_COOKIES_MIGRATED = "dashboard_cookies_migrated"
fun prefsName(connectionId: String): String =
"hermes_dashboard_${connectionId.take(8)}"
@@ -864,8 +907,11 @@ class DashboardCookieJar(
override fun loadForRequest(url: HttpUrl): List<Cookie> {
val now = clockMillis()
val stored = store.load().filterNot { it.isExpired(now) }
if (stored.size != store.load().size) {
// Load once (each load() is a decrypt + JSON decode); prune expired
// entries back to disk only when something actually expired.
val all = store.load()
val stored = all.filterNot { it.isExpired(now) }
if (stored.size != all.size) {
store.save(stored)
}
return stored.mapNotNull { it.toCookie() }
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import android.util.Log
import com.hermesandroid.relay.util.AppForegroundTracker
@@ -188,6 +188,68 @@ class GatewayChatClient(
private val _connectionState = MutableStateFlow(GatewayConnectionState.Idle)
val connectionState: StateFlow<GatewayConnectionState> = _connectionState.asStateFlow()
/**
* Active personality the gateway is applying, as a config value ("none" when
* the overlay is cleared, otherwise the personality name). Tracks the
* upstream `display.personality` the way the desktop/TUI do: updated from the
* [setPersonality] / [getPersonality] round-trips AND from connection-level
* `session.info` events, so a change made via `/personality`, the desktop, or
* the TUI reflects in the app. Null until first observed.
*/
private val _serverPersonality = MutableStateFlow<String?>(null)
val serverPersonality: StateFlow<String?> = _serverPersonality.asStateFlow()
/**
* Active model / provider the gateway reports for our session, tracked off
* `session.info` the same way as [serverPersonality]. Lets a `/model` switch
* made on the desktop/TUI (or our own dispatch) reflect in the app's model
* pill without an app reload. Null until first observed; only ever set to a
* non-blank value.
*/
private val _serverModel = MutableStateFlow<String?>(null)
val serverModel: StateFlow<String?> = _serverModel.asStateFlow()
private val _serverProvider = MutableStateFlow<String?>(null)
val serverProvider: StateFlow<String?> = _serverProvider.asStateFlow()
/**
* Active reasoning EFFORT from `session.info` (string; "" when reasoning is
* disabled). The reasoning DISPLAY mode is NOT on session.info — it stays a
* `config.get reasoning` concern ([getReasoningSettings]). Only ever set to a
* non-blank value so a disabled-reasoning "" never clobbers the chip.
*/
private val _serverReasoningEffort = MutableStateFlow<String?>(null)
val serverReasoningEffort: StateFlow<String?> = _serverReasoningEffort.asStateFlow()
/**
* Server-reported credential warning (upstream `session.info.credential_warning`)
* — present ONLY when the active provider's key is missing/invalid, absent
* (→ null here) when healthy. Cleared on absence so it self-resolves when the
* key is fixed.
*/
private val _serverCredentialWarning = MutableStateFlow<String?>(null)
val serverCredentialWarning: StateFlow<String?> = _serverCredentialWarning.asStateFlow()
/**
* Effective approval-bypass (YOLO) + fast-mode state from `session.info`
* (`yolo`/`fast` booleans). YOLO has NO `config.get` upstream — session.info
* is the only read. Null until first observed.
*/
private val _serverYolo = MutableStateFlow<Boolean?>(null)
val serverYolo: StateFlow<Boolean?> = _serverYolo.asStateFlow()
private val _serverFast = MutableStateFlow<Boolean?>(null)
val serverFast: StateFlow<Boolean?> = _serverFast.asStateFlow()
/**
* Context-window usage `(used, max)` from `session.info`'s `usage` block
* (upstream `_get_usage`). `session.info` is emitted on session resume, so
* this lets the context bar paint immediately on resume instead of waiting
* for the first turn's usage event. Null until observed / when omitted.
*/
private val _serverContext = MutableStateFlow<Pair<Int, Int>?>(null)
val serverContext: StateFlow<Pair<Int, Int>?> = _serverContext.asStateFlow()
/** Serializes connect / session-establish so concurrent sends share one socket. */
private val connectMutex = Mutex()
@@ -223,6 +285,24 @@ class GatewayChatClient(
private fun currentSessionProfile(): String? =
sessionProfileProvider().takeIf { !it.isNullOrBlank() }
/**
* Supplies the explicit in-chat overrides to bind onto each fresh
* `session.create` (upstream honors `model`/`provider`/`reasoning_effort`/
* `fast` → the new session's per-session overrides). Pulled live so it
* always reflects the current picker + safety/speed controls; null (or all
* fields null) = no explicit override, so the new session inherits the
* profile / server default. Wired by ChatViewModel. A live session keeps its
* agent config, so this only affects session creation — mid-session switches
* go through [setModel]/[setReasoning]/[setFast] (`config.set`).
*/
@Volatile
var sessionModelProvider: () -> GatewaySessionModel? = { null }
private fun currentSessionModel(): GatewaySessionModel? =
sessionModelProvider()?.takeIf {
!it.model.isNullOrBlank() || !it.reasoningEffort.isNullOrBlank() || it.fast != null
}
@Volatile
private var activeTurn: GatewayTurn? = null
@@ -418,16 +498,31 @@ class GatewayChatClient(
* send.
*/
fun prewarm(storedSessionId: String?) {
scope.launch {
try {
connectMutex.withLock {
ensureConnected()
if (storedSessionId != null) resumeForPrewarm(storedSessionId)
}
} catch (e: Exception) {
Log.d(TAG, "Gateway prewarm skipped: ${e.message}")
scope.launch { prewarmAwait(storedSessionId) }
}
/**
* Suspending [prewarm]: establishes the socket and (when [storedSessionId]
* is non-null) resumes the existing session, returning only once that work
* has settled. Returns true when a live session is available afterwards.
*
* An in-chat model/effort/fast switch MUST await this before its
* `config.set`. Otherwise the switch races the fire-and-forget [prewarm]
* and runs with `liveSessionId == null`, which upstream applies as a GLOBAL
* config write instead of a per-session one — so the pick never lands on
* the session the next turn actually uses (a fresh chat pre-creates a
* session, so this path is the common case, not the edge case).
*/
suspend fun prewarmAwait(storedSessionId: String?): Boolean {
try {
connectMutex.withLock {
ensureConnected()
if (storedSessionId != null) resumeForPrewarm(storedSessionId)
}
} catch (e: Exception) {
Log.d(TAG, "Gateway prewarm skipped: ${e.message}")
}
return liveSessionId != null
}
/**
@@ -565,6 +660,57 @@ class GatewayChatClient(
},
)
/**
* Read the active personality (`config.get {key:"personality"}`). Returns the
* upstream config value — `"none"` when the overlay is cleared, otherwise the
* personality name. Connects on demand. Used to seed [serverPersonality] when
* a gateway connection comes up so the app reflects whatever the server
* (config / desktop / TUI) currently has active.
*/
suspend fun getPersonality(): Result<String> {
if (webSocket == null || readySignal?.isCompleted != true) {
try {
connectMutex.withLock { ensureConnected() }
} catch (e: Exception) {
return Result.failure(e)
}
}
return rpc("config.get", buildJsonObject { put("key", "personality") })
.map { result ->
(result.stringField("value") ?: "none").ifBlank { "none" }
.also { _serverPersonality.value = it }
}
}
/**
* Set the personality the way the desktop + TUI do (`config.set
* {key:"personality"}`). The gateway persists `display.personality` +
* `agent.system_prompt` to the active profile's config AND applies the
* overlay live to the current session (no history reset). Pass `"none"`
* (or `"default"`/`"neutral"`) to clear the overlay. Returns the resolved
* active value (`"none"` or the name); also updates [serverPersonality]
* directly so observers don't have to wait on the `session.info` echo (which
* only fires when a live session exists).
*/
suspend fun setPersonality(value: String): Result<String> {
if (webSocket == null || readySignal?.isCompleted != true) {
try {
connectMutex.withLock { ensureConnected() }
} catch (e: Exception) {
return Result.failure(e)
}
}
val params = buildJsonObject {
put("key", "personality")
put("value", value)
liveSessionId?.let { put("session_id", it) }
}
return rpc("config.set", params).map { result ->
(result.stringField("value") ?: value).ifBlank { "none" }
.also { _serverPersonality.value = it }
}
}
/**
* Fetch the curated provider/model list (`model.options`) — the same RPC
* the upstream desktop + TUI model picker uses (grok / kimi / gpt-5.5 …,
@@ -592,6 +738,11 @@ class GatewayChatClient(
.mapNotNull { (it as? JsonPrimitive)?.contentOrNull },
isCurrent = (obj["is_current"] as? JsonPrimitive)?.booleanOrNull ?: false,
warning = obj.stringField("warning"),
authenticated = (obj["authenticated"] as? JsonPrimitive)?.booleanOrNull ?: true,
unavailableModels = (obj["unavailable_models"] as? JsonArray).orEmpty()
.mapNotNull { (it as? JsonPrimitive)?.contentOrNull },
freeTier = (obj["free_tier"] as? JsonPrimitive)?.booleanOrNull ?: false,
totalModels = (obj["total_models"] as? JsonPrimitive)?.contentOrNull?.toIntOrNull() ?: 0,
)
}
GatewayModelOptions(
@@ -659,6 +810,58 @@ class GatewayChatClient(
},
)
/**
* Toggle per-session approval bypass (YOLO) via `config.set {key:"yolo"}` —
* the same session-scoped flag the desktop's setSessionYolo and the TUI's
* Shift+Tab use (`value` "1"/"0", `scope` "session" = ephemeral, never writes
* config.yaml). Requires a live session for the per-session flag. Updates
* [serverYolo] from the echo so observers don't wait on `session.info`.
* Returns the resolved enabled state. There is deliberately NO `getYolo()` —
* upstream has no `config.get yolo`; session.info is the only read.
*/
suspend fun setYolo(enabled: Boolean, scope: String = "session"): Result<Boolean> {
if (webSocket == null || readySignal?.isCompleted != true) {
try {
connectMutex.withLock { ensureConnected() }
} catch (e: Exception) {
return Result.failure(e)
}
}
val params = buildJsonObject {
put("key", "yolo")
put("value", if (enabled) "1" else "0")
put("scope", scope)
liveSessionId?.let { put("session_id", it) }
}
return rpc("config.set", params).map { result ->
(result.stringField("value") == "1").also { _serverYolo.value = it }
}
}
/**
* Toggle fast mode (priority service tier) via `config.set {key:"fast"}` —
* desktop parity (`value` "fast"/"normal", session-scoped). Capability-gated
* upstream: enabling fails (error 4002) when the current model has no fast
* tier. Updates [serverFast]; returns the resolved enabled state.
*/
suspend fun setFast(enabled: Boolean): Result<Boolean> {
if (webSocket == null || readySignal?.isCompleted != true) {
try {
connectMutex.withLock { ensureConnected() }
} catch (e: Exception) {
return Result.failure(e)
}
}
val params = buildJsonObject {
put("key", "fast")
put("value", if (enabled) "fast" else "normal")
liveSessionId?.let { put("session_id", it) }
}
return rpc("config.set", params).map { result ->
(result.stringField("value") == "fast").also { _serverFast.value = it }
}
}
fun shutdown() {
activeTurn?.cancel()
activeTurn = null
@@ -757,10 +960,52 @@ class GatewayChatClient(
currentSessionProfile()?.let { put("profile", it) }
},
)
val live = resumed.getOrNull()?.stringField("session_id")
val result = resumed.getOrNull()
val live = result?.stringField("session_id")
if (live != null) {
liveSessionId = live
storedSessionId = storedId
// Paint the session's real model/provider/effort/etc NOW from the
// resume result's embedded `info` (same shape session.info carries),
// so a reopened session shows its ACTUAL model immediately instead of
// a misleading default until the first turn's async session.info.
(result["info"] as? JsonObject)?.let { applySessionInfo(it) }
}
}
/**
* Apply connection-level session info (model / provider / reasoning effort /
* personality / yolo / fast / context usage) into the `_server*` state flows.
* Shared by the `session.info` event handler and the `session.resume` RPC
* result — the resume response embeds the same `info` object, so reopening a
* session can paint its real model up front rather than waiting for a turn.
*/
private fun applySessionInfo(info: JsonObject) {
if (info.containsKey("personality")) {
_serverPersonality.value =
(info.stringField("personality") ?: "").ifBlank { "none" }
}
info.stringField("model")?.takeIf { it.isNotBlank() }?.let { _serverModel.value = it }
info.stringField("provider")?.takeIf { it.isNotBlank() }?.let { _serverProvider.value = it }
// reasoning effort: ignore "" (reasoning disabled) so it can't clobber
// the chip; display mode is config.get-only, not here.
info.stringField("reasoning_effort")?.takeIf { it.isNotBlank() }
?.let { _serverReasoningEffort.value = it }
// credential_warning: present only when the provider key is missing/
// invalid. ABSENT means healthy — clear to null so it self-resolves.
_serverCredentialWarning.value =
info.stringField("credential_warning")?.takeIf { it.isNotBlank() }
(info["yolo"] as? JsonPrimitive)?.booleanOrNull?.let { _serverYolo.value = it }
(info["fast"] as? JsonPrimitive)?.booleanOrNull?.let { _serverFast.value = it }
// Context usage: require used > 0 — a COLD resume resets counters and
// reports 0 until the first turn rebuilds the prompt; painting 0 would
// mislead on a session that actually has history.
(info["usage"] as? JsonObject)?.let { usage ->
val used = (usage["context_used"] as? JsonPrimitive)?.intOrNull
val max = (usage["context_max"] as? JsonPrimitive)?.intOrNull
if (used != null && used > 0 && max != null && max > 0) {
_serverContext.value = used to max
}
}
}
@@ -783,10 +1028,12 @@ class GatewayChatClient(
currentSessionProfile()?.let { put("profile", it) }
},
)
val live = resumed.getOrNull()?.stringField("session_id")
val result = resumed.getOrNull()
val live = result?.stringField("session_id")
if (live != null) {
liveSessionId = live
storedSessionId = requestedStoredId
(result["info"] as? JsonObject)?.let { applySessionInfo(it) }
return
}
Log.w(
@@ -802,6 +1049,23 @@ class GatewayChatClient(
put("cols", DEFAULT_COLS)
if (!newSessionTitle.isNullOrBlank()) put("title", newSessionTitle)
currentSessionProfile()?.let { put("profile", it) }
// Bind the in-chat overrides to the new session as its
// per-session overrides. Upstream tui_gateway session.create
// reads `model`/`provider` (→ model_override), `reasoning_effort`
// (→ create_reasoning_override) and `fast` (→ priority service
// tier) — verified server.py:4175-4191. Without this a fresh
// chat ignores the picker/safety controls and builds the agent
// from the global default, and worse, setting effort/fast before
// the first message runs a SESSIONLESS config.set that upstream
// applies as a GLOBAL config write. A live session keeps its own
// config — this is create-only; mid-session switches use
// config.set (setModel/setReasoning/setFast).
currentSessionModel()?.let { sm ->
sm.model?.takeIf { it.isNotBlank() }?.let { put("model", it) }
sm.provider?.takeIf { it.isNotBlank() }?.let { put("provider", it) }
sm.reasoningEffort?.takeIf { it.isNotBlank() }?.let { put("reasoning_effort", it) }
sm.fast?.let { put("fast", it) }
}
},
).getOrElse { e ->
throw GatewayPreflightException("session.create failed: ${e.message}")
@@ -902,6 +1166,20 @@ class GatewayChatClient(
return
}
// `session.info` is connection-level (personality / model / context
// usage), emitted on a config change even with no turn in flight. Capture
// the active personality here — for our own session only — so a
// `/personality`, desktop, or TUI change keeps the app in sync. Falls
// through to the turn dispatch below so an in-flight turn still sees it.
if (type == "session.info" &&
(eventSessionId == null || liveSessionId == null || eventSessionId == liveSessionId)
) {
// Connection-level session info (model / provider / effort / persona /
// yolo / fast / usage) — shared with the session.resume result via
// applySessionInfo so both paths stay in lockstep.
payload?.let { applySessionInfo(it) }
}
val turn = activeTurn ?: return
// Foreign-session events (another client's chat on the same gateway) are not ours.
if (eventSessionId != null && liveSessionId != null && eventSessionId != liveSessionId) {
@@ -1263,6 +1541,13 @@ class GatewayChatClient(
onToolGenerating = { v -> callbackDispatcher { callbacks.onToolGenerating(v) } },
onSubagentEvent = { v -> callbackDispatcher { callbacks.onSubagentEvent(v) } },
onInteractionRequest = { v -> callbackDispatcher { callbacks.onInteractionRequest(v) } },
// MUST be wrapped like every other member: GatewayTurnCallbacks gives
// onStatusUpdate a default no-op, so omitting it here silently swallows
// EVERY gateway status line — the ❌ terminal-error lifecycle update
// included. Without it markError never fires, the turn isn't badged
// "Error", and onComplete's history reload wipes the error bubble (the
// "reply appears then vanishes" bug).
onStatusUpdate = { kind, text -> callbackDispatcher { callbacks.onStatusUpdate(kind, text) } },
)
}
@@ -1,6 +1,6 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import com.hermesandroid.relay.network.models.UsageInfo
import com.hermesandroid.relay.network.upstream.models.UsageInfo
import kotlinx.serialization.json.JsonArray
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.JsonPrimitive
@@ -217,8 +217,15 @@ class GatewayEventMapper(private val callbacks: GatewayTurnCallbacks) {
),
)
// Known-but-unrendered (notification.show, status.update, …) and
// unknown types alike: ignore.
"status.update" -> {
val text = payload.string("text")
if (!text.isNullOrBlank()) {
callbacks.onStatusUpdate(payload.string("kind"), text)
}
}
// Known-but-unrendered (notification.show, …) and unknown types
// alike: ignore.
else -> Unit
}
}
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import android.annotation.SuppressLint
import android.app.NotificationChannel
@@ -1,6 +1,6 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import com.hermesandroid.relay.network.models.UsageInfo
import com.hermesandroid.relay.network.upstream.models.UsageInfo
/**
* Shared types for the Gateway chat transport — upstream hermes-agent's
@@ -146,6 +146,14 @@ data class GatewayModelProvider(
val models: List<String>,
val isCurrent: Boolean,
val warning: String?,
// Picker hints from upstream `model.options` (build_models_payload,
// picker_hints=True). Default to "usable" so older servers that omit them
// don't gray everything out.
val authenticated: Boolean = true,
/** Paid models the current account can't pick (free-tier / no credits). */
val unavailableModels: List<String> = emptyList(),
val freeTier: Boolean = false,
val totalModels: Int = 0,
)
/** Result of the gateway `model.options` RPC. */
@@ -155,6 +163,33 @@ data class GatewayModelOptions(
val currentProvider: String,
)
/**
* The explicit in-chat overrides to bind onto a gateway `session.create` as the
* new session's PER-SESSION overrides. Matches the upstream desktop client,
* whose `session.create` carries `model`/`provider`/`reasoning_effort`/`fast`
* (tui_gateway honors them → `session_model_override` / `create_reasoning_override`
* / `create_service_tier_override`; verified `tui_gateway/server.py:4175-4191`).
* Supplied live by ChatViewModel from the picker + safety/speed controls.
*
* Every field is nullable = "no explicit override for this new chat", so the
* fresh session inherits the profile / server default rather than the picker
* (or a stale local value) silently clobbering it. Crucially this keeps these
* picks OFF the sessionless `config.set` path, which upstream applies as GLOBAL
* writes (and `yolo` even leaks to other sessions via `os.environ`).
*
* [model] is the model id (e.g. `grok-4.3`); [provider] is the authenticated
* provider slug (e.g. `xai`). [reasoningEffort] is the upstream effort string
* (`low`/`medium`/`high`/…). [fast] pins the priority service tier when true.
* Note `yolo` is intentionally absent — upstream `session.create` does NOT
* accept it as a per-session override, so it is applied post-create instead.
*/
data class GatewaySessionModel(
val model: String?,
val provider: String?,
val reasoningEffort: String? = null,
val fast: Boolean? = null,
)
/** Result of the gateway `config.get {key:"reasoning"}` RPC. */
data class GatewayReasoningSettings(
val effort: String,
@@ -196,4 +231,10 @@ class GatewayTurnCallbacks(
* cancelled.
*/
val onInteractionRequest: (GatewayAsk) -> Unit,
/**
* Gateway `status.update` lifecycle line — model fallback, retries, and
* errors (often emoji-prefixed: 🔄 fallback, ⏳ retry, ❌ error). Default
* no-op so non-gateway/legacy constructors don't need to provide it.
*/
val onStatusUpdate: (kind: String?, text: String) -> Unit = { _, _ -> },
)
@@ -1,21 +1,21 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import android.os.Handler
import android.os.Looper
import android.util.Log
import com.hermesandroid.relay.data.AgentDisplay
import com.hermesandroid.relay.data.AppAnalytics
import com.hermesandroid.relay.network.models.CreateSessionRequest
import com.hermesandroid.relay.network.models.HermesSseEvent
import com.hermesandroid.relay.network.models.MessageItem
import com.hermesandroid.relay.network.models.MessageListResponse
import com.hermesandroid.relay.network.models.RenameSessionRequest
import com.hermesandroid.relay.network.models.SessionItem
import com.hermesandroid.relay.network.models.SessionListResponse
import com.hermesandroid.relay.network.models.SessionResponse
import com.hermesandroid.relay.network.models.SkillInfo
import com.hermesandroid.relay.network.models.SkillListResponse
import com.hermesandroid.relay.network.models.UsageInfo
import com.hermesandroid.relay.network.upstream.models.CreateSessionRequest
import com.hermesandroid.relay.network.upstream.models.HermesSseEvent
import com.hermesandroid.relay.network.upstream.models.MessageItem
import com.hermesandroid.relay.network.upstream.models.MessageListResponse
import com.hermesandroid.relay.network.upstream.models.RenameSessionRequest
import com.hermesandroid.relay.network.upstream.models.SessionItem
import com.hermesandroid.relay.network.upstream.models.SessionListResponse
import com.hermesandroid.relay.network.upstream.models.SessionResponse
import com.hermesandroid.relay.network.upstream.models.SkillInfo
import com.hermesandroid.relay.network.upstream.models.SkillListResponse
import com.hermesandroid.relay.network.upstream.models.UsageInfo
import com.hermesandroid.relay.util.TurnLatencyTracer
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import com.hermesandroid.relay.data.AgentDisplay
import com.hermesandroid.relay.data.Attachment
@@ -1,10 +1,10 @@
package com.hermesandroid.relay.network
package com.hermesandroid.relay.network.upstream
import android.content.Context
import com.hermesandroid.relay.data.VoiceAudioRoute
import com.hermesandroid.relay.network.shared.VoiceAudioClient
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.encodeToString
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.JsonPrimitive
@@ -21,102 +21,12 @@ import java.io.IOException
import java.util.Base64
import java.util.concurrent.TimeUnit
interface VoiceAudioClient {
val route: VoiceAudioRoute
suspend fun transcribe(audioFile: File): Result<String>
suspend fun synthesize(text: String): Result<File>
}
class RelayVoiceAudioClientAdapter(
private val relayVoiceClient: RelayVoiceClient,
) : VoiceAudioClient {
override val route: VoiceAudioRoute = VoiceAudioRoute.Relay
override suspend fun transcribe(audioFile: File): Result<String> =
relayVoiceClient.transcribe(audioFile)
override suspend fun synthesize(text: String): Result<File> =
relayVoiceClient.synthesize(text)
}
/**
* Routes each STT/TTS call to the Standard (dashboard) or Relay voice client.
*
* Auto preference order is **Relay first, then Standard**: a paired Relay is
* the purpose-built mobile facade — profile-aware voice config, no dashboard
* sign-in dependency — so users who installed the plugin keep the richer
* path. Standard is the zero-plugin route for vanilla Hermes installs and is
* used whenever Relay isn't configured/paired (or fails mid-call). Power
* users can force either route in Voice Settings.
*/
class AutoVoiceAudioClient(
private val standardClient: VoiceAudioClient,
private val relayClient: VoiceAudioClient,
private val routeProvider: () -> VoiceAudioRoute,
private val standardReadyProvider: () -> Boolean,
private val relayReadyProvider: () -> Boolean,
) : VoiceAudioClient {
override val route: VoiceAudioRoute
get() = routeProvider()
override suspend fun transcribe(audioFile: File): Result<String> =
runWithSelectedRoute { it.transcribe(audioFile) }
override suspend fun synthesize(text: String): Result<File> =
runWithSelectedRoute { it.synthesize(text) }
private suspend fun <T> runWithSelectedRoute(
block: suspend (VoiceAudioClient) -> Result<T>,
): Result<T> {
return when (routeProvider()) {
VoiceAudioRoute.Standard -> {
if (!standardReadyProvider()) {
Result.failure(
IllegalStateException(
"Standard Hermes voice is not available — check dashboard sign-in in Manage",
),
)
} else {
block(standardClient)
}
}
VoiceAudioRoute.Relay -> {
if (!relayReadyProvider()) {
Result.failure(IllegalStateException("Relay voice is not available"))
} else {
block(relayClient)
}
}
VoiceAudioRoute.Auto -> runAuto(block)
}
}
private suspend fun <T> runAuto(
block: suspend (VoiceAudioClient) -> Result<T>,
): Result<T> {
var relayFailure: Result<T>? = null
if (relayReadyProvider()) {
val result = block(relayClient)
if (result.isSuccess || !standardReadyProvider()) return result
relayFailure = result
}
if (standardReadyProvider()) {
val result = block(standardClient)
if (result.isSuccess) return result
return relayFailure ?: result
}
return relayFailure ?: Result.failure(
IllegalStateException("Voice needs a reachable Hermes dashboard or Relay voice route"),
)
}
}
/**
* Standard (no-plugin) voice client — speaks the upstream **dashboard web
* server** contract that hermes-desktop's voice mode uses:
*
* POST {dashboard}/api/audio/transcribe {data_url, mime_type} → {ok, transcript}
* POST {dashboard}/api/audio/speak {text} → {ok, data_url, mime_type}
* POST {dashboard}/api/audio/transcribe {data_url, mime_type} to {ok, transcript}
* POST {dashboard}/api/audio/speak {text} to {ok, data_url, mime_type}
*
* These routes live on `hermes_cli/web_server.py` (:9119 by convention), NOT
* on the API server (:8642) — current upstream api_server advertises
@@ -124,13 +34,19 @@ class AutoVoiceAudioClient(
* cookie session (gated_auth_middleware), so [okHttpClient] must carry the
* same per-connection cookie jar the Manage tab signs in with; an API bearer
* header is meaningless on this surface. Revisit when upstream PR #8199
* lands the `/v1/audio` routes on the API server (docs/upstream-contributions.md §6).
* lands the `/v1/audio` routes on the API server (docs/upstream-contributions.md section 6).
* (No glob spellings in block comments — Kotlin block comments nest.)
*/
class StandardHermesVoiceClient(
private val context: Context,
private val okHttpClient: OkHttpClient,
private val dashboardUrlProvider: () -> String?,
// Active chat profile name (null = default/launch). Sent DEFENSIVELY on
// /api/audio/speak: upstream `TTSSpeakRequest` is text-only and Pydantic
// ignores extra fields, so this is harmless today and forward-compatible if
// upstream ever adds profile-aware TTS. Until then, standard voice remains
// the host's global TTS (see VoiceViewModel's standard-voice profile notice).
private val profileProvider: () -> String? = { null },
private val json: Json = Json {
ignoreUnknownKeys = true
isLenient = true
@@ -150,6 +66,14 @@ class StandardHermesVoiceClient(
if (!audioFile.exists() || audioFile.length() == 0L) {
return@withContext Result.failure(IOException("Audio file missing or empty: ${audioFile.name}"))
}
// Upstream caps decoded transcription audio at 25 MB (web_server.py
// _MAX_TRANSCRIPTION_UPLOAD_BYTES → HTTP 413). The decoded size equals
// the file size, so guard here to avoid a wasted ~33 MB base64 upload.
if (audioFile.length() > MAX_TRANSCRIBE_BYTES) {
return@withContext Result.failure(
IOException("Recording too long for Hermes - try a shorter utterance"),
)
}
val dataUrl = buildAudioDataUrl(audioFile)
val payload = buildJsonObject {
@@ -181,7 +105,12 @@ class StandardHermesVoiceClient(
return@withContext Result.failure(IllegalArgumentException("Cannot synthesize blank text"))
}
val payload = buildJsonObject { put("text", cleanText) }
val payload = buildJsonObject {
put("text", cleanText)
// Defensive only — upstream /api/audio/speak ignores it (text-only
// TTSSpeakRequest). Omitted for the default profile.
profileProvider()?.trim()?.takeIf { it.isNotBlank() }?.let { put("profile", it) }
}
val request = Request.Builder()
.url("$baseUrl/api/audio/speak")
.post(json.encodeToString(JsonObject.serializer(), payload).toRequestBody(JSON_MEDIA))
@@ -241,8 +170,10 @@ class StandardHermesVoiceClient(
val body = runCatching { response.body.string() }.getOrDefault("")
val detail = body.takeIf { it.isNotBlank() } ?: response.message
val message = when (response.code) {
400 -> "$operation rejected that input - ${detail.ifBlank { "bad request" }}"
401, 403 -> "$operation needs dashboard sign-in - open Manage to sign in"
404 -> "$operation unavailable on this Hermes build - update hermes-agent or use Relay"
413 -> "Recording too long for Hermes - try a shorter utterance"
in 500..599 -> "$operation failed - server error HTTP ${response.code}"
else -> "$operation failed - HTTP ${response.code}: $detail"
}
@@ -292,5 +223,9 @@ class StandardHermesVoiceClient(
private companion object {
val JSON_MEDIA = "application/json".toMediaType()
// Matches upstream _MAX_TRANSCRIPTION_UPLOAD_BYTES (web_server.py): the
// dashboard rejects decoded transcription audio above 25 MB with 413.
const val MAX_TRANSCRIBE_BYTES = 25L * 1024 * 1024
}
}
@@ -1,4 +1,4 @@
package com.hermesandroid.relay.network.models
package com.hermesandroid.relay.network.upstream.models
import kotlinx.serialization.ExperimentalSerializationApi
import kotlinx.serialization.KSerializer
@@ -6,8 +6,8 @@ import android.provider.Settings
import android.service.notification.NotificationListenerService
import android.service.notification.StatusBarNotification
import android.util.Log
import com.hermesandroid.relay.network.ChannelMultiplexer
import com.hermesandroid.relay.network.models.Envelope
import com.hermesandroid.relay.network.relay.ChannelMultiplexer
import com.hermesandroid.relay.network.relay.models.Envelope
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
@@ -0,0 +1,86 @@
package com.hermesandroid.relay.permissions
import android.Manifest
import android.content.ComponentName
import android.content.Context
import android.content.pm.PackageManager
import android.os.Build
import android.provider.Settings
import androidx.core.content.ContextCompat
import com.hermesandroid.relay.accessibility.HermesAccessibilityService
import com.hermesandroid.relay.accessibility.MediaProjectionHolder
/**
* One snapshot of Android grants and special-access switches that Hermes-Relay
* features can consume. Standard Chat and Manage do not need dangerous runtime
* permissions, so they are intentionally not represented as a "required"
* Android grant here.
*/
data class AppPermissionStatus(
val notificationsPermitted: Boolean = false,
val microphonePermitted: Boolean = false,
val cameraPermitted: Boolean = false,
val notificationListenerPermitted: Boolean = false,
val accessibilityServiceEnabled: Boolean = false,
val screenCapturePermitted: Boolean = false,
val overlayPermitted: Boolean = false,
val contactsPermitted: Boolean = false,
val smsPermitted: Boolean = false,
val phonePermitted: Boolean = false,
val locationPermitted: Boolean = false,
)
object AppPermissionStatusProbe {
fun snapshot(context: Context): AppPermissionStatus {
val appContext = context.applicationContext
return AppPermissionStatus(
notificationsPermitted = hasPostNotifications(appContext),
microphonePermitted = hasPermission(appContext, Manifest.permission.RECORD_AUDIO),
cameraPermitted = hasPermission(appContext, Manifest.permission.CAMERA),
notificationListenerPermitted = isNotificationListenerEnabled(appContext),
accessibilityServiceEnabled = isAccessibilityServiceEnabled(appContext),
screenCapturePermitted = MediaProjectionHolder.projection != null,
overlayPermitted = Settings.canDrawOverlays(appContext),
contactsPermitted = hasPermission(appContext, Manifest.permission.READ_CONTACTS),
smsPermitted = hasPermission(appContext, Manifest.permission.SEND_SMS),
phonePermitted = hasPermission(appContext, Manifest.permission.CALL_PHONE),
locationPermitted = hasPermission(appContext, Manifest.permission.ACCESS_FINE_LOCATION),
)
}
private fun hasPostNotifications(context: Context): Boolean {
return if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.TIRAMISU) {
hasPermission(context, Manifest.permission.POST_NOTIFICATIONS)
} else {
true
}
}
private fun hasPermission(context: Context, permission: String): Boolean {
return ContextCompat.checkSelfPermission(
context,
permission,
) == PackageManager.PERMISSION_GRANTED
}
private fun isAccessibilityServiceEnabled(context: Context): Boolean {
val enabled = Settings.Secure.getString(
context.contentResolver,
Settings.Secure.ENABLED_ACCESSIBILITY_SERVICES,
) ?: return false
val expected = ComponentName(
context.packageName,
HermesAccessibilityService::class.java.name,
).flattenToString()
return enabled.split(':').any { it.equals(expected, ignoreCase = true) } ||
enabled.contains(context.packageName, ignoreCase = true)
}
private fun isNotificationListenerEnabled(context: Context): Boolean {
val enabled = Settings.Secure.getString(
context.contentResolver,
"enabled_notification_listeners",
) ?: return false
return enabled.contains(context.packageName, ignoreCase = true)
}
}
@@ -43,6 +43,7 @@ import androidx.compose.runtime.LaunchedEffect
import androidx.compose.runtime.collectAsState
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.produceState
import androidx.compose.runtime.remember
import androidx.compose.runtime.rememberUpdatedState
import androidx.compose.runtime.rememberCoroutineScope
@@ -66,14 +67,31 @@ import androidx.navigation.compose.composable
import androidx.navigation.compose.currentBackStackEntryAsState
import androidx.navigation.compose.rememberNavController
import androidx.navigation.navArgument
import com.hermesandroid.relay.ui.components.MorphingSphere
import com.hermesandroid.relay.ui.components.CrashReportGate
import com.hermesandroid.relay.ui.components.LocalAgentIconPath
import com.hermesandroid.relay.ui.components.LocalAvailableSphereSkins
import com.hermesandroid.relay.ui.components.LocalSphereSkin
import com.hermesandroid.relay.ui.components.SphereRegistry
import com.hermesandroid.relay.ui.components.SphereSkinLoader
import com.hermesandroid.relay.ui.components.SphereState
import com.hermesandroid.relay.ui.components.avatar.AgentAvatar
import com.hermesandroid.relay.ui.components.avatar.AvatarRenderState
import com.hermesandroid.relay.ui.components.avatar.LocalAgentAvatar
import com.hermesandroid.relay.ui.components.avatar.LocalAvailableAvatars
import com.hermesandroid.relay.ui.components.avatar.LocalPetPlaybackSpeed
import com.hermesandroid.relay.ui.components.avatar.LocalPetStabilize
import com.hermesandroid.relay.ui.components.avatar.PetLoader
import com.hermesandroid.relay.ui.components.avatar.SphereAvatar
import com.hermesandroid.relay.ui.components.ConnectionStatusToast
import com.hermesandroid.relay.ui.components.ConnectionSwitcherSheet
import com.hermesandroid.relay.ui.components.ChatTransportStatusBadge
import com.hermesandroid.relay.ui.components.ChatTransportTier
import com.hermesandroid.relay.ui.components.PowerFeatureGateScreen
import com.hermesandroid.relay.ui.components.PowerFeatureGateStatus
import com.hermesandroid.relay.ui.components.RelayStatusStrip
import com.hermesandroid.relay.ui.components.UnattendedGlobalBanner
import com.hermesandroid.relay.ui.components.UpdateBanner
import com.hermesandroid.relay.ui.components.resolveChatTransportStatus
import com.hermesandroid.relay.update.UpdateCheckResult
import com.hermesandroid.relay.viewmodel.UpdateViewModel
import com.hermesandroid.relay.ui.components.WhatsNewDialog
@@ -81,13 +99,16 @@ import com.hermesandroid.relay.data.AgentDisplay
import com.hermesandroid.relay.data.BridgePreferencesRepository
import com.hermesandroid.relay.data.BridgeSafetyPreferencesRepository
import com.hermesandroid.relay.data.BuildFlavor
import com.hermesandroid.relay.data.EnhancedVoiceOverrides
import com.hermesandroid.relay.data.VoiceAudioRoute
import com.hermesandroid.relay.data.VoicePreferencesRepository
import com.hermesandroid.relay.data.VoiceSettings
import com.hermesandroid.relay.data.displayLabel
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.flow.map
import kotlinx.coroutines.flow.mapNotNull
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
import com.hermesandroid.relay.util.HumanError
import kotlinx.coroutines.delay
import com.hermesandroid.relay.ui.onboarding.OnboardingScreen
@@ -106,6 +127,7 @@ import com.hermesandroid.relay.ui.screens.DeveloperSettingsScreen
import com.hermesandroid.relay.ui.screens.MediaSettingsScreen
import com.hermesandroid.relay.ui.screens.PairedDevicesScreen
import com.hermesandroid.relay.ui.screens.ConnectionsSettingsScreen
import com.hermesandroid.relay.ui.screens.PermissionsStatusScreen
import com.hermesandroid.relay.ui.screens.ProfileInspectorScreen
import com.hermesandroid.relay.ui.screens.RealtimeVoiceTestScreen
import com.hermesandroid.relay.ui.screens.SettingsScreen
@@ -113,16 +135,17 @@ import com.hermesandroid.relay.ui.screens.TerminalScreen
import com.hermesandroid.relay.ui.screens.NotificationCompanionSettingsScreen
import com.hermesandroid.relay.ui.screens.VoiceSettingsScreen
import com.hermesandroid.relay.ui.screens.prewarmDashboardManage
import com.hermesandroid.relay.ui.theme.AppThemes
import com.hermesandroid.relay.ui.theme.HermesRelayTheme
import com.hermesandroid.relay.ui.theme.RelayRefresh
import com.hermesandroid.relay.ui.theme.relayGridTexture
import com.hermesandroid.relay.diagnostics.DiagnosticCategory
import com.hermesandroid.relay.diagnostics.DiagnosticSeverity
import com.hermesandroid.relay.diagnostics.DiagnosticsLog
import com.hermesandroid.relay.network.RelayProfileInspectorClient
import com.hermesandroid.relay.network.AutoVoiceAudioClient
import com.hermesandroid.relay.network.DynamicDashboardCookieJar
import com.hermesandroid.relay.network.RelayVoiceAudioClientAdapter
import com.hermesandroid.relay.network.relay.RelayProfileInspectorClient
import com.hermesandroid.relay.network.shared.AutoVoiceAudioClient
import com.hermesandroid.relay.network.upstream.DynamicDashboardCookieJar
import com.hermesandroid.relay.network.relay.RelayVoiceAudioClientAdapter
import com.hermesandroid.relay.viewmodel.ChatViewModel
import com.hermesandroid.relay.viewmodel.ConnectionViewModel
import com.hermesandroid.relay.viewmodel.ProfileInspectorViewModel
@@ -132,8 +155,8 @@ import com.hermesandroid.relay.audio.VoicePlayer
import com.hermesandroid.relay.audio.VoiceRecorder
import com.hermesandroid.relay.audio.VoiceSfxPlayer
import com.hermesandroid.relay.audio.RealtimePcmPlayer
import com.hermesandroid.relay.network.RelayVoiceClient
import com.hermesandroid.relay.network.StandardHermesVoiceClient
import com.hermesandroid.relay.network.relay.RelayVoiceClient
import com.hermesandroid.relay.network.upstream.StandardHermesVoiceClient
import com.hermesandroid.relay.auth.AuthState
import androidx.lifecycle.viewModelScope
@@ -233,6 +256,7 @@ sealed class Screen(
data object NotificationCompanionSettings :
Screen("settings/notifications", "Notification companion", Icons.Filled.Settings)
// === END PHASE3-notif-listener-followup ===
data object PermissionsSettings : Screen("settings/permissions", "Permissions", Icons.Filled.Settings)
// === PHASE3-safety-rails: bridge safety route ===
data object BridgeSafetySettings :
Screen("settings/bridge_safety", "Bridge safety", Icons.Filled.Settings)
@@ -387,6 +411,9 @@ fun RelayApp() {
val relayVoiceReady by connectionViewModel.relayVoiceReady.collectAsState()
val standardVoiceReadyState = rememberUpdatedState(standardVoiceReady)
val relayVoiceReadyState = rememberUpdatedState(relayVoiceReady)
// Latest enhanced-voice overrides (null when nothing is set). Read lazily by
// the relay TTS adapter so changes apply without rebuilding it.
val enhancedOverridesState = rememberUpdatedState(EnhancedVoiceOverrides.fromSettings(voiceSettings))
// Voice pipeline wiring — mirrors ChatViewModel.initializeMedia (above).
// We build a dedicated OkHttpClient so voice requests don't contend with
@@ -433,12 +460,21 @@ fun RelayApp() {
.connectTimeout(15, java.util.concurrent.TimeUnit.SECONDS)
.build(),
dashboardUrlProvider = { connectionViewModel.activeDashboardUrl() },
// Live read (null for the default profile) — sent defensively on
// /api/audio/speak; upstream ignores it, so standard voice stays the
// host's global TTS. Same live source the relay voice client uses.
profileProvider = {
AgentDisplay.profileRequestName(connectionViewModel.selectedProfile.value?.name)
},
)
}
val voiceAudioClient = remember {
AutoVoiceAudioClient(
standardClient = standardVoiceClient,
relayClient = RelayVoiceAudioClientAdapter(voiceClient),
relayClient = RelayVoiceAudioClientAdapter(
voiceClient,
enhancedOverridesProvider = { enhancedOverridesState.value },
),
routeProvider = { selectedAudioRouteState.value },
standardReadyProvider = { standardVoiceReadyState.value },
relayReadyProvider = { relayVoiceReadyState.value },
@@ -692,6 +728,12 @@ fun RelayApp() {
connectionViewModel.chatHandler.parseToolAnnotations = parseAnnotations
}
// Sync "show system messages" debug toggle to ChatHandler
val showSystemMessages by connectionViewModel.showSystemMessages.collectAsState()
LaunchedEffect(showSystemMessages) {
connectionViewModel.chatHandler.showSystemMarkers = showSystemMessages
}
// Sync streaming endpoint preference to chat. Resolves "auto" against the
// current server capabilities so vanilla upstream + bootstrap-injected
// sessions API picks /v1/chat/completions for portable SSE chat while
@@ -732,9 +774,74 @@ fun RelayApp() {
// Observe theme preference
val themePreference by connectionViewModel.theme.collectAsState()
val appThemeId by connectionViewModel.appTheme.collectAsState()
val fontScale by connectionViewModel.fontScale.collectAsState()
HermesRelayTheme(themePreference = themePreference, fontScale = fontScale) {
// Resolve the active sphere skin (built-in / adaptive / user-loaded) and
// publish it + the full available set so every MorphingSphere picks it up
// via LocalSphereSkin without per-call-site threading. Adaptive skins read
// the brand lazily inside MorphingSphere, so this can sit outside the theme.
val sphereSkinId by connectionViewModel.sphereSkin.collectAsState()
val sphereContext = androidx.compose.ui.platform.LocalContext.current
val availableSphereSkins by produceState(
initialValue = SphereRegistry.builtIns,
key1 = sphereContext,
) {
value = SphereRegistry.builtIns +
withContext(Dispatchers.IO) { SphereSkinLoader.loadUserSkins(sphereContext) }
}
val activeSphereSkin = remember(sphereSkinId, appThemeId, availableSphereSkins) {
SphereRegistry.resolve(
selectedId = sphereSkinId,
themeDefaultSkinId = AppThemes.byId(appThemeId).defaultSphereSkinId,
available = availableSphereSkins,
)
}
// Agent avatar seam (P2/P3): the built-in sphere plus any user-loaded "pets"
// (P3). The sphere nests the skin system one level below (avatar → skin).
// Published beside the skin locals so every avatar call site resolves it via
// LocalAgentAvatar without per-call-site threading. An unknown selected id
// (e.g. a pet pack was removed) falls back to the sphere.
val agentAvatarId by connectionViewModel.agentAvatar.collectAsState()
// Re-scans the pets/ dir whenever the tick bumps (in-app import/delete, or the
// Appearance screen opening), so newly added/removed pets appear everywhere
// without an app restart.
val avatarsRefreshTick by connectionViewModel.avatarsRefreshTick.collectAsState()
val availableAgentAvatars by produceState(
initialValue = listOf<AgentAvatar>(SphereAvatar),
key1 = sphereContext,
key2 = avatarsRefreshTick,
) {
value = listOf<AgentAvatar>(SphereAvatar) +
withContext(Dispatchers.IO) { PetLoader.loadPets(sphereContext) }
}
val activeAgentAvatar = remember(agentAvatarId, availableAgentAvatars) {
availableAgentAvatars.firstOrNull { it.id == agentAvatarId } ?: SphereAvatar
}
val petSpeed by connectionViewModel.petSpeed.collectAsState()
val petStabilize by connectionViewModel.petStabilize.collectAsState()
val agentIconPath by connectionViewModel.profileIcon.collectAsState()
CompositionLocalProvider(
LocalSphereSkin provides activeSphereSkin,
LocalAvailableSphereSkins provides availableSphereSkins,
LocalAgentAvatar provides activeAgentAvatar,
LocalAvailableAvatars provides availableAgentAvatars,
LocalPetPlaybackSpeed provides petSpeed,
LocalPetStabilize provides petStabilize,
LocalAgentIconPath provides agentIconPath,
) {
HermesRelayTheme(
appThemeId = appThemeId,
themePreference = themePreference,
fontScale = fontScale,
) {
// Surface a crash report from a previous session, if any. Renders a
// platform Dialog (own window) so tree position is z-order-agnostic;
// it just needs to be inside the theme for Material colors.
CrashReportGate()
val navController = rememberNavController()
var postOnboardingRoute by remember { mutableStateOf<String?>(null) }
@@ -963,7 +1070,7 @@ fun RelayApp() {
// (or there was none), and the checklist has visibly
// finished ticking. Anything weaker (e.g. the resolver's
// earlier health evidence) reveals a chat screen that still
// shows "Connect Standard Hermes" for the few hundred ms
// shows "Connect Vanilla Hermes" for the few hundred ms
// until the client-based verdict catches up.
(chatReady && initialChatSettled && startupNarrationComplete) ||
// Error path: a settled unreachable reveals the normal UI,
@@ -1232,19 +1339,19 @@ fun RelayApp() {
snackbarHost = { SnackbarHost(snackbarHostState) },
bottomBar = {
if (!suppressGlobalChrome && !isKeyboardVisible && !showStartupSphere && !voiceUiState.voiceMode) {
val leading = when {
apiReachable -> "api online"
relayReady -> "relay connected"
else -> "offline"
}
val leadingColor = when {
apiReachable -> RelayRefresh.Green
relayReady -> RelayRefresh.Relay
else -> RelayRefresh.Danger
}
val routeLabel = activeEndpoint?.displayLabel()
?: activeConnection?.label
?: "no route"
val transportStatus = resolveChatTransportStatus(
streamingEndpoint = streamingEndpoint,
gatewayAvailability = gatewayAvailability,
serverCapabilities = serverCapabilities,
)
val transportRouteLabel = if (transportStatus.tier == ChatTransportTier.Offline) {
""
} else {
routeLabel
}
val profileLabel = selectedProfile?.name?.takeIf { it.isNotBlank() } ?: "default"
val displayProfile = AgentDisplay.effectiveDisplayProfile(
selectedProfile = selectedProfile,
@@ -1259,10 +1366,24 @@ fun RelayApp() {
} else {
"profile: $profileLabel"
}
val openConnections = {
navController.navigate(Screen.ConnectionsSettings.route) {
launchSingleTop = true
}
}
RelayStatusStrip(
leading = "$leading / $routeLabel",
leadingBadge = {
ChatTransportStatusBadge(
status = transportStatus,
onClick = openConnections,
)
},
routeLabel = transportRouteLabel,
trailing = "$modelLabel / $safetyLabel",
leadingColor = leadingColor,
// Tap the persistent status/route readout to open
// Connections — preserves the affordance the dropped
// header endpoint chip used to provide.
onClick = openConnections,
)
}
}
@@ -1311,7 +1432,10 @@ fun RelayApp() {
onManageSignIn = {
postOnboardingRoute = Screen.Manage.route
connectionViewModel.completeOnboarding()
}
},
onOpenPermissions = {
navController.navigate(Screen.PermissionsSettings.route)
},
)
}
composable(
@@ -1397,6 +1521,11 @@ fun RelayApp() {
launchSingleTop = true
}
},
onNavigateToVoiceSettings = {
navController.navigate(Screen.VoiceSettings.route) {
launchSingleTop = true
}
},
onNavigateToProfileInspector = { profileName ->
navController.navigate(Screen.ProfileInspector.route(profileName)) {
launchSingleTop = true
@@ -1613,6 +1742,9 @@ fun RelayApp() {
onNavigateToNotificationCompanion = {
navController.navigate(Screen.NotificationCompanionSettings.route)
},
onNavigateToPermissions = {
navController.navigate(Screen.PermissionsSettings.route)
},
// === PHASE3-safety-rails: bridge safety route ===
onNavigateToBridgeSafety = {
navController.navigate(Screen.BridgeSafetySettings.route)
@@ -1663,6 +1795,16 @@ fun RelayApp() {
)
}
// === END PHASE3-notif-listener-followup ===
composable(Screen.PermissionsSettings.route) {
PermissionsStatusScreen(
onBack = { navController.popBackStack() },
onOpenBridge = {
navController.navigate(Screen.Bridge.route) {
launchSingleTop = true
}
},
)
}
// === PHASE3-safety-rails: bridge safety route ===
composable(Screen.BridgeSafetySettings.route) {
if (BuildFlavor.isSideload) {
@@ -2099,8 +2241,12 @@ fun RelayApp() {
.relayGridTexture(alpha = 0.14f),
contentAlignment = Alignment.Center
) {
// Sphere fills background
MorphingSphere(modifier = Modifier.fillMaxSize())
// Avatar fills background (sphere by default; routed through the
// seam so a future pet appears on the startup screen too).
LocalAgentAvatar.current.Render(
state = AvatarRenderState(state = SphereState.Idle),
modifier = Modifier.fillMaxSize(),
)
// Branding overlaid at bottom third
Column(
@@ -2160,6 +2306,7 @@ fun RelayApp() {
}
} // end Box
}
} // end CompositionLocalProvider (sphere skin)
}
/** One line of the startup sphere's progress narration. */
@@ -2,8 +2,10 @@ package com.hermesandroid.relay.ui.components
import android.content.ClipData
import android.widget.Toast
import androidx.compose.foundation.background
import androidx.compose.foundation.clickable
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.PaddingValues
import androidx.compose.foundation.layout.Row
@@ -48,6 +50,7 @@ import androidx.compose.runtime.saveable.rememberSaveable
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.clip
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.platform.ClipEntry
import androidx.compose.ui.platform.LocalClipboard
@@ -56,15 +59,19 @@ import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.text.input.ImeAction
import androidx.compose.ui.text.input.PasswordVisualTransformation
import androidx.compose.ui.text.input.VisualTransformation
import androidx.compose.ui.text.style.TextOverflow
import androidx.compose.ui.unit.dp
import com.hermesandroid.relay.auth.AuthState
import com.hermesandroid.relay.network.ConnectionState
import com.hermesandroid.relay.network.RelayUrlDeriver
import com.hermesandroid.relay.data.EndpointCandidate
import com.hermesandroid.relay.data.hasSecureProxy
import com.hermesandroid.relay.network.relay.ConnectionState
import com.hermesandroid.relay.network.relay.RelayUrlDeriver
import com.hermesandroid.relay.ui.LocalSnackbarHost
import com.hermesandroid.relay.ui.showHumanError
import com.hermesandroid.relay.util.classifyError
import com.hermesandroid.relay.viewmodel.ConnectionViewModel
import com.hermesandroid.relay.viewmodel.RelayUiState
import com.hermesandroid.relay.viewmodel.StandardVoiceAvailability
import com.hermesandroid.relay.viewmodel.asBadgeState
import com.hermesandroid.relay.viewmodel.statusText
import kotlinx.coroutines.flow.first
@@ -211,6 +218,248 @@ fun ActiveCardRelayStatusSection(
)
}
/**
* Capability overview for the active connection. Features are intentionally
* separate from routes: users can see what Hermes can do without reading the
* selected network path as the feature boundary.
*/
@Composable
fun ActiveCardFeaturesSection(
connectionViewModel: ConnectionViewModel,
relayEnabled: Boolean,
onOpenApiInfo: () -> Unit,
onOpenDashboard: () -> Unit,
onOpenRelayInfo: () -> Unit,
onOpenSessionInfo: () -> Unit,
) {
val apiReachable by connectionViewModel.apiServerReachable.collectAsState()
val apiHealth by connectionViewModel.apiServerHealth.collectAsState()
val activeConnection by connectionViewModel.activeConnection.collectAsState()
val standardVoiceAvailability by
connectionViewModel.standardVoiceAvailability.collectAsState()
val relayConfigured by connectionViewModel.relayConfigured.collectAsState()
val relayReady by connectionViewModel.relayReady.collectAsState()
val relayUiState by connectionViewModel.relayUiState.collectAsState()
val authState by connectionViewModel.authState.collectAsState()
val dashboardStatus = activeConnection?.dashboardLastStatus
val dashboardSignInRequired =
dashboardStatus?.authRequired == true && dashboardStatus.authenticated != true
val secureProxyAdvertised =
activeConnection?.routeCandidates.orEmpty().any { it.hasSecureProxy() }
val apiValue = when {
apiHealth == ConnectionViewModel.HealthStatus.Probing -> "Checking"
apiReachable -> "Ready"
activeConnection?.apiServerUrl.isNullOrBlank() -> "Missing"
else -> "Offline"
}
val apiTone = when (apiValue) {
"Ready" -> CapabilityTone.Good
"Offline", "Missing" -> CapabilityTone.Warning
else -> CapabilityTone.Neutral
}
val dashboardValue = when {
activeConnection?.resolvedDashboardUrl.isNullOrBlank() -> "Missing"
dashboardStatus == null -> "Unchecked"
!dashboardStatus.reachable -> "Offline"
dashboardSignInRequired -> "Sign in"
dashboardStatus.authenticated == true -> "Signed in"
else -> "Available"
}
val dashboardTone = when (dashboardValue) {
"Signed in", "Available" -> CapabilityTone.Good
"Sign in" -> CapabilityTone.Info
"Offline", "Missing" -> CapabilityTone.Warning
else -> CapabilityTone.Neutral
}
val voiceValue = when (standardVoiceAvailability) {
StandardVoiceAvailability.Ready -> "Ready"
StandardVoiceAvailability.SignInRequired -> "Sign in"
StandardVoiceAvailability.Unsupported -> "Unsupported"
StandardVoiceAvailability.Unreachable -> "Offline"
StandardVoiceAvailability.Unknown -> "Checking"
}
val voiceTone = when (standardVoiceAvailability) {
StandardVoiceAvailability.Ready -> CapabilityTone.Good
StandardVoiceAvailability.SignInRequired -> CapabilityTone.Info
StandardVoiceAvailability.Unsupported,
StandardVoiceAvailability.Unreachable -> CapabilityTone.Warning
StandardVoiceAvailability.Unknown -> CapabilityTone.Neutral
}
val relayValue = when {
!relayEnabled -> "Disabled"
!relayConfigured -> "Optional"
relayReady -> "Ready"
relayUiState == RelayUiState.Stale -> "Reconnect"
else -> "Configured"
}
val relayTone = when {
!relayEnabled || !relayConfigured -> CapabilityTone.Neutral
relayReady -> CapabilityTone.Good
relayUiState == RelayUiState.Stale -> CapabilityTone.Warning
else -> CapabilityTone.Info
}
val terminalValue = when {
!relayEnabled -> "Disabled"
authState is AuthState.Paired -> "Ready"
relayConfigured -> "Pair Relay"
else -> "Optional"
}
val terminalTone = when (terminalValue) {
"Ready" -> CapabilityTone.Good
"Pair Relay" -> CapabilityTone.Info
else -> CapabilityTone.Neutral
}
val proxyValue = if (secureProxyAdvertised) "Available" else "Not advertised"
val proxyTone = if (secureProxyAdvertised) CapabilityTone.Good else CapabilityTone.Neutral
// Lighter than the old six-filled-tile grid: one subtle grouped surface
// with a status dot + value per capability, dividers between rows. The
// header/glance pills used to duplicate API/Dashboard/Voice/Relay state;
// this list is now the single place those facts live on the active card.
Surface(
color = MaterialTheme.colorScheme.surface.copy(alpha = 0.5f),
shape = RoundedCornerShape(12.dp),
modifier = Modifier.fillMaxWidth(),
) {
Column(modifier = Modifier.padding(horizontal = 4.dp, vertical = 4.dp)) {
CapabilityRow(
label = "Vanilla Hermes API",
value = apiValue,
tone = apiTone,
onClick = onOpenApiInfo,
)
CapabilityDivider()
CapabilityRow(
label = "Dashboard",
value = dashboardValue,
tone = dashboardTone,
onClick = onOpenDashboard,
)
CapabilityDivider()
CapabilityRow(
label = "Vanilla Hermes voice",
value = voiceValue,
tone = voiceTone,
onClick = if (standardVoiceAvailability ==
StandardVoiceAvailability.SignInRequired
) {
onOpenDashboard
} else {
null
},
)
CapabilityDivider()
CapabilityRow(
label = "Relay tools",
value = relayValue,
tone = relayTone,
onClick = onOpenRelayInfo,
)
CapabilityDivider()
CapabilityRow(
label = "Terminal",
value = terminalValue,
tone = terminalTone,
onClick = onOpenSessionInfo,
)
CapabilityDivider()
CapabilityRow(
label = "Secure proxy",
value = proxyValue,
tone = proxyTone,
)
}
}
}
private enum class CapabilityTone { Neutral, Good, Info, Warning }
/** Hairline divider between capability rows — inset so it reads as a list. */
@Composable
private fun CapabilityDivider() {
HorizontalDivider(
modifier = Modifier.padding(horizontal = 12.dp),
color = MaterialTheme.colorScheme.outlineVariant.copy(alpha = 0.4f),
)
}
/**
* One capability line: a status dot, the feature name, and its current
* value (right-aligned, colored by tone). Replaces the old filled
* [CapabilityChip] tile — status now reads as a dot + value, so the row
* stays light and the six features chunk as a scannable list rather than a
* dense grid of dark-blue blocks. Honesty principle: every feature is shown
* even when unavailable, with its short reason (e.g. "Not advertised") as
* the value rather than being hidden.
*/
@Composable
private fun CapabilityRow(
label: String,
value: String,
tone: CapabilityTone,
modifier: Modifier = Modifier,
onClick: (() -> Unit)? = null,
) {
val dotColor = when (tone) {
CapabilityTone.Good -> Color(0xFF4CAF50)
CapabilityTone.Info -> MaterialTheme.colorScheme.primary
CapabilityTone.Warning -> MaterialTheme.colorScheme.error
CapabilityTone.Neutral -> MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.5f)
}
val valueColor = when (tone) {
CapabilityTone.Good -> Color(0xFF4CAF50)
CapabilityTone.Info -> MaterialTheme.colorScheme.primary
CapabilityTone.Warning -> MaterialTheme.colorScheme.error
CapabilityTone.Neutral -> MaterialTheme.colorScheme.onSurfaceVariant
}
val rowModifier = modifier
.fillMaxWidth()
.then(if (onClick != null) Modifier.clickable(onClick = onClick) else Modifier)
.padding(horizontal = 8.dp, vertical = 10.dp)
Row(
modifier = rowModifier,
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.spacedBy(10.dp),
) {
Box(
modifier = Modifier
.size(8.dp)
.clip(RoundedCornerShape(50))
.background(dotColor),
)
Text(
text = label,
style = MaterialTheme.typography.bodyMedium,
color = MaterialTheme.colorScheme.onSurface,
modifier = Modifier.weight(1f),
maxLines = 1,
overflow = TextOverflow.Ellipsis,
)
Text(
text = value,
style = MaterialTheme.typography.bodySmall,
color = valueColor,
maxLines = 1,
overflow = TextOverflow.Ellipsis,
)
if (onClick != null) {
Icon(
imageVector = Icons.Filled.ChevronRight,
contentDescription = null,
tint = MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.6f),
modifier = Modifier.size(16.dp),
)
}
}
}
/**
* Advanced expandable section — three subsections:
* - Manual URL configuration (API URL + key + Save & Test,
@@ -416,7 +665,7 @@ private fun ManualUrlSubsection(
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Text(
text = "Relay is optional for voice. Standard voice uses the Hermes API; Relay voice uses this route when selected or needed.",
text = "Relay is optional for voice. Vanilla Hermes voice uses the Hermes API; Relay voice uses this route when selected or needed.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
@@ -862,6 +1111,10 @@ fun ActiveCardSecurityPosture(
onNavigateToPairedDevices: () -> Unit,
) {
val relayUrl by connectionViewModel.relayUrl.collectAsState()
val effectiveApiServerUrl by connectionViewModel.effectiveApiServerUrl.collectAsState()
val effectiveDashboardUrl by connectionViewModel.effectiveDashboardUrl.collectAsState()
val effectiveRelayUrl by connectionViewModel.effectiveRelayUrl.collectAsState()
val relayConfigured by connectionViewModel.relayConfigured.collectAsState()
val insecureReason by connectionViewModel.insecureReason.collectAsState()
val isTailscaleDetected by connectionViewModel.isTailscaleDetected.collectAsState()
val currentPairedSession by connectionViewModel.currentPairedSession.collectAsState()
@@ -870,14 +1123,43 @@ fun ActiveCardSecurityPosture(
// say "Plain (on LAN)" instead of "Insecure (network unknown)" when
// the resolver already knows which candidate we're on.
val activeEndpoint by connectionViewModel.activeEndpoint.collectAsState()
val selectedRouteUrls = buildList {
effectiveApiServerUrl.trim().takeIf { it.isNotBlank() }?.let(::add)
effectiveDashboardUrl.trim().takeIf { it.isNotBlank() }?.let(::add)
val selectedRelayUrl = effectiveRelayUrl.ifBlank { relayUrl }
if (relayConfigured || selectedRelayUrl.isNotBlank()) {
selectedRelayUrl.trim().takeIf { it.isNotBlank() }?.let(::add)
}
}
val secureUrlCount = selectedRouteUrls.count { url ->
isSelectedRouteUrlSecure(
url = url,
activeEndpoint = activeEndpoint,
isTailscaleDetected = isTailscaleDetected,
)
}
val transportState = when {
selectedRouteUrls.isEmpty() -> null
secureUrlCount == selectedRouteUrls.size -> TransportSecurityState.AllSecure
secureUrlCount > 0 -> TransportSecurityState.Mixed
else -> TransportSecurityState.AllInsecure
}
TransportSecurityBadge(
isSecure = isUrlSecure(relayUrl),
reason = insecureReason.ifBlank { null },
size = TransportSecuritySize.Row,
modifier = Modifier.fillMaxWidth(),
activeRole = activeEndpoint?.role,
)
if (transportState != null) {
TransportSecurityBadge(
state = transportState,
size = TransportSecuritySize.Row,
modifier = Modifier.fillMaxWidth(),
)
} else {
TransportSecurityBadge(
isSecure = isUrlSecure(relayUrl),
reason = insecureReason.ifBlank { null },
size = TransportSecuritySize.Row,
modifier = Modifier.fillMaxWidth(),
activeRole = activeEndpoint?.role,
)
}
if (isTailscaleDetected) {
Row(
@@ -948,6 +1230,29 @@ fun ActiveCardSecurityPosture(
}
}
private fun isSelectedRouteUrlSecure(
url: String,
activeEndpoint: EndpointCandidate?,
isTailscaleDetected: Boolean,
): Boolean {
if (isUrlSecure(url)) return true
return activeEndpoint.isEncryptedOverlayRoute(isTailscaleDetected)
}
private fun EndpointCandidate?.isEncryptedOverlayRoute(isTailscaleDetected: Boolean): Boolean {
if (this == null) return false
val role = role.lowercase()
val securityHint = security.orEmpty().lowercase()
return role == "tailscale" ||
(isTailscaleDetected && securityHint.contains("tailscale")) ||
role == "plugin_proxy" ||
role == "plugin-proxy" ||
hasSecureProxy() ||
securityHint.contains("wireguard") ||
securityHint.contains("https") ||
securityHint.contains("tls")
}
/**
* Numbered step row for the Manual pairing code fallback. Tightly
* coupled to its Card 3 layout — step badge sizing + content shape —
@@ -0,0 +1,40 @@
package com.hermesandroid.relay.ui.components
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Text
import androidx.compose.runtime.Composable
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.layout.ContentScale
import androidx.compose.ui.text.TextStyle
import coil3.compose.AsyncImage
import java.io.File
/**
* The agent's "face" for a circular avatar badge: the active profile's local
* icon ([LocalAgentIconPath]) if one is set, otherwise the first letter of [name]
* on the badge's primary background. Fills its container — wrap it in the
* circular `Surface`/`Box` that owns the shape and color.
*/
@Composable
fun AgentAvatarFace(name: String, letterStyle: TextStyle, modifier: Modifier = Modifier) {
val iconPath = LocalAgentIconPath.current
if (!iconPath.isNullOrBlank()) {
AsyncImage(
model = File(iconPath),
contentDescription = null,
contentScale = ContentScale.Crop,
modifier = modifier.fillMaxSize(),
)
} else {
Box(modifier = modifier.fillMaxSize(), contentAlignment = Alignment.Center) {
Text(
text = name.firstOrNull()?.uppercase() ?: "H",
style = letterStyle,
color = MaterialTheme.colorScheme.onPrimary,
)
}
}
}
@@ -0,0 +1,85 @@
package com.hermesandroid.relay.ui.components
import android.net.Uri
import androidx.activity.compose.rememberLauncherForActivityResult
import androidx.activity.result.contract.ActivityResultContracts
import androidx.compose.foundation.background
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.shape.CircleShape
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.OutlinedButton
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.runtime.Composable
import androidx.compose.runtime.collectAsState
import androidx.compose.runtime.getValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.clip
import androidx.compose.ui.layout.ContentScale
import androidx.compose.ui.unit.dp
import coil3.compose.AsyncImage
import com.hermesandroid.relay.viewmodel.ConnectionViewModel
import java.io.File
/**
* Per-profile agent-icon picker — the visual twin of the local-name (alias) row.
* The chosen image is copied into app storage and shown beside the agent's name
* in chat. Client-side only: never sent to Hermes. Keyed per `(connection,
* profile)` by [ConnectionViewModel.setProfileIcon] / `ProfileIconStore`.
*/
@Composable
fun AgentIconRow(connectionViewModel: ConnectionViewModel) {
val iconPath by connectionViewModel.profileIcon.collectAsState()
val launcher = rememberLauncherForActivityResult(
ActivityResultContracts.OpenDocument()
) { uri: Uri? -> uri?.let { connectionViewModel.setProfileIcon(it) } }
Column(verticalArrangement = Arrangement.spacedBy(6.dp)) {
Text(
text = "Agent icon",
style = MaterialTheme.typography.labelLarge,
color = MaterialTheme.colorScheme.onSurface,
)
Row(
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.spacedBy(12.dp),
) {
Box(
modifier = Modifier
.size(44.dp)
.clip(CircleShape)
.background(MaterialTheme.colorScheme.surfaceVariant),
contentAlignment = Alignment.Center,
) {
val path = iconPath
if (!path.isNullOrBlank()) {
AsyncImage(
model = File(path),
contentDescription = "Agent icon",
contentScale = ContentScale.Crop,
modifier = Modifier.fillMaxSize(),
)
}
}
OutlinedButton(onClick = { launcher.launch(arrayOf("image/*")) }) {
Text(if (iconPath.isNullOrBlank()) "Set image" else "Change")
}
if (!iconPath.isNullOrBlank()) {
TextButton(onClick = { connectionViewModel.clearProfileIcon() }) {
Text("Clear")
}
}
}
Text(
text = "Shown beside this profile's name in chat. Stays on this device — never sent to Hermes.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
@@ -0,0 +1,605 @@
package com.hermesandroid.relay.ui.components
import android.content.Context
import android.provider.Settings
import android.view.accessibility.AccessibilityManager
import androidx.activity.compose.BackHandler
import androidx.compose.animation.AnimatedVisibility
import androidx.compose.animation.core.MutableTransitionState
import androidx.compose.animation.core.tween
import androidx.compose.animation.fadeIn
import androidx.compose.animation.fadeOut
import androidx.compose.animation.slideInVertically
import androidx.compose.foundation.background
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.verticalScroll
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.heightIn
import androidx.compose.foundation.layout.imePadding
import androidx.compose.foundation.layout.navigationBarsPadding
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.size
import androidx.compose.foundation.layout.statusBarsPadding
import androidx.compose.foundation.layout.widthIn
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.text.BasicTextField
import androidx.compose.foundation.text.KeyboardActions
import androidx.compose.foundation.text.KeyboardOptions
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.automirrored.filled.Send
import androidx.compose.material.icons.filled.Close
import androidx.compose.material3.Icon
import androidx.compose.material3.IconButton
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Surface
import androidx.compose.material3.Text
import androidx.compose.runtime.Composable
import androidx.compose.runtime.DisposableEffect
import androidx.compose.runtime.LaunchedEffect
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateListOf
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.rememberUpdatedState
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.alpha
import androidx.compose.ui.draw.drawWithContent
import androidx.compose.ui.graphics.BlendMode
import androidx.compose.ui.graphics.Brush
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.graphics.CompositingStrategy
import androidx.compose.ui.graphics.SolidColor
import androidx.compose.ui.graphics.graphicsLayer
import androidx.compose.ui.platform.LocalConfiguration
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.semantics.LiveRegionMode
import androidx.compose.ui.semantics.clearAndSetSemantics
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.liveRegion
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.text.input.ImeAction
import androidx.compose.ui.text.style.TextOverflow
import androidx.compose.ui.unit.Dp
import androidx.compose.ui.unit.dp
import com.hermesandroid.relay.data.ChatMessage
import com.hermesandroid.relay.data.MessageRole
import com.hermesandroid.relay.ui.components.avatar.AvatarRenderState
import com.hermesandroid.relay.ui.components.avatar.LocalAgentAvatar
import com.hermesandroid.relay.ui.theme.RelayRefresh
import kotlinx.coroutines.delay
// --- Text-flow tuning constants -------------------------------------------
//
// All time-based numbers stay inside the ranges WP-C1 prescribes so the
// "clean text flowing in and fading out" reads calm rather than frantic.
/** Soft word-wrap width for a flow line — keeps each buffer entry to ~one
* visual line so the bounded buffer maps cleanly to "≤6 lines". */
private const val FLOW_MAX_CHARS = 42
/** Soft-wrap target only — the visible buffer is now bounded by the ~1/3
* screen viewport + scroll, not a hard line count. */
private const val FLOW_MAX_LINES = 6
/** Memory ceiling for the persistent line buffer. Lines past this (already
* scrolled well above the faded top edge) are dropped silently so a very long
* turn can't grow the list without bound. */
private const val FLOW_BUFFER_MAX = 80
/** How long a settled line lingers after it stops growing, before it begins
* fading. Inside the 2.5–4s band from the spec. */
private const val FLOW_DWELL_MS = 3_000L
private const val FLOW_FADE_IN_MS = 180
private const val FLOW_FADE_OUT_MS = 600
/** Buffer maintenance cadence. Cheap list bookkeeping only — it mutates
* observed state (and so triggers recomposition) only when something
* actually changes, so an idle clean mode does not churn the UI. */
private const val FLOW_TICK_MS = 80L
/**
* One ephemeral line in the text flow.
*
* [text] and [visibility] are snapshot-observed so a growing tail or a
* fade-out re-renders just that line. [settledAt]/[hiddenAt] are plain
* bookkeeping read only by the maintenance loop, so they intentionally do
* NOT trigger recomposition.
*
* [visibility] starts `currentState = false, targetState = true`; handing
* that to `AnimatedVisibility(visibleState = …)` plays the enter transition
* the first time the line is composed — the idiomatic "animate on appear".
*/
private class FlowLine(val key: Int, initialText: String) {
var text by mutableStateOf(initialText)
val visibility = MutableTransitionState(false).apply { targetState = true }
/** Wall-clock millis at which the line stopped growing (null while it is
* still the active streaming tail). Starts the dwell countdown. */
var settledAt: Long? = null
/** Wall-clock millis at which the fade-out was requested. */
var hiddenAt: Long? = null
}
/**
* Split [text] into short, append-only flow segments.
*
* Explicit newlines hard-break; long paragraphs greedily soft-wrap at word
* boundaries to [maxChars]. Because the source content only ever grows
* (streaming appends), every segment except the last is final the moment the
* next word/line exists — which is exactly what lets the caller treat the
* last segment as the "growing tail" and everything before it as settled,
* and key each line by its stable index.
*/
private fun segmentFlowLines(text: String, maxChars: Int): List<String> {
if (text.isBlank()) return emptyList()
val out = ArrayList<String>()
for (rawLine in text.split('\n')) {
val line = rawLine.trim()
if (line.isEmpty()) continue
val current = StringBuilder()
for (word in line.split(' ')) {
if (word.isEmpty()) continue
val candidate = if (current.isEmpty()) word.length else current.length + 1 + word.length
if (candidate > maxChars && current.isNotEmpty()) {
out.add(current.toString())
current.setLength(0)
current.append(word)
} else {
if (current.isNotEmpty()) current.append(' ')
current.append(word)
}
}
if (current.isNotEmpty()) out.add(current.toString())
}
return out
}
/**
* Soft fade on the TOP edge so lines that scroll up dissolve cleanly into the
* background instead of hard-clipping — the "slides up and clears" look — while
* the avatar above stays unobstructed. Renders the content into an offscreen
* layer and masks the top [fade] dp with a transparent->opaque gradient.
*/
private fun Modifier.topFadeEdge(fade: Dp = 28.dp): Modifier = this
.graphicsLayer { compositingStrategy = CompositingStrategy.Offscreen }
.drawWithContent {
drawContent()
val fadePx = fade.toPx().coerceAtMost(size.height)
if (fadePx <= 0f) return@drawWithContent
drawRect(
brush = Brush.verticalGradient(
0f to Color.Transparent,
(fadePx / size.height) to Color.Black,
),
blendMode = BlendMode.DstIn,
)
}
/** Resolved motion/accessibility posture for clean mode. */
private data class CleanMotionState(
/** OS animator scale is non-zero (i.e. system animations are ON). */
val osAnimations: Boolean,
/** TalkBack-style touch exploration is active — faded text is unreadable
* to it, so the text path must fall back to a static, announced mirror. */
val touchExploration: Boolean,
)
@Composable
private fun rememberCleanMotionState(): CleanMotionState {
val context = LocalContext.current
// ANIMATOR_DURATION_SCALE == 0 is the platform "remove animations" / many
// OEM "reduce motion" toggles. Read once on entry; a mid-mode toggle is
// rare and recovered by leaving + re-entering the mode.
val osAnimations = remember {
runCatching {
Settings.Global.getFloat(
context.contentResolver,
Settings.Global.ANIMATOR_DURATION_SCALE,
1f,
) != 0f
}.getOrDefault(true)
}
val a11y = remember {
context.getSystemService(Context.ACCESSIBILITY_SERVICE) as? AccessibilityManager
}
var touchExploration by remember {
mutableStateOf(a11y?.isTouchExplorationEnabled == true)
}
DisposableEffect(a11y) {
val listener = AccessibilityManager.TouchExplorationStateChangeListener { enabled ->
touchExploration = enabled
}
a11y?.addTouchExplorationStateChangeListener(listener)
onDispose { a11y?.removeTouchExplorationStateChangeListener(listener) }
}
return CleanMotionState(osAnimations = osAnimations, touchExploration = touchExploration)
}
/**
* Ephemeral, themed text flow bound to the agent's streaming reply.
*
* New segments materialize with `fadeIn + slideInVertically`; a settled line
* dwells ~[FLOW_DWELL_MS], then `fadeOut`s and is **removed from the buffer**
* (it leaves the composition tree, so it stops composing — not merely
* alpha-0). The still-growing tail never fades; its dwell starts only once
* [streaming] flips false. The buffer is hard-capped at [FLOW_MAX_LINES].
*
* Accessibility: when [motionEnabled] is false (animations disabled, OS
* reduce-motion, or TalkBack touch exploration) the flow renders the recent
* lines **statically** inside a polite live region — never gating the
* conversation on animation. Even on the animated path a visually-hidden
* polite mirror carries the readable words, since faded glyphs are
* unreadable to assistive tech.
*
* @param content the last assistant message's (streaming) content.
* @param streaming whether that message is still growing this turn.
* @param messageId stable id of the bound message; a new id resets the buffer.
*/
@Composable
fun AgentTextFlow(
content: String,
streaming: Boolean,
messageId: String?,
motionEnabled: Boolean,
modifier: Modifier = Modifier,
) {
val flowStyle = MaterialTheme.typography.bodyMedium.copy(fontFamily = FontFamily.Monospace)
val flowColor = MaterialTheme.colorScheme.onSurfaceVariant
// Readable, non-faded mirror of the visible tail — used as the live-region
// text on both paths so assistive tech hears the words.
val mirrorText = remember(content) {
segmentFlowLines(content, FLOW_MAX_CHARS).takeLast(FLOW_MAX_LINES).joinToString(" ")
}
// --- Static / reduced-motion path -------------------------------------
if (!motionEnabled) {
val staticLines = remember(content) {
segmentFlowLines(content, FLOW_MAX_CHARS).takeLast(FLOW_BUFFER_MAX)
}
val staticScroll = rememberScrollState()
// Pin the latest line to the bottom of the bounded viewport.
LaunchedEffect(staticLines.size) { staticScroll.scrollTo(staticScroll.maxValue) }
// No contentDescription — the merged child Text content IS the readable
// content; liveRegion announces it on change. Lines persist + scroll
// (bounded + top-faded like the animated path) — they never vanish.
Column(
modifier = modifier
.semantics { liveRegion = LiveRegionMode.Polite }
.topFadeEdge()
.verticalScroll(staticScroll),
verticalArrangement = Arrangement.Bottom,
) {
staticLines.forEach { line ->
Text(
text = line,
style = flowStyle,
color = flowColor,
maxLines = 2,
overflow = TextOverflow.Ellipsis,
modifier = Modifier.fillMaxWidth(),
)
}
}
return
}
// --- Animated path ----------------------------------------------------
val flowLines = remember(messageId) { mutableStateListOf<FlowLine>() }
val currentContent by rememberUpdatedState(content)
val currentStreaming by rememberUpdatedState(streaming)
LaunchedEffect(messageId) {
flowLines.clear()
// Largest segment index ever materialized — guards against re-adding a
// line that was dropped from the front by the memory cap.
var maxKeyAdded = -1
while (true) {
val text = currentContent
val isStreamingNow = currentStreaming
val segs = segmentFlowLines(text, FLOW_MAX_CHARS)
// Add new lines (they slide in) and grow the still-streaming tail.
// Lines PERSIST — they never fade out; older ones simply scroll up
// within the bounded ~1/3-height viewport and dissolve at the top
// fade edge. (No dwell / fade-out / removal anymore.)
segs.forEachIndexed { i, s ->
val existing = flowLines.firstOrNull { it.key == i }
if (existing == null) {
if (i > maxKeyAdded) {
flowLines.add(FlowLine(key = i, initialText = s))
maxKeyAdded = i
}
} else if (existing.text != s) {
existing.text = s
}
}
// Memory guard: drop the oldest lines once well past the viewport
// (already scrolled above the fade — invisible to the user).
while (flowLines.size > FLOW_BUFFER_MAX) flowLines.removeAt(0)
// Nothing left to do once the turn ended and every segment is in.
if (!isStreamingNow && maxKeyAdded >= segs.lastIndex) return@LaunchedEffect
delay(FLOW_TICK_MS)
}
}
val scrollState = rememberScrollState()
// Pin the latest line to the bottom as content streams in / lines slide up.
LaunchedEffect(flowLines.size, flowLines.lastOrNull()?.text) {
scrollState.scrollTo(scrollState.maxValue)
}
Box(modifier = modifier) {
// Visually-hidden, readable, politely-announced mirror. Present even
// with motion on, so non-touch assistive tech still receives the words
// the faded glyphs can't convey. The Text's own content is its
// semantics text, so liveRegion alone announces it on change.
Text(
text = mirrorText,
maxLines = 1,
modifier = Modifier
.fillMaxWidth()
.heightIn(max = 1.dp)
.alpha(0f)
.semantics { liveRegion = LiveRegionMode.Polite },
style = flowStyle,
)
Column(
modifier = Modifier
.align(Alignment.BottomStart)
.fillMaxWidth()
.topFadeEdge()
.verticalScroll(scrollState),
verticalArrangement = Arrangement.Bottom,
) {
flowLines.forEach { line ->
androidx.compose.runtime.key(line.key) {
AnimatedVisibility(
visibleState = line.visibility,
enter = fadeIn(tween(FLOW_FADE_IN_MS)) +
slideInVertically(tween(FLOW_FADE_IN_MS)) { it / 6 },
exit = fadeOut(tween(FLOW_FADE_OUT_MS)),
) {
Text(
text = line.text,
style = flowStyle,
color = flowColor,
maxLines = 2,
overflow = TextOverflow.Ellipsis,
// The visible glyphs fade; the mirror above owns
// accessibility, so keep AT off these duplicates.
modifier = Modifier
.fillMaxWidth()
.clearAndSetSemantics {},
)
}
}
}
}
}
}
/**
* Thin single-line composer for clean mode.
*
* Deliberately stripped: no model/effort pills, no attachments, no slash
* palette — just a pill field plus a send affordance, calling [onSend] with
* the same [com.hermesandroid.relay.viewmodel.ChatViewModel.sendMessage]
* contract the full composer uses. Internal text state is UI-local.
*/
@Composable
private fun CleanModeComposer(
enabled: Boolean,
onSend: (String) -> Unit,
modifier: Modifier = Modifier,
) {
var text by remember { mutableStateOf("") }
val canSend = enabled && text.isNotBlank()
val submit = {
val trimmed = text.trim()
if (enabled && trimmed.isNotEmpty()) {
onSend(trimmed)
text = ""
}
}
Surface(
shape = RoundedCornerShape(28.dp),
color = MaterialTheme.colorScheme.surfaceContainerHigh,
modifier = modifier.fillMaxWidth(),
) {
Row(
modifier = Modifier.padding(start = 18.dp, end = 6.dp, top = 4.dp, bottom = 4.dp),
verticalAlignment = Alignment.CenterVertically,
horizontalArrangement = Arrangement.spacedBy(6.dp),
) {
BasicTextField(
value = text,
onValueChange = { text = it },
modifier = Modifier
.weight(1f)
.heightIn(min = 40.dp)
.padding(vertical = 8.dp),
enabled = enabled,
singleLine = true,
textStyle = MaterialTheme.typography.bodyLarge.copy(
color = MaterialTheme.colorScheme.onSurface,
),
cursorBrush = SolidColor(MaterialTheme.colorScheme.primary),
keyboardOptions = KeyboardOptions(imeAction = ImeAction.Send),
keyboardActions = KeyboardActions(onSend = { submit() }),
decorationBox = { inner ->
Box(contentAlignment = Alignment.CenterStart) {
if (text.isEmpty()) {
Text(
text = "Message",
style = MaterialTheme.typography.bodyLarge,
color = RelayRefresh.Dim,
)
}
inner()
}
},
)
IconButton(
onClick = submit,
enabled = canSend,
modifier = Modifier.size(44.dp),
) {
Icon(
imageVector = Icons.AutoMirrored.Filled.Send,
contentDescription = "Send",
tint = if (canSend) {
MaterialTheme.colorScheme.primary
} else {
MaterialTheme.colorScheme.onSurfaceVariant.copy(alpha = 0.5f)
},
)
}
}
}
}
/**
* Clean text-flow chat mode — a full-screen, minimalist third presentation of
* the agent surface (alongside normal chat and the voice overlay).
*
* Centered morphing sphere, a calm themed text flow ([AgentTextFlow]) instead
* of a persistent transcript, and a thin composer. Mirrors the voice overlay's
* centered-sphere + bottom-content skeleton (`VoiceModeOverlay.kt:262-280`).
*
* Exit is an **explicit control** (top-corner dismiss + system back) — never
* any-tap, because the in-mode composer needs taps. All mode state lives in
* the caller as plain UI-local state; this is a presentation over the same
* conversation, not new ViewModel state.
*
* Honors [animationEnabled], OS reduce-motion, and TalkBack: the sphere
* renders a static frame and the text stays readable + announced when motion
* is suppressed.
*
* The avatar is rendered through the [LocalAgentAvatar] seam (WP-C2), so clean
* mode gets future "pets" for free alongside chat and the voice overlay.
*/
@Composable
fun CleanChatMode(
messages: List<ChatMessage>,
isStreaming: Boolean,
sphereState: SphereState,
streamingIntensity: Float,
toolCallBurst: Float,
animationEnabled: Boolean,
enabled: Boolean,
onSend: (String) -> Unit,
onExit: () -> Unit,
modifier: Modifier = Modifier,
) {
val motion = rememberCleanMotionState()
val sphereAnimated = animationEnabled && motion.osAnimations
// Faded text is unreadable to touch exploration, so the text path goes
// static (readable + announced) whenever TalkBack is exploring.
val textMotionEnabled = sphereAnimated && !motion.touchExploration
val lastAssistant = remember(messages) {
messages.lastOrNull { it.role == MessageRole.ASSISTANT }
}
val flowContent = lastAssistant?.content.orEmpty()
val flowStreaming = lastAssistant?.isStreaming == true && isStreaming
// Cap the flow at ~1/3 of the screen so lines can slide up and accumulate
// without ever climbing into / blocking the avatar above them.
val maxFlowHeight = (LocalConfiguration.current.screenHeightDp * 0.34f).dp
BackHandler(enabled = true) { onExit() }
val sphereDescription = remember(sphereState) {
"Agent ${sphereState.name.lowercase()}"
}
Box(
modifier = modifier
.fillMaxSize()
// Opaque so the chat underneath is fully hidden — this is a mode,
// not a translucent overlay.
.background(RelayRefresh.Background),
) {
Column(
modifier = Modifier
.fillMaxSize()
.statusBarsPadding()
.navigationBarsPadding()
.imePadding()
.padding(horizontal = 20.dp),
) {
// Explicit dismiss — the only way out besides system back.
Row(
modifier = Modifier.fillMaxWidth(),
horizontalArrangement = Arrangement.End,
) {
IconButton(onClick = onExit) {
Icon(
imageVector = Icons.Filled.Close,
contentDescription = "Exit clean mode",
tint = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
// Centered sphere — takes the slack so the flow + composer keep a
// stable bottom anchor as lines come and go.
Box(
modifier = Modifier
.fillMaxWidth()
.weight(1f),
contentAlignment = Alignment.Center,
) {
Box(
modifier = Modifier
.fillMaxSize()
.semantics { contentDescription = sphereDescription },
) {
LocalAgentAvatar.current.Render(
state = AvatarRenderState(
state = sphereState,
intensity = streamingIntensity,
toolCallBurst = toolCallBurst,
// Pin to a still frame when motion is suppressed.
paused = !sphereAnimated,
),
modifier = Modifier.fillMaxSize(),
)
}
}
AgentTextFlow(
content = flowContent,
streaming = flowStreaming,
messageId = lastAssistant?.id,
motionEnabled = textMotionEnabled,
modifier = Modifier
.fillMaxWidth()
.widthIn(max = 560.dp)
.heightIn(min = 96.dp, max = maxFlowHeight)
.padding(bottom = 12.dp),
)
CleanModeComposer(
enabled = enabled,
onSend = onSend,
modifier = Modifier.padding(bottom = 12.dp),
)
}
}
}
@@ -0,0 +1,53 @@
package com.hermesandroid.relay.ui.components
import androidx.compose.runtime.Composable
import androidx.compose.runtime.LaunchedEffect
import androidx.compose.runtime.mutableFloatStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.withFrameNanos
import kotlinx.coroutines.delay
/**
* Frame-throttled stand-in for `rememberInfiniteTransition` for slow, ambient
* effects — heartbeat dots, banner glows, drifting orbs. Returns a phase that
* loops `0f → 1f` every [periodMillis], advanced at roughly [fps] instead of
* the display refresh rate.
*
* Why this exists: on Android 15 the platform logs `setRequestedFrameRate` on
* every Compose draw pass (and on Samsung builds at INFO level). An
* always-visible `rememberInfiniteTransition` pins the entire window at the
* panel's refresh (e.g. 120Hz) for as long as it's composed — flooding logcat
* and burning battery to animate motion the eye can't resolve at full rate
* anyway. ~30fps is imperceptible for a multi-second pulse.
*
* When [running] is false the phase holds at `0f` and the loop parks (no frames
* requested), so a hidden/idle effect costs nothing.
*
* Linear by design (matches the `LinearEasing` + `RepeatMode.Restart` the old
* infinite transitions used). Map the phase to your value range at the call
* site, e.g. `1f + 0.8f * phase` for a 1f→1.8f scale.
*/
@Composable
fun rememberAmbientPhase(
periodMillis: Int,
fps: Int = 30,
running: Boolean = true,
): Float {
val phase = remember { mutableFloatStateOf(0f) }
LaunchedEffect(periodMillis, fps, running) {
if (!running || periodMillis <= 0) {
phase.floatValue = 0f
return@LaunchedEffect
}
val frameIntervalMs = (1000L / fps.coerceAtLeast(1)).coerceAtLeast(1L)
var lastNanos = withFrameNanos { it }
while (true) {
val now = withFrameNanos { it }
val dtMs = (now - lastNanos).coerceAtLeast(0L) / 1_000_000f
lastNanos = now
phase.floatValue = (phase.floatValue + dtMs / periodMillis) % 1f
delay(frameIntervalMs)
}
}
return phase.floatValue
}

Some files were not shown because too many files have changed in this diff Show More