dev
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a5f3d0a2c3 | fix(android): add bridge media sharing and mms handoff | ||
|
|
fd77da043b |
feat(bridge): v0.5.0 agent-aware phone status, auto-return, activity log
Three cross-layer additions that close longstanding visibility gaps
between the phone, the relay, and the host-side agent:
Agent awareness — unattended/screen/credential-lock state
-----------------------------------------------------------
PhoneSnapshot gains unattendedEnabled, credentialLockDetected, and
screenOn fields. PhoneStatusPromptBuilder.buildBridgeLine() now
appends explicit guidance — e.g. "Unattended access: off — commands
only land when the screen is already on" or "Unattended access: on,
but the device has a credential lock — commands will return
keyguard_blocked and fail."
BridgeStatusReporter.emitTick() emits a parallel `unattended` group
in the bridge.status WSS envelope so the host-side `/bridge/status`
cache (and the `android_phone_status` tool that reads it) sees the
same state for non-phone frontends like Discord. Push triggers fire
on toggle flip via ConnectionViewModel so the host cache updates in
~1s instead of waiting up to 30s for the periodic tick.
android_phone_status tool description updated so the LLM proactively
checks unattended.* fields and warns the user when commands will
hit keyguard_blocked.
Auto-return to Hermes-Relay
---------------------------
android_return_to_hermes is an LLM-called tool that the agent
routinely forgets, leaving the user stranded on Starbucks/Chrome/etc
after the run. Two safety nets:
1. Tightened tool descriptions (REQUIRED FINAL STEP, MANDATORY
CLEANUP framing in android_open_app + android_return_to_hermes).
2. New BridgeRunTracker singleton — coordinates two completion
signals: Chat-tab SSE run.completed (fast, phone-only) and a
12s bridge-idle timer (universal, works for Discord/CLI/web).
Whichever fires first dispatches a local /return_to_hermes via
handleLocalCommand. markReturnedToHermes() prevents double-fires
when the LLM does call return explicitly.
BridgeCommandHandler tracks foreground-shifting paths (/open_app,
/send_intent) at respond() and arms/resets the idle timer accordingly.
Reset hooks at both dispatch start and respond finish so slow-
executing commands (screenshots, big tree reads) don't eat the idle
budget.
Bridge activity log wiring
--------------------------
The Activity Log card on the Bridge tab was scaffolded in Phase 3 but
never wired — recordActivity() existed, the UI rendered the flow, but
no code ever called it. BridgeCommandHandler now emits a
BridgeActivityEntry per dispatched command (Success/Failed/Blocked)
via a new onActivity callback. ConnectionViewModel pipes it through
to BridgePreferencesRepository.appendEntry. High-frequency polls
(/ping, /events, /current_app, /screen_hash) suppressed so the log
shows user-meaningful activity, not noise. Per-route summarizers
produce natural-looking entries: "tap (540, 1200)", "open_app
com.starbucks.mobilecard", etc. resultText surfaces error strings
on Failed/Blocked so users can see WHY without digging through logs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
b7dab4c496 |
fix(voice-intents): close UI-automation bypass + structured contact phones
Follow-up pass from the P0/P1 audit after Bailey's 2026-04-15 on-device
test where Victor hit two distinct gaps: (1) he denied an SMS modal
and the agent immediately fell back to driving the Messages app UI by
android_open_app + android_tap, bypassing the denial, (2) the contact
Hannah Dixon had a Galaxy Watch phone AND a mobile, and the voice
auto-picker chose the Watch.
== P0: Close the android_tap destructive-verb bypass ==
BridgeSafetyManager.requiresConfirmation now fires on /tap and
/long_press in addition to /tap_text and /type. BridgeCommandHandler
adds extractDestructiveVerbText() which resolves the tapped node's
text via ScreenReader.findNodeById when the caller passes a nodeId
(the common LLM pattern after android_read_screen). The text is
pattern-matched against the user's destructive-verb list the same
way /tap_text does, so tapping a button whose label is "Send" /
"Delete" / "Pay" / "Confirm" now fires the safety modal regardless
of which tool the agent used to locate it.
Fail-open for coordinate-only taps (no nodeId) and for any case
where the snapshot/findNodeById fails — matches pre-0.4.0 semantics
for /tap so we're strictly ADDING coverage, not converting taps to
fail-closed. Coordinate hit-testing is a P0.5 follow-up and rarely
matters in practice because modern LLMs prefer nodeId after
android_read_screen.
Recycles window roots + the resolved node on the gate path to avoid
leaking AccessibilityNodeInfo handles.
== P1: Server-side phone preference sort + Galaxy Watch label heuristic ==
ActionExecutor.searchContacts now sorts each contact's phones list
in preference order before returning: mobile > main > home > work >
other > everything else (custom, watch, fax, pager). Critically, a
label-substring heuristic detects "watch" / "pager" / "smartwatch" /
"fax" in the phone's human-readable label and demotes those entries
to the bottom tier REGARDLESS of their canonical type. This catches
Galaxy Watch entries that Samsung Health registers as TYPE_MOBILE
with a "Watch" label, which the type-only ranker from the previous
commit couldn't distinguish from real mobiles.
Stable sort via rank*1000 + originalIndex so within a tier the
insertion order from the content provider is preserved.
The voice handler's pickPreferredPhone simplifies to "return the
first non-empty entry" since the list is pre-sorted server-side.
This centralizes ranking in one place and — crucially — means the
LLM tool-calling path benefits without needing tool-description
cooperation: the LLM can just use phones[0] blindly and get the
right number.
== P1: Denial-retry warnings in UI-automation tool descriptions ==
android_open_app, android_tap, android_tap_text, android_type, and
android_call all gained explicit DENIAL-RETRY GUARD paragraphs
telling the LLM not to use these tools to replicate a destructive
action that android_send_sms or android_call just received a
user_denied response for. android_send_sms already had this in the
previous commit; now all the fallback-path tools carry the same
warning so the LLM sees it no matter which alternate route it
considers. Belt and braces — the real enforcement is in the code
(the /tap verb gate now fires on the Messages app's Send button),
but the tool-description guidance helps the LLM fail gracefully
instead of hammering the modal.
== Structural cleanup: recursive JSON serialization ==
BridgeCommandHandler.respondFromResult went from "primitives-only
with toString fallback" to fully recursive anyToJsonElement for
nested Lists and Maps. Pre-fix, the search_contacts response
serialized the `contacts` field as a Kotlin list-repr string
("[{id=9, name=Hannah, phones=...}]") — the LLM had to parse
pseudo-JSON. Now it gets real JSON with the structured phones
list as actual arrays of objects. Same recursive serializer
handles future ActionResult fields that carry nested structure.
== userDeniedResponse helper ==
Three denial paths in BridgeCommandHandler (/tap_text verb gate,
/call, /send_sms) all now return the same canonical shape via a
userDeniedResponse(contextText) helper:
{
"error": "<context>. This is a FINAL denial. Do NOT retry...",
"error_code": "user_denied",
"reason": "confirmation_denied_or_timeout",
"final": true,
"instruction": "Do not retry via UI automation. Denial is terminal."
}
Structured error_code + explicit instruction field give the LLM
both machine-readable classification and a literal no-fallback
directive in the JSON payload. The free-text error carries the
contextual details (SMS recipient, call number, which verb fired
the tap gate).
== sideload_only 403s carry error_code + flavor ==
The four sideload-only 403s (/location, /search_contacts, /call,
/send_sms on a googlePlay build) now include error_code="sideload_only"
and flavor="googlePlay" so the LLM can classify them cleanly and
(per the android_send_sms tool description) ask the user for
explicit consent before falling back to UI automation, rather than
silently switching paths. This fixes the secondary bug where Victor
paraphrased a denial as "Direct SMS is blocked on this build" — on
sideload it was actually a user_denied, but Victor's loose free-text
interpretation let it slide into "build limitation" territory and
justified a UI-automation fallback.
== classifyBridgeError ==
Expanded to catch "user denied" (broader than "user denied
destructive action"), "this is a final denial", and "sideload-only"
substrings. Loosened user_denied pattern in case future error text
varies.
== Files ==
- app/src/main/kotlin/.../accessibility/ActionExecutor.kt
- fetchPhonesForContact returns List<Map> with {number, type, label}
- sortPhonesByPreference helper (new)
- phoneTypeKey + phoneTypeDisplayLabel helpers (new)
- searchContacts applies sortPhonesByPreference server-side
- app/src/main/kotlin/.../network/handlers/BridgeCommandHandler.kt
- extractDestructiveVerbText helper (new) — resolves node text
for /tap + /long_press by nodeId via ScreenReader
- userDeniedResponse helper (new) — canonical 403 shape
- anyToJsonElement helper (new) — recursive JSON serialization
- respondFromResult uses recursive serializer
- Three awaitConfirmation denial paths use userDeniedResponse
- Four sideload_only 403s carry error_code + flavor fields
- classifyBridgeError expanded patterns
- app/src/main/kotlin/.../bridge/BridgeSafetyManager.kt
- requiresConfirmation accepts /tap and /long_press paths
- app/src/sideload/kotlin/.../voice/VoiceBridgeIntentHandlerImpl.kt
- pickPreferredPhone simplified (server pre-sorts)
- ContactResolution.Found gained phoneType + phoneLabel +
totalPhones fields for richer voice preview
- SendSms branch speaks phone qualifier when totalPhones > 1
("texting Hannah's Mobile at ...")
- plugin/android_tool.py
- android_search_contacts description documents new structured
phones shape + disambiguation rules
- android_send_sms description carries TRUST MODEL + DENIAL
IS FINAL + sideload_only fallback guidance (~3000 chars)
- android_call gained DENIAL IS FINAL block
- android_open_app / android_tap / android_tap_text / android_type
gained DENIAL-RETRY GUARD warnings
Server synced + gateway restarted, all 18 tools still register,
check_requirements still returns True when phone_connected=true.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
ddd6f706df |
fix(voice-intents): trust-model language in tool descriptions
Victor was double-confirming before every SMS action, asking the user
"who is Hannah Dixon, is she in your contacts, confirm the wording" in
chat BEFORE calling any tool, even though:
1. android_search_contacts exists specifically so the LLM can resolve
names autonomously
2. The phone's destructive-verb safety modal is a hardcoded final
on-device Allow/Deny checkpoint that blocks every SMS send
Claude-family models have a trained safety reflex around messaging
actions — especially emotionally loaded content and especially to
names the model doesn't recognize. That reflex is generally good but
redundant here because the on-device modal is the real checkpoint.
The fix is to make the trust model explicit in the tool descriptions
so the model knows:
- The user gave it phone control explicitly via the master toggle
- The on-device modal is sufficient confirmation on its own
- Chat-side double-confirmation is redundant AND frustrating
- Contact names should be resolved via search_contacts, not by
bouncing lookup questions back to the user
android_search_contacts, android_send_sms, and android_call all got
expanded descriptions with a TRUST MODEL block explaining the
on-device checkpoint and explicitly prohibiting the double-confirm
anti-pattern.
Tool description lengths:
- android_search_contacts: 754 chars
- android_send_sms: 1837 chars
- android_call: 991 chars
These are long for tool descriptions, but Claude reads them carefully
and the behavioral override is worth the token cost. Alternative
approaches (personality prompt, plugin.yaml directive, system-level
instructions) all couple us to specific hermes-agent deploy config
or violate the "guidance travels with the plugin" principle we
established in
|
||
|
|
1323dc3dc3 |
fix(voice-intents): single-gate _check_requirements + structured 503
Reverting the three-gate rule from the previous commit. Bailey hit the
real-world cost on-device: with the master toggle off, Victor saw no
android_* tools and hallucinated a reason for their absence — "Phone
bridge isn't connected right now, pair the Hermes-Relay app" — then
asked the user for a pairing code when the phone was already paired.
The phone WAS connected. Victor filled the absence-of-tools with a
guessed explanation and steered the user to the wrong action.
Root cause is a fundamental LLM-interaction principle: hidden tools
don't mean "no capability" to the model, they mean "invent a narrative
for why this capability isn't here." The clean-no-tools signal I
rationalized in the last commit sounded good in theory but fails in
practice because LLMs reason about presence/absence by making things
up.
Fix: gate check_fn ONLY on phone_connected. Let downstream return
structured errors for the other failure modes:
- No a11y → HTTP 503 + error_code=service_unavailable + explicit
"the phone IS paired, this is NOT a pairing problem" text +
required_action field naming the exact Settings destination
- Master toggle off → HTTP 403 + error_code=bridge_disabled +
similar structured fields (already landed in
|
||
|
|
c5cdb45cdb |
fix(voice-intents): UX polish, chat parity, OEM hardening, portable guidance
Agent-team follow-up after Bailey's 2026-04-15 voice SMS test session.
Fifteen files touched, ranging from voice-mode UX to Android manifest
hardening to plugin tool descriptions. See DEVLOG for the full story
and "known residual items" list.
Voice mode UX:
- TTS speaks the post-dispatch result ("Text sent", "Cancelled",
"Permission needed") via the existing ttsQueue pipeline
- Voice-in-voice cancel: "cancel", "stop", "never mind", "abort",
"forget it", "wait" spoken DURING the 5s countdown now routes
straight to cancelPending() instead of being classified as a
fresh turn that goes to the LLM. VoiceBridgeIntentHandler interface
gained hasPendingDestructive() so VoiceViewModel can intercept
cancel utterances before the classifier runs
- Visual countdown: DestructiveCountdownState on VoiceUiState +
onCountdownStart callback + LinearProgressIndicator in the overlay
synced to the real delay, so the user sees the 5s window tick down
instead of staring at a static UI
- Phone-number literal bypass: "text +1 555 123 4567 saying hi" no
longer fails contact resolution; PHONE_NUMBER_REGEX detects
phone-shaped contacts and skips resolveContactPhone
- Multi-contact match hint: "Found 3 contacts matching John. Using
John Smith." when resolveContactPhone returns more than one hit
so the user knows one of several got picked (full multi-turn
disambiguation deferred to Wave 3)
Chat parity:
- ChatHandler intercepts tool.completed SSE events on /v1/runs for
android_* action tools (send_sms, call, search_contacts, open_app,
return_to_hermes, screenshot, press_key, setup) and emits a
structured follow-up bubble with the same per-category formatting
voice mode uses. Read-only + UI-micro-action tools skipped to
avoid doubling up with the existing ToolProgressCard
- New appendLocalVoiceIntentResult(description, agentName) signature
so chat-originated action bubbles get a distinct "Phone action"
caption vs voice-originated "Voice action"
- MessageBubble renders a subtle visual marker (tertiary-tinted
leading border) for both action-bubble types so they read as
distinct from LLM replies when interleaved
Plugin + phone-side:
- _check_requirements gates on all three: phone_connected AND
bridge.accessibility_granted AND bridge.master_enabled. Tools
vanish from the LLM's schema entirely when the master toggle
is off, giving the model a clean "no tools" signal instead of
an error-interpretation race. Trade-off: tools disappear mid-
session if the user flips the toggle — desired per "stop the
agent from controlling my phone" intent
- /return_to_hermes short-circuits when service.currentApp ==
service.packageName (already foreground), returning 200 with
a "note: already foreground" instead of re-firing the launch
intent. Benign for voice mode where Hermes is always foreground
- Tool descriptions carry the "prefer direct dispatch" +
"call return_to_hermes as final step" guidance inline on
android_open_app / android_send_sms / android_call. Portable
across hermes-agent installs; replaces the earlier Victor
personality prompt edit (reverted — only worked for one install)
Android OEM hardening:
- BridgeForegroundService.onTaskRemoved revokes MediaProjection
and downgrades the FGS type back to SPECIAL_USE-only when the
user swipes the app from recents. Fixes the bug where the system
screen-cast icon persisted in the status bar indefinitely after
app close because foreground services legitimately survive task
removal and the projection stayed bound. Bridge itself keeps
running so agent phone control over WSS still works
- MainActivity manifest entry gained launchMode="singleTask" and
configChanges="uiMode|fontScale|locale|density|orientation|
screenSize|screenLayout|keyboardHidden". Before this the activity
was being recreated on config changes and certain resume paths,
triggering installSplashScreen() every time and showing the
splash on warm reopen. Samsung OneUI's aggressive backgrounded-
app kills (even with FGS) is the residual case that needs
user-side battery whitelisting and can't be fixed in code
DEVLOG.md updated with the full session narrative.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
5c763eb265 |
fix(voice-intents): honest errors, modal lifecycle, and missing LLM tools
End-to-end fix for the voice -> bridge -> agent pipeline. Twelve changes
across phone, plugin, and bridge layers so the voice fast-path, the LLM
tool-calling path, and the safety modal all actually work.
Voice fast-path:
- Remove Scroll voice intent (nobody says it aloud; /scroll route stays
for LLM android_scroll tool calls).
- Add SMS_INDIRECT regex to catch "send Hannah a text saying hi" phrasing
and "message" as a direct verb ("message Sam saying hi").
- Classify resolver failures via ContactResolution / AppResolution sealed
types so voice speaks specific per-category messages (permission,
service, not-found, no-phone, other) instead of "couldn't find" for
every failure mode.
- Pre-check SEND_SMS before the 5s countdown so permission-denied doesn't
silent-fail at the end of the confirmation flow.
- Flip voice state to Thinking immediately on classifier fall-through so
the UI shows progress during SSE connect latency.
- Enrich IntentResult.Handled with details: Map<String, String> for
structured chat-trace rendering (app label, package, match tier,
contact, resolved number, body, error code).
- Rewrite chat-trace formatter to render markdown per category.
Bridge safety modal:
- Fix showConfirmation threading -- ComposeView setContent must run on
Main, was called from Dispatchers.Default via voice local-dispatch,
threw, and got swallowed as "likely overlay permission missing".
- Fix OverlayLifecycleOwner.start() init order -- current androidx.savedstate
asserts performAttach runs while lifecycle is still INITIALIZED; the
old code advanced to CREATED first and tripped the assertion, killing
every destructive-verb modal attempt silently.
Voice mode UI:
- Replace single responseText slot with compact rolling transcript
observing ChatViewModel.messages (last 6), rendered via new
CompactTranscriptRow + StreamingResponseRow composables.
- Voice-action traces render via MarkdownContent. User messages keep
the "YOU" caption (fix mislabeling where voice-intent user messages
were captioned "ACTION").
- Preserve local voice-intent trace messages across
ChatHandler.loadMessageHistory reloads so "Opened Chrome" bubbles
don't vanish when session_end reload fires after fall-through.
LLM bridge path -- honest errors:
- respondFromResult emits structured error_code + required_permission
alongside the existing free-text error when ActionExecutor errors
match known patterns (permission_denied, service_unavailable,
user_denied). LLM gets both human-readable text AND a machine-readable
classification.
LLM bridge path -- missing tools:
- Fix plugin _check_requirements: was hitting /ping which returns
{pong, ts}, looking for phone_connected and accessibilityService
fields that do not exist there. Result: the gate always returned
False and hid all 13 non-setup tools from every gateway platform.
Now hits /bridge/status and requires BOTH phone_connected AND
bridge.accessibility_granted -- tools vanish from the LLM's schema
when a11y is revoked (common post-Studio-reinstall) instead of
letting the LLM confidently dispatch commands that 503.
- Add 4 new plugin tools: android_search_contacts, android_send_sms,
android_call, android_return_to_hermes. The first three wrap phone
routes that were fully implemented (with safety modals and direct
SmsManager / CALL_PHONE dispatch) but never exposed to the agent,
forcing it to drive the Messages app UI step-by-step. The fourth
lets the agent foreground Hermes Relay as the final step of a
phone-control task so the user sees the reply in-context without
manually switching apps.
- New /return_to_hermes route on BridgeCommandHandler, exempt from
the master-toggle check so the agent can always wrap up cleanly.
Tool count goes from 14 to 18. Plugin + gateway verified on the server
(systemctl --user restart hermes-gateway, _check_requirements() returns
True, all 18 tools register with no missing or orphan handlers).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
7a5e7e31f8 |
chore(rename): replace Greek-letter agent codenames with ASCII slugs
Bulk rename across the repo + Obsidian plan: agents previously identified by α β γ δ ε ζ η θ now use descriptive ASCII slugs (bridge-server, flavor-split, accessibility, bridge-ui, notif-listener, safety-rails, voice-intents, vision-nav). Greek letters were a math-paper convention that sorts nicely but renders badly in some terminals, can't be typed without a special keyboard, and made commit-history search awkward. Single sed pass per file, applied via bash for-loop across 42 repo files + the canonical Obsidian plan. Verified with grep -rc '[αβγδεζηθ]' = 0. Git history is not rewritten — existing commit subjects keep their Greek letters because rewriting them would require a force push and invalidate every commit hash since the divergence point. Going forward, branches, commit messages, and marker blocks all use ASCII. Marker block convention going forward: // === PHASE3-<slug>: ... === // === END PHASE3-<slug> === Followup blocks use PHASE3-<slug>-followup. |
||
|
|
985632dafe |
feat(phase3-α): migrate legacy bridge relay into unified relay (port 8767)
Retires the standalone bridge relay (plugin/tools/android_relay.py +
the duplicate top-level plugin/android_relay.py, both listening on port
8766) and folds its functionality into the unified Hermes-Relay on port
8767 as the bridge channel. Wire protocol (bridge.command /
bridge.response / bridge.status) is unchanged — only the transport moved.
- plugin/relay/channels/bridge.py: real BridgeHandler with phone_ws +
pending[request_id]→Future, asyncio.Lock-protected. handle_command()
mints request_id, sends bridge.command, awaits bridge.response with
30s timeout (matches legacy android_relay._RESPONSE_TIMEOUT).
detach_ws() fails all pending futures with ConnectionError on phone
disconnect so HTTP callers fail fast instead of hanging to timeout.
- plugin/relay/server.py: 14 HTTP routes (/ping, /screen, /screenshot,
/get_apps, /apps legacy alias, /current_app, /tap, /tap_text, /type,
/swipe, /open_app, /press_key, /scroll, /wait, /setup) delegate to
_bridge_dispatch → BridgeHandler.handle_command. BridgeError →
503/504/502 based on message. server.bridge.detach_ws(ws) wired into
_on_disconnect so phone drops instantly fail in-flight commands.
All additions bracketed by # === PHASE3-α: ... === / # === END
PHASE3-α === markers for mechanical merges with Agent ε (notification-
listener).
- plugin/android_tool.py + plugin/tools/android_tool.py: BRIDGE_URL
default 8766 → 8767. _relay_port() falls back through
ANDROID_RELAY_PORT → RELAY_PORT → 8767. android_setup() rewritten —
no longer imports the deleted android_relay module, instead probes
/health to verify the unified relay is up.
- plugin/android_relay.py + plugin/tools/android_relay.py: DELETED.
- plugin/tests/test_bridge_channel.py: new unittest suite (7 tests,
all passing) covering envelope routing, future resolution, timeout
cleanup, disconnect cleanup, send-failure cleanup, and the legacy-
timeout regression guard. Run with:
python -m unittest plugin.tests.test_bridge_channel
- DEVLOG.md + CLAUDE.md: Phase 3 / Wave 1 / α entry + Key Files row for
plugin/relay/channels/bridge.py. Repo layout block trimmed
android_relay.py.
Judgment call: bridge HTTP routes are unauthenticated at the HTTP layer,
matching the legacy relay. Trust boundary unchanged (localhost-only);
disconnected phone naturally 503s every call; bridge grant already
tracked in Session.grants["bridge"] so Wave 2 safety-rails can wrap
without touching this handler.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
8f61262cf0 |
feat(chat): inbound media pipeline — relay MediaRegistry + phone fetcher + Discord-style rendering
Closes the gap where tool-produced files (screenshots today, video/audio/PDF
in the future) were leaking into chat as raw "MEDIA:/tmp/..." text instead
of rendering inline.
Root cause is upstream: APIServerAdapter.send() in hermes-agent's
gateway/platforms/api_server.py is an explicit no-op, and
_write_sse_chat_completion streams raw deltas without ever invoking
extract_media(). Upstream's extract_media() / send_document() pipeline
only fires for push platforms (Telegram, Feishu, WeChat). Our HTTP pull
adapter has no file-delivery path at all.
Fix: plugin-owned file-serving on the relay, opaque-token markers in chat
text, out-of-band bearer-auth'd HTTPS fetch for bytes. Zero LLM context
cost (token is ~25 chars; bytes never travel through the chat stream).
## Server (Python)
- plugin/relay/media.py — MediaRegistry (asyncio.Lock-guarded OrderedDict
LRU, 24h TTL, 500-entry cap, 100 MB per-file cap), _MediaEntry,
MediaRegistrationError. Path sandboxing: absolute, realpath-resolves
under an allowed root (default tmpdir + HERMES_WORKSPACE +
RELAY_MEDIA_ALLOWED_ROOTS), exists, regular file, under size cap.
- plugin/relay/client.py — stdlib urllib.request register_media() helper
for in-process tool callers.
- plugin/relay/server.py — handle_media_register (loopback-only, mirrors
/pairing/register trust model) + handle_media_get (Bearer auth via the
existing SessionManager, web.FileResponse stream with Content-Type +
Content-Disposition). Routes registered in create_app.
- plugin/relay/config.py — 4 new env vars: RELAY_MEDIA_MAX_SIZE_MB,
RELAY_MEDIA_TTL_SECONDS, RELAY_MEDIA_LRU_CAP, RELAY_MEDIA_ALLOWED_ROOTS.
- plugin/tools/android_tool.py + plugin/android_tool.py — android_screenshot()
calls register_media() after writing the temp file, emits
MEDIA:hermes-relay://<token> on success. On any failure (relay down,
timeout) falls back to bare MEDIA:<path> with a warning; the phone
handles the bare form as an "unavailable" placeholder.
- plugin/tests/test_media_registry.py — 11 tests (happy path, TTL expiry,
LRU eviction + reorder on get, relative/nonexistent/directory path
rejection, allowed-roots whitelist, symlink-escape rejection, oversize,
empty content_type).
- plugin/tests/test_relay_media_routes.py — 8 tests (/media/register
loopback gate, happy path, validation 400, bad JSON; /media/{token}
401 without/with bad bearer, 200 streams bytes, 404 expired, 404 unknown).
## Phone (Kotlin)
- network/RelayHttpClient.kt — OkHttp GET /media/{token}, Bearer auth,
ws→http URL rewrite, Content-Disposition filename parse.
- data/ChatMessage.kt — AttachmentState (LOADING/LOADED/FAILED),
AttachmentRenderMode (IMAGE/VIDEO/AUDIO/PDF/TEXT/GENERIC), extended
Attachment with state/errorMessage/relayToken/cachedUri. Outbound
attachments default to LOADED (backward-compat).
(NB: also fixes a Kotlin block-comment nesting bug — the KDoc used
"text/*" as a MIME wildcard, which Kotlin's nested-comment lexer treated
as an unclosed /* opening and swallowed the rest of the file. Rephrased.)
- data/MediaSettings.kt — DataStore-backed: maxInboundSizeMb (25),
autoFetchThresholdMb (2, persisted-not-enforced placeholder),
autoFetchOnCellular (off), cachedMediaCapMb (200).
- util/MediaCacheWriter.kt — LRU-capped cache at cacheDir/hermes-media/,
mtime eviction, MIME→extension map, returns FileProvider content:// URIs.
- network/handlers/ChatHandler.kt — mediaRelayRegex + mediaBarePathRegex,
scanForMediaMarkers called unconditionally from onTextDelta (not gated
on parseToolAnnotations), dispatchedMediaMarkers dedupe set, new
onMediaAttachmentRequested + onUnavailableMediaMarker var callbacks,
mutateMessage helper exposed so ChatViewModel can flip attachment state.
- viewmodel/ChatViewModel.kt — initializeMedia wiring, LOADING placeholder
insertion on marker dispatch, fetch via RelayHttpClient, size cap check
post-download, cache via MediaCacheWriter, state flip to LOADED/FAILED.
Cellular gate encoded as LOADING + errorMessage="Tap to download" (no
new enum value needed); manualFetchAttachment() retries ignoring the gate.
- viewmodel/ConnectionViewModel.kt — owns media singletons, shared
OkHttpClient, cached-cap mirror loop so the writer's cap lambda is
synchronous.
- ui/RelayApp.kt — initializeMedia wired inside the existing
LaunchedEffect(apiClient) block.
- ui/components/MessageBubble.kt — attachments loop dispatches through
InboundAttachmentCard regardless of direction (no separate outbound
render path).
- ui/components/InboundAttachmentCard.kt — single Compose component
dispatched on (state × renderMode). IMAGE decodes from cachedUri via
BitmapFactory + asImageBitmap (matches existing outbound image path,
no Coil/Glide added). VIDEO/AUDIO/PDF/TEXT/GENERIC render as tap-to-open
cards firing ACTION_VIEW with FLAG_GRANT_READ_URI_PERMISSION.
- ui/screens/ChatScreen.kt — empty-bubble skip respects
attachments.isNotEmpty, wires manualFetchAttachment to retry + manual-fetch.
- ui/screens/SettingsScreen.kt — new InboundMediaSection between Chat and
Appearance (max size, auto-fetch threshold, cellular toggle, cached cap,
clear button). Coexists with the Connection UX landed in the previous
commit: both features share this file and this commit is where their
shared edits land.
- res/xml/file_provider_paths.xml + AndroidManifest.xml — FileProvider
declaration with authority ${applicationId}.fileprovider.
## Docs
- DEVLOG.md — full session entry with root cause, design, files, known
gaps, test plan.
- docs/decisions.md — ADR 14 on plugin-owned media endpoint, trust model,
resource bounds, alternatives rejected.
- docs/spec.md — new §6.2a Inbound Media covering wire contract, server
routes, phone parse/fetch/render flow, known gaps.
- docs/relay-server.md + user-docs/reference/relay-server.md — new
RELAY_MEDIA_* env vars, new /media/register + /media/{token} routes.
- user-docs/reference/configuration.md — new "Inbound Media Settings"
section with honest notes on the two known gaps (auto-fetch threshold
not enforced, session replay breaks across relay restarts).
- CLAUDE.md — Key Files + Integration Points updated.
## Known gaps (filed as DEVLOG follow-ups)
- Auto-fetch threshold slider is persisted but not enforced today —
only the cellular toggle + the hard max cap actually gate fetches.
- Session replay breaks across relay restarts. MediaRegistry is in-memory;
phone-side persistent cache (indexed by token or content hash) is the
right layer for durability and is out of scope for this pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
e9c3f824ba |
feat: /hermes-relay-pair skill + rename plugin hermes-android → hermes-relay
Three parallel workstreams landed together.
## 1. Skill authoring — new /hermes-relay-pair slash command
skills/hermes-relay-pair/SKILL.md (98 lines, agentskills.io-compatible):
- Proper YAML frontmatter (name, description, version, author, license,
platforms: [linux, macos], metadata.hermes tags/category/homepage)
- Body sections: When to Use / Prerequisites / Procedure / Pitfalls /
Verification — terse imperatives suitable for the agent's context window
- Tells the agent to run `python -m plugin.pair`, explains the venv path
trap, walks through host-side verification (curl /health, clients
count), warns about code expiration + QR terminal rendering gotchas
- Once installed to ~/.hermes/skills/, auto-registers as the
/hermes-relay-pair slash command in every chat session + messaging
platform.
## 2. Rename hermes-android → hermes-relay (user-facing, not Python)
The repo was rebranded on 2026-04-08 but the Python *package* / plugin
distribution name still said `hermes-android`. It's visible in
`hermes plugins list`, pypi-style imports, and a dozen user-docs /
README / install-snippet references. Renamed everywhere it's a package
name or user-visible string; Python module path `plugin` stays exactly
as-is (imports unchanged).
- pyproject.toml: name "hermes-android" → "hermes-relay",
version "0.4.0" → "0.5.0", description rewritten
- plugin/plugin.yaml: name + version aligned with pyproject
- plugin/__init__.py: docstring updated
- install.sh: PLUGIN_NAME target dir is now ~/.hermes/plugins/hermes-relay
- README.md, AGENTS.md, docs/{security,relay-server,upstream-contributions}.md,
user-docs/guide/getting-started.md, user-docs/reference/{relay-server,configuration}.md:
install snippets + package references swept
- plugin/{android_tool,tools/android_tool}.py: package name strings updated
- relay_server/SKILL.md, skills/hermes-pairing-qr/*: deprecation notices
updated to reference the new name
- Historical DEVLOG entries, plan.md build plan, and Python import
paths left untouched on purpose (history and correctness)
## 3. user-docs additive pass for the slash command
user-docs/guide/getting-started.md and README.md now mention
/hermes-relay-pair as the primary "from a Hermes chat session" path
alongside the `hermes pair` CLI. Narrow additive edits only — no
section rewrites.
## Upstream gap discovered
Hermes v0.8.0 has `PluginContext.register_cli_command()` for third-party
plugins but main.py only wires `plugins.memory.discover_plugin_cli_commands()`
into the top-level argparser — the generic `_cli_commands` dict is
populated correctly but never consulted. Result: our `hermes pair` and
`hermes relay` sub-commands register cleanly but aren't callable from
the CLI until an upstream fix lands. The `register_cli_command` calls
in plugin/__init__.py are left in place so they'll start working the
moment main.py is patched.
Practical workaround: use /hermes-relay-pair (this skill) or a shell
shim at ~/.local/bin/hermes-pair that execs `python -m plugin.pair`
(deploy step — not in this commit).
Verified locally:
pip install -e . → hermes-relay 0.5.0
python -c "import plugin.pair"
python -m plugin.pair --help
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
5c03efc4ee |
refactor: restructure repo — promote Android project to root
Layout change: - companion-app/ → root (app/, build.gradle.kts, gradlew, etc.) - hermes-android-bridge/ → removed (absorbed into Compose rewrite) - hermes-android-plugin/ → plugin/ - root tools/, skills/, tests/ → plugin/tools/, plugin/skills/, plugin/tests/ - root setup.py, pyproject.toml, requirements.txt → plugin/ - companion_relay/ → unchanged (Python package name = dir name) - CI workflows updated for new paths Android Studio now opens the repo root directly. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |