From 2b700730e553ca984b7dc347f08b4c4472443b92 Mon Sep 17 00:00:00 2001 From: ruvnet Date: Sun, 19 Jul 2026 05:32:32 +0000 Subject: [PATCH] deploy: 47dbdb29c040d81dd57b4aed8a2a7f88cd8810a3 --- ...021-vital-sign-detection-rvdna-pipeline.md | 6 + ...R-046-android-tv-box-armbian-deployment.md | 2 +- api-docs/adr/ADR-116-cog-ha-matter-seed.md | 51 ++ .../ADR-127-homecore-state-machine-rust.md | 74 +++ .../adr/ADR-129-homecore-automation-engine.md | 17 + ...R-131-homecore-ui-operational-dashboard.md | 444 ++++++++++++++++++ ...mecore-recorder-history-semantic-search.md | 36 ++ api-docs/adr/ADR-133-homecore-assist-ruflo.md | 68 +++ ...-fusion-engine-quality-scoring-evidence.md | 31 ++ ...privacy-control-plane-modes-attestation.md | 50 ++ ...51-room-calibration-specialist-training.md | 48 ++ .../adr/ADR-154-signal-dsp-beyond-sota.md | 2 + .../ADR-161-homecore-server-layer-security.md | 71 +++ ...65-homecore-migrate-from-home-assistant.md | 21 +- ...i-core-csi-deserialiser-security-review.md | 117 +++++ ...etric-locked-pck-mpjpe-accuracy-harness.md | 123 +++++ ...ci-bench-regression-compile-verify-gate.md | 110 +++++ ...8-quantization-half-pose-model-measured.md | 172 +++++++ ...uview-swarm-nan-fail-open-safety-review.md | 103 ++++ ...DR-177-nvsim-degenerate-input-hardening.md | 92 ++++ ...njection-and-capability-least-privilege.md | 87 ++++ ...cworld-candle-checkpoint-load-hardening.md | 81 ++++ ...-182-npx-ruview-harness-via-metaharness.md | 279 +++++++++++ ...onboard-led-gamma-stimulus-csi-colormap.md | 98 ++++ .../adr/ADR-262-rufield-ruview-integration.md | 12 +- ...263-rtl8720f-2-4ghz-fmcw-radar-platform.md | 171 +++++++ .../ADR-263-ruview-npm-harness-deep-review.md | 191 ++++++++ .../ADR-264-rtl8720f-radar-wire-protocol.md | 148 ++++++ ...264-rvagent-mcp-and-cli-npm-deep-review.md | 169 +++++++ ...DR-265-ruview-npm-distribution-strategy.md | 124 +++++ .../ADR-266-mediatek-filogic-csi-platform.md | 75 +++ ...ADR-267-mediatek-mimo-csi-wire-protocol.md | 61 +++ .../ADR-268-qualcomm-atheros-csi-platform.md | 41 ++ .../adr/ADR-269-qualcomm-csi-wire-protocol.md | 40 ++ ...0-vendor-rf-sensing-integration-program.md | 122 +++++ api-docs/adr/README.md | 10 +- .../ddd/deployment-platform-domain-model.md | 6 +- api-docs/releases/v0.9.0-realtek-beta.1.md | 45 ++ api-docs/releases/v0.9.1-mediatek-beta.1.md | 32 ++ api-docs/releases/v0.9.2-qualcomm-beta.1.md | 22 + .../v0.9.3-vendor-providers-beta.1.md | 21 + .../research/sota-nn-train-benchmark-brief.md | 147 ++++++ api-docs/user-guide.md | 36 ++ api-docs/vendor-rf-providers.md | 60 +++ 44 files changed, 3708 insertions(+), 8 deletions(-) create mode 100644 api-docs/adr/ADR-131-homecore-ui-operational-dashboard.md create mode 100644 api-docs/adr/ADR-172-cli-core-csi-deserialiser-security-review.md create mode 100644 api-docs/adr/ADR-173-metric-locked-pck-mpjpe-accuracy-harness.md create mode 100644 api-docs/adr/ADR-174-ci-bench-regression-compile-verify-gate.md create mode 100644 api-docs/adr/ADR-175-int8-quantization-half-pose-model-measured.md create mode 100644 api-docs/adr/ADR-176-ruview-swarm-nan-fail-open-safety-review.md create mode 100644 api-docs/adr/ADR-177-nvsim-degenerate-input-hardening.md create mode 100644 api-docs/adr/ADR-178-desktop-ipc-injection-and-capability-least-privilege.md create mode 100644 api-docs/adr/ADR-179-occworld-candle-checkpoint-load-hardening.md create mode 100644 api-docs/adr/ADR-182-npx-ruview-harness-via-metaharness.md create mode 100644 api-docs/adr/ADR-183-onboard-led-gamma-stimulus-csi-colormap.md create mode 100644 api-docs/adr/ADR-263-rtl8720f-2-4ghz-fmcw-radar-platform.md create mode 100644 api-docs/adr/ADR-263-ruview-npm-harness-deep-review.md create mode 100644 api-docs/adr/ADR-264-rtl8720f-radar-wire-protocol.md create mode 100644 api-docs/adr/ADR-264-rvagent-mcp-and-cli-npm-deep-review.md create mode 100644 api-docs/adr/ADR-265-ruview-npm-distribution-strategy.md create mode 100644 api-docs/adr/ADR-266-mediatek-filogic-csi-platform.md create mode 100644 api-docs/adr/ADR-267-mediatek-mimo-csi-wire-protocol.md create mode 100644 api-docs/adr/ADR-268-qualcomm-atheros-csi-platform.md create mode 100644 api-docs/adr/ADR-269-qualcomm-csi-wire-protocol.md create mode 100644 api-docs/adr/ADR-270-vendor-rf-sensing-integration-program.md create mode 100644 api-docs/releases/v0.9.0-realtek-beta.1.md create mode 100644 api-docs/releases/v0.9.1-mediatek-beta.1.md create mode 100644 api-docs/releases/v0.9.2-qualcomm-beta.1.md create mode 100644 api-docs/releases/v0.9.3-vendor-providers-beta.1.md create mode 100644 api-docs/research/sota-nn-train-benchmark-brief.md create mode 100644 api-docs/vendor-rf-providers.md diff --git a/api-docs/adr/ADR-021-vital-sign-detection-rvdna-pipeline.md b/api-docs/adr/ADR-021-vital-sign-detection-rvdna-pipeline.md index fc85b25f..a89a57ad 100644 --- a/api-docs/adr/ADR-021-vital-sign-detection-rvdna-pipeline.md +++ b/api-docs/adr/ADR-021-vital-sign-detection-rvdna-pipeline.md @@ -1092,6 +1092,12 @@ Two robustness bugs were fixed in the on-device edge path (`firmware/esp32-csi-n Both are pinned by host-buildable C99 tests in `firmware/esp32-csi-node/test/test_vitals_count_presence.c` (`make run_vitals`). The exact thresholds are documented constants pending on-device calibration against ground truth. +### 2026-06 — Rust `wifi-densepose-vitals`: IIR filter NaN/inf self-heal (ADR-158 §A1) + +A correctness/safety review of the Rust extraction crate found a real bug parallel to the firmware robustness class above. The 2nd-order resonator `bandpass_filter` in both `breathing.rs` and `heartrate.rs` latches each output `y[n]` into its filter state (`y1`/`y2`). A single non-finite amplitude residual from a corrupt CSI frame produced a NaN `output` that was written into the state; the existing `extract()` `is_finite()` guard dropped that one sample from the history buffer **but never sanitized the poisoned filter state**, so every later output stayed NaN, was rejected too, and the sliding-window history never refilled — breathing **and** heart-rate extraction went silently dead (returning `None` forever) until `reset()`. On the alert path this is a safety-relevant denial of service (one bad frame stops vitals monitoring with no error surfaced). + +Fix: when `bandpass_filter` computes a non-finite `output`, it resets the IIR state to default and returns `0.0`, so the resonator self-heals on the next clean frame (the `0.0` is still dropped by the caller's finite-check, so no spurious sample enters history). Same shape as the calibration NaN bug (ADR-154 §3) — the prior hardening guarded the *history boundary* but not the *filter-state boundary*. Pinned by `breathing::tests::nan_frame_does_not_permanently_poison_filter`, `breathing::tests::inf_mid_stream_does_not_freeze_history`, and `heartrate::tests::nan_frame_does_not_permanently_poison_filter` (all FAIL pre-fix, verified by reverting). The review also de-magicked the HR physiological plausibility band into named `HR_PLAUSIBLE_MIN_BPM`/`HR_PLAUSIBLE_MAX_BPM` consts (value-identical 40/180 BPM) and added a fabricated-vital negative (`pure_noise_is_never_reported_valid` — broadband noise never yields a clinically `Valid` HR; the extractor honestly returns low-confidence `Unreliable`). Clean dimensions confirmed with evidence: flat/silent input → `None`; pure noise → low-confidence `Unreliable`, never `Valid`; harmonic-rich breathing with no cardiac component → low-confidence, not a confident false HR; out-of-band BPM rejected by the plausibility clamp. + ## References - Ramsauer et al. (2020). "Hopfield Networks is All You Need." ICLR 2021. (ModernHopfield formulation) diff --git a/api-docs/adr/ADR-046-android-tv-box-armbian-deployment.md b/api-docs/adr/ADR-046-android-tv-box-armbian-deployment.md index 380d493e..52a85c1e 100644 --- a/api-docs/adr/ADR-046-android-tv-box-armbian-deployment.md +++ b/api-docs/adr/ADR-046-android-tv-box-armbian-deployment.md @@ -83,7 +83,7 @@ This ADR covers Phase 1 (TV box as aggregator) and Phase 2 (custom WiFi firmware |---------|--------|-------------|--------------|--------| | Broadcom BCM43455 | brcmfmac | **Proven** (Nexmon CSI) | Yes | Low — patches exist | | Realtek RTL8822CS | rtw88 | **Moderate** — driver is open-source, CSI hooks need adding | Yes (patched) | Medium | - | MediaTek MT7661 | mt76 | **Unknown** — MediaTek has released CSI tools for some chips | Yes | Medium-High | + | MediaTek MT7661 | mt76 | **Unverified** — no supported public CSI capture API was found in upstream `mt76` or public MediaTek SDK material | Yes | Research only | 2. **CSI extraction architecture** (Linux kernel driver modification): diff --git a/api-docs/adr/ADR-116-cog-ha-matter-seed.md b/api-docs/adr/ADR-116-cog-ha-matter-seed.md index c6919082..354c1a35 100644 --- a/api-docs/adr/ADR-116-cog-ha-matter-seed.md +++ b/api-docs/adr/ADR-116-cog-ha-matter-seed.md @@ -104,6 +104,57 @@ Ranked by build cost × user impact: | **P9** | HACS integration repo (`hass-wifi-densepose`) for HA-side install path | pending | | **P10** | Witness bundle + CSA-style spec compliance check | pending | +## 4.1 Crypto/security review notes (§2.2 witness chain — ADR-262 P2 prerequisite) + +Beyond-SOTA crypto+security review of the SHA-256 + Ed25519 witness chain +(`witness.rs` / `witness_signing.rs`) and the manifest signature surface +(`manifest.rs`), because ADR-262 P2 proposes to **reuse this exact signing +chain**. Top priority was the sibling `wifi-densepose-engine` bug class — +unframed boundary-to-boundary concatenation of operator-influenceable strings +into a signed/hashed digest. + +- **Engine bug class ABSENT (good result, reported with byte evidence).** + `canonical_bytes` is `DOMAIN_TAG ‖ prev_hash[32] ‖ seq:u64-be ‖ ts:u64-be ‖ + kind_len:u32-be ‖ kind ‖ payload_len:u32-be ‖ payload`. The two + variable-length operator-influenceable fields (`kind`, `payload`) are + **length-prefixed**; the fixed-width fields are self-delimiting → the + encoding is injective (no two distinct event tuples share a preimage). The + Ed25519 signature signs the **identical** bytes the SHA-256 chain commits to. + No separate unframed concatenation exists; the manifest `binary_signature` + is signed at build time (Makefile) over a single fixed-length `binary_sha256` + hex value, not in-crate. + +- **CHM-WIT-01 (FIXED) — domain-separation tag added.** The engine fix + prescribed *domain-tag + length-prefix*; length-prefix was present, the + domain tag was not. Added a versioned, NUL-terminated + `WITNESS_DOMAIN_TAG = b"cog-ha-matter/witness-event/v1\x00"` prefix so the + witness message can never be replayed as a message for another Ed25519 + context that shares key infrastructure (notably the manifest signature). + **Witness bytes change by design** (prior on-disk hashes/signatures + invalidated, as with the engine fix); verified safe because no in-repo crate + consumes cog-ha-matter witness bytes programmatically (doc-mentions only). + +- **CHM-WIT-02 (HARDENED) — `verify_signature` now uses `verify_strict`.** For + an audit chain the signature is the attestation, so non-canonical encodings + and small-order keys are rejected (RFC 8032 strict), giving the "one + canonical signature per event" property. Not a forgery fix — the verifying + key is caller-pinned, never read from the event. + +- **Confirmed clean (with evidence):** verify-before-trust + key-pinning + (`verify_signature` takes the verifying key as a parameter; `read_jsonl` + re-derives every hash and chain-verifies); key handling (the crate never + generates/stores/logs/serializes a signing key — only a documented test-only + fixed seed; production keys come from the Seed secure store, out of scope); + determinism (positional bytes, deterministic Ed25519, alphabetically-locked + JSONL field order, sorted TXT records — no HashMap/float nondeterminism feeds + any digest); fail-closed parsing (structured errors, no panics; `main.rs` + reads no untrusted files/paths). + +Tests: `cog-ha-matter --no-default-features` 64 → **68**, 0 failed (CHM-WIT-01 +pinned by 4 fails-on-old tests across `witness.rs`/`witness_signing.rs`; +CHM-WIT-02 guarded by a key-pinning test). Python deterministic proof +unchanged (cog-ha-matter is off the signal proof path). + ## 5. References - ADR-101 — `cog-pose-estimation` packaging precedent (signed binaries on GCS, .cog manifest) diff --git a/api-docs/adr/ADR-127-homecore-state-machine-rust.md b/api-docs/adr/ADR-127-homecore-state-machine-rust.md index 1875b890..1887ed3f 100644 --- a/api-docs/adr/ADR-127-homecore-state-machine-rust.md +++ b/api-docs/adr/ADR-127-homecore-state-machine-rust.md @@ -190,4 +190,78 @@ The entity registry is a `RwLock>` backed by an a - `v2/crates/wifi-densepose-sensing-server/src/main.rs` — Axum + Tokio architecture pattern used throughout the existing server stack - `docs/adr/ADR-126-ruview-native-ha-port-master.md` — HOMECORE master; §5.5 crate naming; §6 compatibility contract; §5.1 RUVIEW-POLICY + +--- + +## 9. Security & concurrency review (P1 core, beyond-SOTA sweep) + +Foundational review of the `homecore` crate — the state store + event bus + +service/entity registries every other HOMECORE module trusts. Same rigor as +the ADR-129/130/132/133/161 sibling reviews. **Three real fixes (one +concurrency, two hardening), each pinned by a fails-on-old test; the bus-lag +and lock-discipline dimensions confirmed clean with evidence.** + +- **HC-RACE-01 (state-set TOCTOU — lost / reordered `state_changed`, the + crux). FIXED.** `StateMachine::set` did `get()` (releasing the DashMap + shard lock) → compute the next snapshot + the no-op / `last_changed` + decision → `insert()` (re-acquiring the lock) → `send()`. The + read-modify-write was **not atomic** w.r.t. a concurrent writer on the + same entity, contradicting §2.1's promise that "the writer atomically + replaces the map entry." A writer that read a stale `old` could + mis-classify a genuine transition as a no-op and **drop its + `state_changed` event** (a missed automation trigger) or fire an event + whose `new_state` duplicated the previously delivered one (a spurious + trigger for any automation keyed on `old_state != new_state`). **Fix:** + hold the shard write-lock across the entire read→decide→insert→fire + sequence via `entry()`/`insert_entry()`; `tx.send` is non-blocking, + non-async, and never re-enters the map, so firing under the shard lock + cannot deadlock and keeps global event order in lock-step with global + commit order. Pinned by `concurrent_set_fires_no_duplicate_adjacent_events` + (4 writers toggling one entity A/B; asserts no two consecutive fired + events carry an identical `new_state` — impossible under correct + serialisation; a probe observed ~93k such duplicate-adjacent events across + 200 trials on the racy code, zero on the fix). +- **HC-EID-LEN-01 (unbounded `entity_id` — memory-DoS at the REST boundary). + FIXED.** `homecore-api/src/rest.rs` parses untrusted path segments + straight through `EntityId::parse`; with no length cap, an + otherwise-valid id (`a.` + many MB of `[a-z0-9_]`) was accepted and a + `POST /api/states/` would persist it into the DashMap state store + (permanent growth across distinct ids). **Fix:** reject ids longer than + `MAX_ENTITY_ID_LEN` (255, HA-compatible) up front in `parse()`, before any + per-char scan, with a new `EntityIdError::TooLong`; fail-closed at the + boundary type protects every caller. Pinned by `entity_id_length_boundary` + (exactly-MAX accepted, MAX+1 and a 4 MiB id rejected — fails on old code). +- **HC-SVC-PANIC-01 (service-handler panic not isolated). HARDENED.** + `ServiceRegistry::call` already ran handlers outside the registry lock (no + `RwLock` poisoning, no blocking of other callers — clean), but a + panicking handler unwound through `call()` into the caller's task. **Fix:** + wrap the handler future in `AssertUnwindSafe` + `catch_unwind`, converting + a panic to `ServiceError::HandlerPanicked`; the registry stays fully + usable. Pinned by `panicking_handler_is_isolated_and_registry_survives`. + +**Dimensions confirmed clean (with evidence):** + +- **Event-bus bounds / lag (same class as the homecore-api WS lag-DoS).** + Both `StateMachine` and `EventBus` use bounded `tokio::sync::broadcast` + (capacity 4,096). A slow subscriber gets a recoverable `Lagged(n)` + (drop-oldest + re-sync); `fire_*` is non-blocking and **never waits on + slow receivers**, so a lagging subscriber cannot block the publisher, grow + the channel without bound, or take down a fast subscriber. Evidenced by + `slow_subscriber_does_not_block_publisher_or_kill_the_bus` (fire 3× + capacity at an idle subscriber; publisher unblocked, bus stays live). +- **Lock ordering / lock-across-await (deadlock).** No code path holds two + of `{state DashMap, registry RwLock, service RwLock}` simultaneously, so + no inconsistent-ordering deadlock can exist. Every `tokio::sync::RwLock` + guard in `registry.rs`/`service.rs` is used in a single synchronous + statement and dropped before any `.await`; `call` explicitly scopes the + read guard out before awaiting the handler. The only guard held across a + send is the DashMap shard lock in `set`, across a synchronous + (non-await) broadcast send — safe. +- **Panic-on-input.** No reachable `unwrap`/`expect`/index in non-test code + beyond the safe `send().unwrap_or(0)` and the dead-but-harmless + `split_once(...).unwrap_or(...)` fallbacks on already-validated ids. + +`cargo test -p homecore --no-default-features`: **20 → 24 passed, 0 failed** +(+4 pins). Workspace green; Python deterministic proof unchanged +(`f8e76f21…46f7a`, bit-exact — `homecore` is off the signal proof path). - `docs/adr/ADR-028-esp32-capability-audit.md` — witness chain pattern (Ed25519 per state transition) diff --git a/api-docs/adr/ADR-129-homecore-automation-engine.md b/api-docs/adr/ADR-129-homecore-automation-engine.md index f3085ce9..ec7fe99a 100644 --- a/api-docs/adr/ADR-129-homecore-automation-engine.md +++ b/api-docs/adr/ADR-129-homecore-automation-engine.md @@ -190,6 +190,23 @@ This is the same Wasmtime host already used for integration plugins (ADR-128) --- +## 8a. Security review (beyond-SOTA sweep, post ADR-154–159) + +A focused security review of `homecore-automation` (the execution/eval surface — triggers → conditions → actions, with templates) was run after the ADR-154–159 sweep, applying the same rigor that the sibling engine/bfld/calibration/vitals/geo reviews used. **Two real DoS findings, each pinned by a fails-on-old test; the condition-bypass, fail-closed-parsing, and action-authorization dimensions were probed and found clean.** + +- **HC-SEC-01 (template-injection / unbounded-expansion DoS, HIGH) — FIXED.** A `template:` condition / `value_template` is user automation config, and was rendered with MiniJinja's defaults: **no instruction budget, no output cap**. A single condition such as `{% for i in range(5000) %}{% for j in range(5000) %}xxxx{% endfor %}{% endfor %}` rendered a **100 MB string over ~11 s on one render call** (measured) — a CPU/memory denial of service (the bfld-class "unbounded expansion"; MiniJinja's per-call `range()` 10k cap does **not** stop nested loops). **Fix:** enable MiniJinja's `fuel` feature and set a per-render budget (`set_fuel(Some(1_000_000))`) so a nested loop burns one unit per iteration — the attack now fails fast (~90 ms) with "engine ran out of fuel"; plus a 64 KiB source-length cap rejecting pathological sources before compilation. Legitimate HA templates (a few dozen instructions) are unaffected. Pinned by `nested_loop_template_is_bounded_not_unbounded_dos`, `single_huge_repeat_template_is_bounded`, `oversized_template_source_is_rejected` (all fail-on-old: unbounded render / no rejection), and `legitimate_template_still_renders_within_fuel` (no regression). +- **HC-SEC-02 (panic-on-config DoS, MEDIUM) — FIXED.** `Action::Delay { seconds }` and `Action::WaitForTrigger { timeout_seconds }` fed the user-supplied float straight into `Duration::from_secs_f64`, which **panics** on negative, NaN, infinite, or overflowing inputs — all reachable from a crafted (or typo'd) YAML (`delay: {seconds: -1}`, `.nan`, `.inf`, `1e308`). One hostile config aborts the spawned automation run task with a panic (measured: "cannot convert float seconds to Duration: value is negative"). **Fix:** a `safe_duration_from_secs` guard that saturates instead of panicking (NaN/±inf/negative → `Duration::ZERO`, matching HA's lenient "non-positive delay = no delay"; absurdly large → clamped to ~100 years). Pinned by `delay_negative_seconds_does_not_panic`, `delay_nan_seconds_does_not_panic`, `delay_infinite_seconds_does_not_panic`, `wait_for_trigger_negative_timeout_does_not_panic`, `safe_duration_saturates_hostile_values` (incl. overflow clamp). + +**Dimensions confirmed clean (with evidence):** +- **Condition bypass / fail-closed eval** — a `Condition::Template` whose render errors evaluates to `false` (`condition.rs` `Err(_) => false`), and a `Choose` branch condition that fails to deserialize is treated as **non-matching** (the branch is skipped), not silently passing (`action.rs` `ChoiceBranch::matches` `Err(_) => return false`). Both fail **closed** (do-not-run), confirmed by the existing `choose_*` tests and template-false-blocks-action behavioral test. No true-by-default-on-parse-error path found. +- **Re-entrancy / livelock (DoS)** — run-mode machinery is bounded and tested: `Single`/`IgnoreFirst` re-entrancy guard, `Restart` cancel-and-replace, `Queued` FIFO serialization, and `max: N` semaphore cap (ADR-162; `restart_mode_cancels_prior_run`, `queued_mode_runs_sequentially_not_concurrently`, `max_two_caps_concurrency_at_two`, `single_mode_does_not_double_fire_on_rapid_triggers`). A self-triggering automation does not livelock the engine — each fire is bounded by its run-mode. +- **Action authorization** — templates are read-only sandboxed (`states`/`state_attr`/`is_state`/`now` globals; no service-call or state-set global is exposed to template scope), so a template cannot escalate into an action. Service authorization itself is enforced at the `homecore` service-registry boundary (out of this crate's scope); no gap found in what the automation crate enforces. +- **Panic-on-config (parse)** — `serde_yaml`/`serde_json` deserialization returns structured `AutomationError` (no `unwrap`/`expect`/index reachable from a crafted config in the eval/exec path); the only remaining panic surface was the `from_secs_f64` path fixed as HC-SEC-02. + +Validation: `cargo test -p homecore-automation --no-default-features` → 54 passed / 0 failed (+14 over baseline). Python deterministic proof unchanged (homecore-automation is off the signal-processing proof path). + +--- + ## 9. References ### HA upstream diff --git a/api-docs/adr/ADR-131-homecore-ui-operational-dashboard.md b/api-docs/adr/ADR-131-homecore-ui-operational-dashboard.md new file mode 100644 index 00000000..3e73efb9 --- /dev/null +++ b/api-docs/adr/ADR-131-homecore-ui-operational-dashboard.md @@ -0,0 +1,444 @@ +# ADR-131: HOMECORE-UI — Operational dashboard for the two-tier Cognitum stack + +| Field | Value | +|-------|-------| +| **Status** | Accepted — UI implemented (§10); full backend wiring specified (§11–§12) | +| **Date** | 2026-06-14 | +| **Deciders** | ruv | +| **Codename** | **HOMECORE-UI** — first-class operator dashboard inside the Cognitum Appliance shell | +| **Relates to** | [ADR-126](ADR-126-ruview-native-ha-port-master.md) (HOMECORE master), [ADR-127](ADR-127-homecore-state-machine-rust.md) (HOMECORE-CORE state machine), [ADR-128](ADR-128-homecore-integration-plugin-system.md) (HOMECORE-PLUGINS), [ADR-129](ADR-129-homecore-automation-engine.md) (automation engine), [ADR-130](ADR-130-homecore-rest-websocket-api.md) (HOMECORE-API), [ADR-132](ADR-132-homecore-recorder-history-semantic-search.md) (recorder/semantic search), [ADR-151](ADR-151-room-calibration-specialist-training.md) (room calibration HTTP API), [ADR-100](ADR-100-cog-packaging-specification.md) (Cog packaging), [ADR-116](ADR-116-cog-ha-matter-seed.md) (cog-ha-matter), [ADR-069](ADR-069-cognitum-seed-csi-pipeline.md) (SEED RVF ingest), [ADR-105](ADR-105-federated-csi-training.md) (federated CSI training) | +| **Tracking issue** | TBD | +| **Parent** | [ADR-126](ADR-126-ruview-native-ha-port-master.md) (sub-ADR, HOMECORE-127…134 family) | + +--- + +## 1. Context + +HOMECORE (ADR-126 through ADR-134) is the native Rust + WASM + TypeScript port of Home Assistant running as the hub on the Cognitum v0 Appliance. As of P2, the state machine ([ADR-127](ADR-127-homecore-state-machine-rust.md)), API ([ADR-130](ADR-130-homecore-rest-websocket-api.md)), and COG runtime ([ADR-128](ADR-128-homecore-integration-plugin-system.md)) are in place. What is missing is a first-class dashboard UI that operators, integrators, and residents can use to manage the full two-tier hardware stack that HOMECORE coordinates. + +### 1.1 The two-tier hardware model this UI must represent + +This is the most important architectural constraint the UI must carry through every panel: + +- **Cognitum SEED** — a Pi Zero 2 W-based edge node. It has its own RVF vector store (8-dim, content-addressed, with kNN queries), Ed25519 witness chain, SHA-256 ingest audit trail, onboard environmental sensors (BME280 temperature/humidity/pressure, PIR motion, reed switch, ADS1115 4-channel ADC, vibration), 13 drift detectors, an MCP proxy (114 tools, JSON-RPC 2.0, default-deny policy), 98 HTTPS API endpoints, and epoch-based swarm sync for multi-SEED deployments. SEEDs sit close to the ESP32 sensing nodes and receive feature vectors from them at 1 Hz. Multiple SEEDs can form a peer mesh. **This is the sensing and memory tier.** +- **Cognitum v0 Appliance** — a Pi 5 + Hailo-10H hub, running at `:9000`. It hosts the COG runtime (`/var/lib/cognitum/apps/`), the HOMECORE state machine and event bus, the calibration service, `ruview-mcp-brain:9876`, `cognitum-rvf-agent:9004`, `ruvector-hailo-worker:50051`, and acts as the fleet coordinator for multi-room correlation and federated training. The Appliance is where HOMECORE runs, and it is what the dashboard user is sitting in front of. **This is the computation and orchestration tier.** + +SEEDs are **subordinate nodes that the Appliance supervises** — they are not peers. The UI navigation hierarchy must reflect this: the Appliance is the root, SEEDs are children, ESP32 nodes are leaves. + +### 1.2 What the UI is not + +HOMECORE-UI is **not** a re-skin of the existing Cognitum Cog Store. It is a full operational dashboard that **extends** the Cognitum platform's shell — the Cog Store, API Explorer, and Guide already exist and must remain intact, with the HOMECORE dashboard added as a first-class navigation section alongside them. + +--- + +## 2. Decision + +Build HOMECORE-UI as a **complete** TypeScript + Rust→WASM frontend (per this ADR's §3 and the HOMECORE-127…134 family) that: + +1. Lives at `http://cognitum-v0:9000/homecore` (or as a dedicated nav item in the Cognitum Appliance shell). +2. Is visually and stylistically seamless with the existing Cognitum platform — same dark theme, same design tokens, same component patterns as `https://seed.cognitum.one/store`. +3. Drives the HOMECORE REST + WebSocket API ([ADR-130](ADR-130-homecore-rest-websocket-api.md)) and the calibration HTTP API ([ADR-151](ADR-151-room-calibration-specialist-training.md)) for all data. +4. Updates in real-time via the homecore `subscribe_events` WebSocket channel. **The UI must never poll for entity state.** + +**This is a decision to deliver the complete operational dashboard — every panel in §4.1 through §4.10, every navigation section in §5, fully wired to live data — not a design-system scaffold or a partial first cut.** A static layout shell with placeholder data is explicitly **out of scope as a deliverable**: the design system (§3) is a means to the complete UI, not an end in itself. The acceptance bar for this ADR is that an operator can drive the full two-tier stack — fleet, entities, rooms, COGs, calibration, events, audit, and settings — from the dashboard, against real APIs, with no panel left as a stub. + +### 2.1 `homecore-server` is the single backend-for-frontend (BFF) gateway + +The data the dashboard needs is spread across **three backend tiers that are not one process**: (a) `homecore-api` (`/api/*` REST + `/api/websocket`, mounted in `homecore-server`); (b) the **calibration API** (`/api/v1/*`, served by a *separate* binary — `wifi-densepose calibrate-serve` / `wifi-densepose-sensing-server`); and (c) the **SEED device tier + appliance daemons** (RVF vector store, witness chain, onboard sensors, reflex rules, COG supervisor, federation), which are physically separate HTTPS services on the SEED nodes and the appliance. + +The browser must talk to **exactly one origin.** Therefore `homecore-server` is promoted to the **single BFF / API gateway** for HOMECORE-UI: it serves the static assets at `/homecore`, serves `homecore-api` at `/api/*`, and **adds a new `/api/homecore/*` namespace** that proxies and aggregates the calibration API and the SEED/appliance tiers server-side. The UI only ever issues same-origin requests; cross-service auth (SEED bearer tokens, calibration tokens) is held by the gateway and **never exposed to the browser**. This collapses the CORS/multi-port problem and gives one place to enforce the long-lived-access-token auth (§4.10). + +### 2.2 No mock data in production + +The in-browser mock layer that the first UI cut shipped behind DEMO banners (§7.1, prior revision) is **demoted to a dev-only fixture** gated behind an explicit `?demo=1` / `HOMECORE_UI_DEMO=1` flag. The production build wires **every** panel to a real gateway endpoint. The full endpoint contract and the backend work each panel needs are specified in **§11**; the staged path to get there is **§12**. A panel may show an empty/typed-error state when its upstream is down, but it must never silently render fabricated data. + +--- + +## 3. Design system — Cognitum platform conventions + +The implementor **must study `https://seed.cognitum.one/store` as the definitive design reference before writing a single line of CSS.** The existing platform's design tokens, extracted from production, are: + +### 3.1 Colour palette (CSS custom properties) + +| Token | Value | Role | +|---|---|---| +| `--bg` | `#0a0e1a` | page background (very dark navy) | +| `--bg2` | `#111627` | secondary background / nav strip | +| `--card` | `#171d30` | card / panel surface | +| `--card-h` | `#1e2540` | card hover state | +| `--border` | `#252d45` | all border strokes (≈0.67px, subtle) | +| `--t1` | `#e0e4f0` | primary text (near-white) | +| `--t2` | `#8890a8` | secondary / muted text | +| `--t3` | `#505872` | tertiary / disabled text | +| `--cyan` | `#4ecdc4` | primary action colour (Install buttons, live indicators, accents) | +| `--cyan-d` | `rgba(78,205,196,0.15)` | cyan tint background for status badges | +| `--green` | `#6bcb77` | success / online / healthy states | +| `--green-d` | `rgba(107,203,119,0.15)` | green tint background | +| `--amber` | `#d4a574` | warning / stale / degraded states | +| `--amber-d` | `rgba(212,165,116,0.15)` | amber tint background | +| `--red` | `#e06060` | error / offline / veto states | +| `--red-d` | `rgba(224,96,96,0.15)` | red tint background | +| `--purple` | `#a78bfa` | informational / epoch / chain indicators | +| `--purple-d` | `rgba(167,139,250,0.15)` | purple tint background | +| `--r` | `10px` | standard border radius on all cards and panels | + +### 3.2 Typography + +- `--font`: `'Segoe UI', system-ui, -apple-system, sans-serif` — all body and heading text. +- `--mono`: `'Cascadia Code', 'Fira Code', Consolas, monospace` — all entity IDs, API endpoints, hex values, JSON payloads, COG binary hashes. + +### 3.3 Component patterns (from the live Cog Store and API Explorer) + +- **Cards**: `background: var(--card)`, `border: 0.67px solid var(--border)`, `border-radius: var(--r)`, `padding: 24px`. +- **Category pills / status badges**: small `border-radius: 4–6px`, uppercase text, coloured background tint (e.g. `background: var(--cyan-d); color: var(--cyan)` for `RUNNING`; `background: var(--amber-d); color: var(--amber)` for `STALE`). +- **Primary action buttons**: `background: var(--cyan)`, `color: var(--bg)`, no border — matching the existing "Install" button style exactly. +- **Secondary / ghost buttons**: transparent background, `border: 1px solid var(--border)`, `color: var(--t1)` — matching the existing "Details" button style. +- **Nav strip**: `background: var(--bg2)`, text items in `--t2`, active item highlighted in `--cyan` with a bottom underline. +- **Featured card gradient borders**: top-edge linear gradient from `var(--cyan)` to `var(--purple)` — replicate for HOMECORE section headers. +- **Live metric cards** (API Explorer status page): icon + large numeric value in `--cyan` or `--green`, label in `--t2` below, on a `var(--card)` background. +- **Method badge pills** on the API Explorer (`GET` in green, `POST` in amber, `AUTH` in purple) — reuse this same pill system for COG status indicators. + +The implementor **must not introduce new colours, typefaces, or border radii.** Every component should feel like it was built by the same team that built the Cog Store and the API Explorer. A user navigating from the Cog Store into the HOMECORE dashboard should not notice a visual seam. + +--- + +## 4. UI sections — required panels + +### 4.1 System Dashboard (the "home screen") + +The always-visible overview panel. Modelled on the API Explorer's live metric cards. All values update in real-time. + +- **v0 Appliance health strip** — reuse the exact metric-card pattern from `seed.cognitum.one/status`: one card each for CPU %, RAM usage, Hailo-10H inference load (% utilisation), Hailo temperature, uptime, and the running services (`ruview-mcp-brain:9876`, `cognitum-rvf-agent:9004`, `ruvector-hailo-worker:50051`). Values in `--cyan`, labels in `--t2`. This strip is always at the top — it represents the machine the user is looking at. +- **SEED Fleet overview** — a grid of SEED node cards (one per paired SEED) on the `var(--card)` surface with `var(--border)`. Each card shows: online/offline status pill (green/red), firmware version, epoch number, current vector count, last ingest timestamp, and witness-chain validity badge. A collapsed row shows the SEED's 5 onboard sensors in summary (PIR: yes/no, door: open/closed, temperature from BME280). Offline SEEDs render the entire card with a `--red-d` background tint. Clicking a SEED card navigates to the SEED Detail view (§4.2). +- **ESP32 Node summary** — count of active ESP32 nodes per SEED, current frame rate (target: 100 Hz CSI + 1 Hz feature vectors), and a compact warning list for nodes with known issues (presence_score normalisation anomaly, stale firmware version). +- **COG Runtime status row** — a horizontal strip of status pills for each installed COG on the v0 Appliance. Pill colours follow the existing badge convention: `--green-d`/`--green` for running, `--red-d`/`--red` for failed, `--t3`/`--t2` for stopped. COG name in `--mono`. Clicking a pill navigates to COG Management (§4.6). +- **Event Bus activity indicator** — a small real-time sparkline showing the homecore broadcast channel event rate (events/sec). Indicate channel lag if a subscriber is falling behind the 4,096-event capacity. + +### 4.2 SEED Detail View (per-SEED drill-down) + +Accessible from the fleet grid. Full-page panel for a single SEED node, using the card + section-header pattern from the Cog Store's detail views. + +- **SEED identity header** — `device_id` in `--mono`, firmware version, paired status in green, USB vs WiFi connection mode. A section-header gradient border (cyan → purple, matching the featured card style) visually separates this from Appliance content. +- **Vector Store panel** — current vector count, dimension (8), last kNN query latency, current epoch number, a small sparkline of ingest rate over the last hour, and a storage budget bar showing usage against the 100K working-set target. A "Compact now" button (`POST /api/v1/store/compact`) in ghost style. When usage exceeds 80%, the bar renders in `--amber`. +- **Witness Chain panel** — chain length (SHA-256 entries), last verification timestamp, a one-click "Verify chain" button (`POST /api/v1/witness/verify`), and an "Export attestation bundle" button for regulated deployments. The Ed25519 custody attestation (device-bound keypair, epoch + vector count + witness head) renders here. Chain length in `--purple`, following the existing epoch/chain colour convention. +- **Onboard Sensors panel** — live readings from all 5 sensors in individual sub-cards: BME280 (temperature °C, humidity %, pressure hPa), PIR (motion boolean with last-triggered timestamp), reed switch (open/closed with last-changed timestamp), ADS1115 (4 analog channels with configurable labels), vibration (boolean with last-triggered). These are ground-truth validators against CSI readings and are critical for diagnosing false positives in the mixture-of-specialists. Sensor values in `--cyan`; sensor names in `--t2`. +- **Reflex Rules panel** — the 3 pre-configured rules with current state: `fragility_alarm` (threshold 0.3 → relay actuator), `drift_cutoff` (threshold 1.0), `hd_anomaly_indicator` (threshold 200 → PWM brightness). Show last-fired time for each. The `fragility_alarm` threshold is the most commonly adjusted field and should be editable inline. Rules that have recently fired render with a `--amber-d` background tint. +- **Cognitive Analysis panel** — boundary fragility score (0.0–1.0, from Stoer-Wagner min-cut on the kNN graph) rendered as a progress bar: green below 0.3, amber 0.3–0.6, red above 0.6. High fragility (>0.3) indicates a regime change in the environment and should be visually prominent. Temporal coherence phase boundaries shown as a labelled timeline of detected environment state transitions. kNN graph rebuild cadence indicator (every 10 s). +- **Ingest pipeline status** — which ESP32 nodes feed this SEED, the packet type each is sending (`0xC5110003` native feature vectors vs `0xC5110002` vitals fallback path — distinguished visually since native is preferred), current ingest batch size, flush interval, and bridge path topology (direct vs host-laptop hop). The bridge-hop warning (known architectural limitation) renders in `--amber` since it adds a network hop. + +### 4.3 SEED Fleet Map (multi-SEED topology) + +For deployments with more than one SEED, a topology view showing the mesh: + +- **Node hierarchy diagram** — v0 Appliance at root, SEEDs as second tier (grouped by room/zone), ESP32 nodes as leaves under each SEED. Lines represent active data flows. ESP-NOW mesh sync links between SEEDs shown as dashed lines. Connection health shown via line colour (green/amber/red). All labels in `--mono`. +- **Cross-SEED event deduplication indicator** — for events that span multiple SEEDs (one fall detected by two rooms; one occupant tracked through room A → hallway → room B), show a fusion badge indicating how many SEEDs contributed to the composite event. +- **Federation config** ([ADR-105](ADR-105-federated-csi-training.md)) — federated-learning round coordinator role (which SEED is the round coordinator), current round number, K healthy nodes selected, delta exchange status. **Model deltas only — never raw CSI** is a design invariant that must be labelled explicitly in the UI. + +### 4.4 Entity & State Browser + +The homecore state machine (`DashMap>`) is the authoritative source of truth. Every COG running on the v0 Appliance contributes entities. + +- **Entity list by domain** — grouped by the `domain.` prefix of `EntityId`, using collapsible section headers. The 21 entities per ESP32 node (11 raw + 10 semantic primitives from `cog-ha-matter`) are the most important set. For each entity: current state string (in `--t1`), last-changed timestamp (in `--t3`), attribute map as collapsible JSON in `--mono`, and the Context (`user_id` + `parent_id` causality chain, critical for care/audit deployments). Entity IDs always in `--mono`. +- **SEED provenance badge** — each entity carries a small badge showing its data lineage: which ESP32 node → which SEED → which COG → homecore state machine. This trace is invaluable for debugging false positives and is a **first-class UI element, not a collapsed detail.** +- **Domain filter + semantic search** — filter by domain prefix and, once [ADR-132](ADR-132-homecore-recorder-history-semantic-search.md) (homecore-recorder) lands, ruvector-backed semantic search: "when did the living room anomaly score last correlate with a door-open event?" A keyword filter across entity IDs and attribute keys ships in the initial release regardless of [ADR-132](ADR-132-homecore-recorder-history-semantic-search.md) status, given entity density; the semantic search layers on top once the recorder lands. +- **Real-time WebSocket feed** — entity states update live via the homecore `subscribe_events` WebSocket command ([ADR-130](ADR-130-homecore-rest-websocket-api.md)). The UI must never poll. Show a broadcast-channel lag indicator; warn visually if the subscriber is falling behind the 4,096-event channel capacity. +- **StateChanged detail panel** — clicking any entity opens a slide-over panel showing the full `StateChangedEvent`: `old_state`, `new_state`, `context.id`, `context.user_id`, and the `context.parent_id` chain rendered as a breadcrumb trail. + +### 4.5 RoomState / Sensing Panel + +Surfaces the mixture-of-specialists output from the calibration service — the highest-level per-room sensing result. Data comes from `GET /api/v1/room/state?bank=` on the v0 Appliance. + +- **Per-room cards** — one card per `room_id` on the `var(--card)` surface. Each card shows live `RoomState` JSON fields as sub-rows: presence (occupied/absent chip in green/red with confidence bar), posture (standing/sitting/lying chip with confidence), breathing BPM (numeric in `--cyan` with range indicator 6–30), heart rate BPM (numeric in `--cyan` with range indicator 40–120), restlessness score (0–1 progress bar), and anomaly score (0–1 with normal/anomalous label, bar turns red above a configurable threshold). +- **STALE warning** — when `stale: true` (the specialist bank was trained against a different baseline), render the entire room card with a `--amber-d` background tint and a prominent amber banner reading "Bank stale — baseline has changed" with a direct "Recalibrate room" link into the calibration wizard (§4.7). This is the most common real-world failure mode and **must never be subtle.** +- **VETO indicator** — when `vetoed: true` (anomaly veto suppressed vitals/posture because the window was physically implausible), render the affected specialist slots in `--red` with a "Veto active" label. Values suppressed by veto **must not render as zeros** — they must render as explicitly withheld. +- **Null specialist placeholders** — specialists not yet trained (`null` in the specialist bank) render as "Not trained" placeholders in `--t3` with a small "Calibrate to enable" prompt in ghost style. They are **not** errors. +- **Confidence bars** — each specialist output has a confidence float, shown as a small inline bar (`--cyan` fill) next to the reading. Low confidence (< 0.4) renders the bar in `--amber`. +- **Multi-SEED fusion indicator** — for rooms served by multiple SEEDs, show a small badge indicating how many SEED nodes contributed to the `MultiNodeMixture` for this room's reading. + +### 4.6 v0 Appliance COG Management + +The v0 Appliance hosts COGs at `/var/lib/cognitum/apps/`. This panel is the operational companion to the existing Cog Store (`seed.cognitum.one/store`). It must match the Cog Store's visual conventions precisely — same card layout, same category pills, same install/detail button pair — because operators will move between the two surfaces. + +- **Installed COGs list** — for each COG: `id` and `version` in `--mono`, architecture badge (`arm`/`hailo10` etc., category-pill pattern), status pill (running/stopped/failed/updating in green/grey/red/amber), `binary_sha256` verified badge (Ed25519 signature verification shown as a shield icon in `--green` or `--red`), and PID from the pid file. Actions: start, stop, restart (ghost style), and view `output.log` / `error.log` in a monospace drawer using `--mono`. Edit `config.json` inline with syntax highlighting. +- **COG Store / App Registry** — browsable `app-registry.json` listing. This panel should visually mirror `seed.cognitum.one/store` as closely as possible — same featured-card hero layout, same icon + title + description + category pill + action button structure. One-click install downloads the binary from GCS, verifies `binary_sha256` + `binary_signature`, writes the manifest, and starts the COG. Show which new homecore entities will appear in the state machine after install, as a preview list before confirming. +- **OTA Updates** — a badge count on installed COGs with available updates, matching the "Installed (N)" tab badge convention from the existing Cog Store. Show a diff panel (version change, new entities, config schema changes) before confirming the update. +- **Hailo HEF status** — for COGs with `arch: hailo10`: loaded HEF files on the Hailo-10H, current inference throughput, and `ruvector-hailo-worker:50051` connection status. The RF Foundation Encoder ([ADR-150](ADR-150-rf-foundation-encoder.md)) and neural pose head display here once available. + +### 4.7 Calibration Wizard + +The full baseline → enroll → train → verify pipeline runs via HTTP against the v0 Appliance ([ADR-151](ADR-151-room-calibration-specialist-training.md)). This is a multi-step guided flow — not a raw API panel. Use a stepped wizard layout with a progress indicator at the top (steps 1–5 as numbered pills, active step in `--cyan`, completed in `--green`, pending in `--t3`). + +- **Step 1 — Select room and SEED** — enter a `room_id` name (validated against `[A-Za-z0-9_-]{1,64}`) and select which SEED(s) and ESP32 nodes serve this room from a dropdown populated from the live fleet. Show current CSI ingest health for the selected nodes inline — if frames are not arriving at the expected rate, display an amber warning **before** allowing the operator to proceed. A broken ingest pipeline will silently fail calibration. +- **Step 2 — Baseline capture** — `POST /api/v1/calibration/start`. A large full-width animated progress bar (cyan fill) reads from `GET /api/v1/calibration/status`: frames recorded vs target, ETA in seconds, `z_median` value. If `motion_flagged` is true, overlay an amber banner: "Room must be empty — movement detected." The baseline UUID produced here is the anchor for all future STALE detection for this room — display it in `--mono` once complete so operators can record it. +- **Step 3 — Anchor enrollment** — the 8 anchor labels in enforced order: `empty`, `stand_still`, `sit`, `lie_down`, `breathe_slow`, `breathe_normal`, `small_move`, `sleep_posture`. For each: a human-readable instruction with an illustration, a countdown timer rendered as a circular progress ring in `--cyan`, and an immediate quality-gate result (accepted in green, retry in amber with a reason string). Drive via `POST /api/v1/enroll/anchor` + `GET /api/v1/enroll/status`. After each accepted anchor, show the extracted feature values (mean, variance, breathing_score, heart_score) in a small `--mono` data row so operators can sanity-check the capture. Show overall progress as "N / 8 anchors accepted." +- **Step 4 — Train** — a single `POST /api/v1/room/train` call. Show the 6 specialist results as a checklist: presence (threshold + occupied_var), posture (prototype count), breathing (min_score), heartbeat (min_score), restlessness (calm/active motion values), anomaly (prototype count + scale). Specialists that returned non-null render in `--green`. Null specialists (insufficient anchor data) render in `--amber` with a "Re-enroll missing anchors" prompt linking back to Step 3 for the specific missing labels. +- **Step 5 — Verify live** — display the live `RoomState` for the just-trained room using the same per-room card layout as §4.5. Prompt the operator to stand in the room and verify presence is detected, try sitting/lying to confirm posture, and breathe normally to confirm vitals are in plausible range. A "Confirm and save" button (cyan, primary) closes the wizard; a "Something's wrong — re-enroll" button (ghost) loops back to Step 3. + +### 4.8 Event Bus & Automation Feed + +- **Live event stream panel** — a virtualized scrolling list of `SystemEvent` variants (`StateChanged`, `EntityRegistered`, `ConfigReloaded`) and notable `DomainEvent`s from the homecore Tokio broadcast channel. Each row shows: event-type pill (coloured by variant), `entity_id` in `--mono`, old state → new state arrow, timestamp, and `context.user_id`. The stream is filterable by entity domain, event type, or source SEED/COG. The filter bar uses the same search-input style as the Cog Store's search field. +- **Context causality breadcrumb** — expanding any event row shows the full Context chain (`context.id` → `parent_id` → `grandparent_id`) as a breadcrumb trail in `--mono`. This is how automation loops become visible without any separate debugging tool. +- **Automation builder** ([ADR-129](ADR-129-homecore-automation-engine.md) scope) — a trigger → condition → action editor on the card surface. The most important RuView-specific trigger types to support are: `state_changed` on `RoomState` entities with a threshold expression (e.g. `anomaly.value > 0.8`), SEED reflex-rule firing events (`fragility_alarm`, `hd_anomaly_indicator`), and custom `domain_event` topics. Actions include calling services in the homecore service registry and firing domain events. The condition expression editor uses `--mono`. + +### 4.9 Witness / Audit Log + +- **Unified witness timeline** — a chronological merged view of events from both tiers: the SEED's SHA-256 ingest chain (every RVF store write attested) and homecore's Ed25519 state-transition chain (biometric crossings, BFLD identity-risk elevations). Each row: `entity_id` in `--mono`, old/new state, timestamp, source SEED `device_id`, signing key fingerprint (first 8 chars in `--mono`). Pagination uses the same "Showing X–Y of Z" convention from the Cog Store's cog grid. +- **Privacy mode banner** — a persistent top-of-panel banner showing current privacy mode: `--green-d`/green text for full-publish mode; `--amber-d`/amber text for audit-only mode (SHA-256 digests on-SEED only, no MQTT state messages). Show the per-SEED privacy mode state, since SEEDs can be individually configured. Toggling privacy mode is a high-stakes action — require an explicit "Confirm" step with a summary of what will change. +- **Export bundle** — an "Export attestation bundle" button (ghost) that packages the SEED witness chain + homecore Ed25519 chain as a downloadable archive for regulated-deployment (care home, hotel, shared office) compliance handoff. + +### 4.10 Settings & Integration Config + +- **SEED fleet management** — add, remove, and reprovision SEEDs. Show the USB-only pairing requirement prominently (the pairing window only opens via `169.254.42.1`, not WiFi — a security invariant). Per-SEED: `device_id` in `--mono`, firmware version, bearer token status, and a "Rotate token" action (ghost) that walks the operator through the secure token rotation flow. +- **ESP32 node provisioning** — per-node NVS config display (target IP, target port, node_id), last-seen firmware version, and a link to the provisioning script. The `node_id` → room/zone assignment is editable here and persists to the room calibration system's `room_id` mapping. +- **MQTT / cog-ha-matter config** ([ADR-116](ADR-116-cog-ha-matter-seed.md)) — broker URL, credentials (masked), MQTT topic prefix, mDNS advertisement status (`_ruview-ha._tcp`), and a live connection indicator (green dot for connected, red for unreachable). The 21 HA-DISCO entities per node are listed here with their `via_device` assignments showing which SEED they belong to in HA's device registry. +- **Long-lived access tokens** — for homecore-api companion-app connections (HA 2025.1 wire-compat, [ADR-130](ADR-130-homecore-rest-websocket-api.md)). Token creation, last-used timestamp, and revocation. The HA companion-app pairing QR-code flow surfaces here. +- **Federation config** — for multi-SEED deployments: ESP-NOW mesh sync status, cross-SEED epoch alignment values, and federated-learning round settings (coordinator SEED, round cadence, Krum aggregation parameters per [ADR-105](ADR-105-federated-csi-training.md)). The design invariant **"model deltas only, never raw CSI"** must be labelled explicitly in this panel. + +--- + +## 5. Navigation structure + +HOMECORE-UI must integrate into the existing Cognitum Appliance nav shell. The top nav should read: + +``` +Framework | Guide | Cog Store | HOMECORE | Status +``` + +— inserting **HOMECORE** as a first-class nav item between the existing "Cog Store" and "Status" entries, using the same nav-item style (text in `--t2`, active state in `--cyan` with bottom underline). + +Within the HOMECORE section, a left sidebar (or top sub-nav on narrow viewports) provides section navigation: + +``` +Dashboard | SEED Fleet | Entities | Rooms | COGs | Calibration | Events | Audit | Settings +``` + +The COG Store panel within HOMECORE (§4.6) links out to `seed.cognitum.one/store` for the full catalog view, ensuring the existing Cog Store remains the canonical browsing experience. + +--- + +## 6. Key UX invariants + +These must be maintained across every panel: + +1. **Always make the tier origin of any data explicit.** A `RoomState` reading traces to an ESP32 node → SEED → COG → v0 Appliance state machine. The provenance badge (§4.4) must appear wherever entity states are displayed. +2. **The `stale` and `vetoed` flags from `RoomState` and the kNN fragility score from SEED cognitive analysis are meaningful diagnostic signals** — they must never be silently hidden, styled grey-on-grey, or collapsed behind an expand toggle. They represent system health operators need to act on. +3. **Values that are `null` because a specialist has not been trained must be visually distinct from values that are unavailable due to an error.** The distinction is operationally important: `null` means "calibrate to enable," unavailable means "investigate." +4. **All entity IDs, hashes, API endpoints, binary signatures, device UUIDs, and JSON payloads must use `--mono` font.** This is already the convention in the API Explorer and must be consistent throughout HOMECORE-UI. +5. **The v0 Appliance Hailo HAT is a separate subsystem from the SEED's edge compute.** Inference results tagged as Hailo-sourced (COGs with `arch: hailo10`) must be visually distinguished from results from CPU-only COGs (`arch: arm`) so operators can triage hardware-specific failures. + +--- + +## 7. Scope — complete UI delivery + +The deliverable is the **entire** dashboard. Every panel below ships fully implemented and wired to its live data source — there is no scaffold-only milestone and no panel left as a placeholder. The table records each panel's authoritative backing API so the build can proceed in whatever order best fits the dependency graph; it is a dependency map, **not** a sequence of partial releases. + +| Panel | Section | Backing API / source | +|---|---|---| +| System Dashboard | §4.1 | [ADR-130](ADR-130-homecore-rest-websocket-api.md) WebSocket + appliance health endpoints | +| SEED Detail View | §4.2 | SEED HTTPS API (vector store, witness, sensors, reflex, cognitive analysis) | +| SEED Fleet Map | §4.3 | fleet topology + federation ([ADR-105](ADR-105-federated-csi-training.md)) | +| Entity & State Browser | §4.4 | [ADR-127](ADR-127-homecore-state-machine-rust.md) state machine via [ADR-130](ADR-130-homecore-rest-websocket-api.md) `subscribe_events`; semantic search via [ADR-132](ADR-132-homecore-recorder-history-semantic-search.md) | +| RoomState / Sensing | §4.5 | [ADR-151](ADR-151-room-calibration-specialist-training.md) `GET /api/v1/room/state` | +| COG Management | §4.6 | [ADR-128](ADR-128-homecore-integration-plugin-system.md) plugin runtime + [ADR-100](ADR-100-cog-packaging-specification.md) app registry | +| Calibration Wizard | §4.7 | [ADR-151](ADR-151-room-calibration-specialist-training.md) calibration HTTP API | +| Event Bus & Automation | §4.8 | [ADR-130](ADR-130-homecore-rest-websocket-api.md) broadcast channel + [ADR-129](ADR-129-homecore-automation-engine.md) automation engine | +| Witness / Audit Log | §4.9 | SEED SHA-256 ingest chain + homecore Ed25519 chain | +| Settings & Integration | §4.10 | SEED provisioning, [ADR-116](ADR-116-cog-ha-matter-seed.md) MQTT/Matter, LLAT, federation | + +### 7.1 Build sequencing within the complete deliverable + +The complete UI depends on backing services that mature on their own timelines. Each panel is built against the **real gateway endpoint** defined in §11; where the upstream is not yet available the panel renders a typed empty/error state, **not** fabricated data (the dev-only `?demo=1` fixture of §2.2 exists for offline development only and is never the shipped behaviour). Concretely, the hard contract dependencies are: [ADR-130](ADR-130-homecore-rest-websocket-api.md) (REST + WebSocket), [ADR-127](ADR-127-homecore-state-machine-rust.md) (state machine), [ADR-151](ADR-151-room-calibration-specialist-training.md) (calibration), [ADR-128](ADR-128-homecore-integration-plugin-system.md) (plugin runtime), [ADR-129](ADR-129-homecore-automation-engine.md) (automation), [ADR-132](ADR-132-homecore-recorder-history-semantic-search.md) (event history + semantic search), [ADR-116](ADR-116-cog-ha-matter-seed.md) (SEED/Matter), [ADR-069](ADR-069-cognitum-seed-csi-pipeline.md) (SEED ingest), and [ADR-105](ADR-105-federated-csi-training.md) (federation). The keyword entity filter (§4.4) ships immediately; semantic search layers on once [ADR-132](ADR-132-homecore-recorder-history-semantic-search.md) lands. The exact panel→endpoint→upstream map and the new gateway code each requires are §11; the staged delivery is §12. + +--- + +## 8. Consequences + +### 8.1 Positive + +- Operators, integrators, and residents get a single coherent surface for the full two-tier stack, replacing the need to SSH into SEEDs or hand-craft API calls. +- The dashboard reuses the proven Cognitum design tokens and component patterns verbatim, so it ships visually consistent with no separate design effort and no perceptible seam between surfaces. +- Diagnostic signals that today are invisible (`stale`/`vetoed` flags, kNN fragility, provenance lineage, channel lag) become first-class, surfacing the system's most common real-world failure modes directly to operators. + +### 8.2 Negative / risks + +- The UI hard-depends on the wire-compat guarantees of ADR-130 and the calibration contract of ADR-151; schema drift in either breaks panels silently. Integration tests against every backing contract in §7 are required. +- Committing to the complete UI in one deliverable is a larger up-front effort and couples the UI's readiness to the maturity of multiple backing services (§7.1, §11). The mitigation is the BFF gateway (§2.1): each panel targets one same-origin endpoint, and the gateway absorbs upstream churn behind a stable contract. +- Promoting `homecore-server` to a gateway means it now **proxies cross-tier traffic** (calibration API, SEED HTTPS, appliance daemons). This adds a network hop, a place for upstream timeouts/partial failures to surface, and a server-side store of SEED bearer tokens that must be protected (§11.10). Each proxied route needs an explicit timeout + typed error mapping so one slow SEED cannot stall the dashboard. +- Several panels depend on data that only exists on **real hardware or new daemons** (SEED device tier, appliance host metrics, COG supervisor). Until those upstreams exist the corresponding gateway routes return `503 upstream_unavailable`; this is honest but means the dashboard is only as "live" as the tiers behind it (§11 classifies every endpoint by what it depends on). +- Faithfully mirroring `seed.cognitum.one/store` couples HOMECORE-UI to the external Cog Store's evolving design; token drift there must be tracked and re-synced. +- The two-tier mental model (Appliance root, SEED children, ESP32 leaves) must be enforced consistently; any panel that flattens or peers the tiers undermines the core architectural constraint. + +--- + +## 9. References + +- `https://seed.cognitum.one/store` — primary design reference for all visual conventions. +- `https://seed.cognitum.one/status` — reference for live metric-card layout. +- [ADR-126](ADR-126-ruview-native-ha-port-master.md) — HOMECORE master ADR. +- [ADR-127](ADR-127-homecore-state-machine-rust.md) — HOMECORE-CORE state machine and entity registry. +- [ADR-128](ADR-128-homecore-integration-plugin-system.md) — HOMECORE-PLUGINS WASM COG substrate. +- [ADR-129](ADR-129-homecore-automation-engine.md) — HOMECORE automation engine. +- [ADR-130](ADR-130-homecore-rest-websocket-api.md) — HOMECORE-API REST + WebSocket wire-compat. +- [ADR-132](ADR-132-homecore-recorder-history-semantic-search.md) — homecore-recorder, history + semantic search. +- [ADR-100](ADR-100-cog-packaging-specification.md) — Cognitum Cog packaging specification (manifest.json, status values, on-device layout). +- [ADR-116](ADR-116-cog-ha-matter-seed.md) — cog-ha-matter (SEED cog, HA-DISCO entity surface, mDNS). +- [ADR-069](ADR-069-cognitum-seed-csi-pipeline.md) — ESP32 CSI → Cognitum SEED RVF ingest pipeline (SEED architecture detail). +- [ADR-105](ADR-105-federated-csi-training.md) — Federated CSI training (multi-SEED federation). +- [ADR-151](ADR-151-room-calibration-specialist-training.md) — Per-room calibration specialist training (calibration HTTP API). +- `v2/crates/homecore/src/` — state machine, entity, event, registry source. +- `docs/integration/calibration-appliance-integration.md` — calibration API contract and RoomState schema. + +--- + +## 10. Implementation status + +Implemented as a zero-dependency, no-build-step vanilla TS/JS + CSS frontend served by `homecore-server` at `/homecore` (the `rufield-viewer` "Axum + vanilla-JS" pattern). The complete deliverable per §2/§7 — all ten panels, fully rendered, wired to live data where the backing service exists and to a contract-conformant DEMO-flagged mock layer (§7.1) where it does not. + +**Location:** `v2/crates/homecore-server/ui/` — `css/tokens.css` (the §3.1 palette, verbatim) + `css/app.css` (§3.3 components); `js/{ui,api,ws,mock,app}.js` (shared helpers, REST client, `subscribe_events` WS client, mock layer, shell+router); `js/panels/*.js` (one module per §4 panel). Mounted via `tower-http` `ServeDir` in `homecore-server::build_app`, gated by `--ui-dir`/`HOMECORE_UI_DIR`. + +**Verification:** +- **Rust** — `#[cfg(test)] mod ui_tests` in `homecore-server/src/main.rs`: 5 integration tests (`tower::oneshot`) covering index, design tokens, all ten panel modules served, API coexistence, and mount-disable. *Written but not compiled in the authoring environment (no Rust toolchain present); run `cargo test -p homecore-server` on a Rust host before merge.* +- **Frontend** — `ui/` test suite under plain `node` (no npm install): `npm test` → import/export graph verifier (15 modules) + render-smoke (executes every panel against a DOM shim; 21 checks) + interaction suite (live WS patch, ws.js handshake/parse, calibration contract; 3 checks). **24/24 green.** +- **Benchmark** — `npm run bench`: total bundle **136.8 KB** uncompressed (**~37× smaller** than HA's ~5 MB Lit bundle, the ADR-126 §1.1 foil); slowest panel **1.5 ms/cold-render**. + +**Honest scope — current vs. target.** *Earlier cut:* the front-end was complete but only §4.4 Entities was wired to a real backend; the rest rendered from an in-browser mock. *This revision implements the §11 wiring:* + +- **Front-end (§11.11) — DONE and verified.** `api.js` rewritten: all data accessors are async and call the §11.2 gateway routes; the mock layer is demoted to a dev-only fixture reachable **only** under `?demo=1` / `HOMECORE_UI_DEMO` (§2.2); every panel `await`s and renders a typed empty/error state on failure (no mock fallback in production). All ten panels converted (3 by hand, 7 via parallel agents). Verified under Node: 5 test files green — import graph, boot, render-smoke (22), interaction (3), **and a new prod-errors suite (13) that runs with demo OFF + gateway unreachable and asserts every panel renders an error state, never mock, never throws** (it caught and fixed a real unhandled-rejection in the events panel). +- **Gateway (§11.1–§11.6) — IMPLEMENTED, COMPILED, TESTED, RUN.** New `homecore-server/src/gateway.rs` (+`reqwest` dep, +CLI/env flags `--calibration-url`/`--calibration-token`/`--apps-dir`/`--gateway-timeout-ms`, merged into `build_app` via `gateway_router`). Real handlers: `/api/cal/*` reverse-proxy (W2), `GET /api/homecore/rooms` with the §11.3 RoomState adapter (W2), `GET /api/homecore/cogs` supervisor over the apps dir (W4), `GET /api/homecore/appliance` from `/proc` + port probes (W6). SEED-device/appliance-daemon routes (seeds, federation, witness, privacy, settings, automations, events-history, hailo, tokens — W3/W5) return a typed `503 upstream_unavailable` per §11.2. **Verified on Rust 1.89: `cargo test -p homecore-server --no-default-features` = 12/12 pass** (6 gateway + 6 UI mount). **Run live:** `GET /api/homecore/appliance` returns real `/proc` metrics + TCP service probes; unauth → `401`; `cogs` → `[]` with no apps dir; SEED-tier → typed `503`; and against a mock calibration upstream the `/api/cal/*` proxy passes through (`200`) and `GET /api/homecore/rooms` correctly adapts `RoomState` to the UI shape (`breathing`→`breathing_bpm`, `heartbeat:null`→`heart_bpm:null`, injected `anomaly.threshold`/`room_id`, `stale` passthrough). **Live testing caught + fixed one real bug** — a double-`v1` path in the `/api/cal/*` proxy URL. + +The endpoint-by-endpoint contract is **§11**; the staged plan and which endpoints depend on real SEED/appliance hardware vs. pure software is **§12**. + +--- + +## 11. Backend wiring — making every panel real + +This section is the authoritative contract for full functionality. It removes the mock layer from the production path (§2.2) by routing every panel through the `homecore-server` BFF gateway (§2.1). Each endpoint is classified by what it depends on: + +- **EXISTS** — backend code already in this repo; gateway only proxies/adapts. +- **NEW-GW** — pure software the gateway itself implements (filesystem, `/proc`, process control, recorder query) — no new external service. +- **NEW-API** — a small HTTP wrapper to add to an existing in-repo crate (`homecore-api`, `homecore-automation`). +- **SEED-DEV** — depends on a SEED node's on-device HTTPS API (separate hardware/firmware). +- **APPLIANCE** — depends on an appliance daemon / accelerator stat source. + +### 11.1 Gateway shape + +`homecore-server` already mounts `homecore-api` at `/api/*` and the UI at `/homecore`. It gains a new **`/api/homecore/*`** namespace (the dashboard-specific aggregation surface) plus a **`/api/cal/*`** reverse-proxy to the calibration service. The browser issues only same-origin requests; the gateway fans out server-side, holding all upstream credentials (§11.10). Every proxied route has an explicit timeout and maps upstream failure to a typed body (`503 upstream_unavailable`, `504 upstream_timeout`) so one slow tier never stalls the dashboard. + +### 11.2 Master endpoint contract (panel → gateway route → upstream → status) + +| Panel | UI method (`api.js`) | Gateway route | Upstream / source | Class | +|---|---|---|---|---| +| §4.4 Entities | `states()` | `GET /api/states` | `homecore` state machine | **EXISTS** ✅ wired | +| §4.4/§4.8 live feed | WS | `GET /api/websocket` (`subscribe_events`) | `homecore` event bus | **EXISTS** ✅ wired | +| §4.8 Event history | `eventHistory(q)` | `GET /api/events?since=…` | `homecore-recorder` ([ADR-132](ADR-132-homecore-recorder-history-semantic-search.md)) | **NEW-API** | +| §4.8 Automations | `automations()` / `saveAutomation()` | `GET/POST/DELETE /api/homecore/automations` | `homecore-automation` ([ADR-129](ADR-129-homecore-automation-engine.md)) | **NEW-API** | +| §4.5 Rooms | `roomStates()` | `GET /api/homecore/rooms` → per-room `GET /api/cal/v1/room/state?bank=` | `calibrate-serve` ([ADR-151](ADR-151-room-calibration-specialist-training.md)) | **EXISTS** (proxy + adapter) | +| §4.7 Calibration | `calibration.*` | `POST /api/cal/v1/calibration/{start,stop}`, `GET …/status`, `POST …/enroll/anchor`, `GET …/enroll/status`, `POST …/room/train` | `calibrate-serve` | **EXISTS** (proxy) | +| §4.6 COGs | `cogs()` / `cogAction()` / `cogLogs()` | `GET /api/homecore/cogs`, `POST …/cogs/:id/{start,stop,restart}`, `GET …/cogs/:id/logs`, `GET/PUT …/cogs/:id/config` | COG supervisor over `/var/lib/cognitum/apps/` ([ADR-100](ADR-100-cog-packaging-specification.md)/[ADR-128](ADR-128-homecore-integration-plugin-system.md)) | **NEW-GW** | +| §4.6 Hailo HEF | `hailo()` | `GET /api/homecore/hailo` | `ruvector-hailo-worker:50051` | **APPLIANCE** | +| §4.1 Appliance health | `appliance()` | `GET /api/homecore/appliance` | host `/proc` + Hailo stats + service probes | **NEW-GW** (+APPLIANCE for Hailo) | +| §4.1/§4.2 Fleet + SEED detail | `seeds()` / `seed(id)` | `GET /api/homecore/seeds`, `GET …/seeds/:id` | SEED device HTTPS API ([ADR-069](ADR-069-cognitum-seed-csi-pipeline.md)) via registry | **SEED-DEV** | +| §4.2 SEED actions | `seedCompact()` / `seedVerify()` | `POST …/seeds/:id/{compact,witness/verify}` | SEED device API | **SEED-DEV** | +| §4.3 Federation | `federation()` | `GET /api/homecore/federation` | federation coordinator ([ADR-105](ADR-105-federated-csi-training.md)) | **SEED-DEV/APPLIANCE** | +| §4.9 Witness/Audit | `witnessLog(p,s)` | `GET /api/homecore/witness?page=…` | merge: `homecore` Ed25519 chain + per-SEED SHA-256 chains | **NEW-API + SEED-DEV** | +| §4.9 Privacy mode | `privacyModes()` / `setPrivacy()` | `GET/POST /api/homecore/privacy` | SEED privacy control plane ([ADR-141](ADR-141-bfld-privacy-control-plane-modes-attestation.md)) + cog-ha-matter | **SEED-DEV** | +| §4.9 Export bundle | `exportAttestation()` | `GET /api/homecore/witness/export` | gateway packages both chains | **NEW-GW** | +| §4.10 Tokens (LLAT) | `tokens()` / `createToken()` / `revokeToken()` | `GET/POST/DELETE /api/homecore/tokens` | `homecore-api` `LongLivedTokenStore` | **NEW-API** | +| §4.10 MQTT/Matter | `mqttConfig()` | `GET /api/homecore/integrations/mqtt` | cog-ha-matter config ([ADR-116](ADR-116-cog-ha-matter-seed.md)) | **NEW-GW/SEED-DEV** | +| §4.10 ESP32 provisioning | `nodes()` / `assignRoom()` | `GET/PUT /api/homecore/nodes` | SEED ingest config ([ADR-069](ADR-069-cognitum-seed-csi-pipeline.md)) | **SEED-DEV** | +| §4.10 SEED mgmt | `pairSeed()` / `rotateToken()` | `POST /api/homecore/seeds/{pair,:id/rotate-token}` | SEED pairing (USB `169.254.42.1`) | **SEED-DEV** | + +### 11.3 Calibration proxy + RoomState adapter + +The calibration service is real but on a different binary/port; the gateway reverse-proxies it under `/api/cal/*` (upstream base from `HOMECORE_CALIBRATION_URL`). Its `RoomState` (`wifi-densepose-calibration/src/runtime.rs`) does **not** match the UI's shape, so the gateway adapts it in `GET /api/homecore/rooms`: + +| Real field (`RoomState`) | UI field | Adapter rule | +|---|---|---| +| `breathing: Option` | `breathing_bpm: {value,confidence}\|null` | rename; `value`=`reading.value`, `confidence`=`reading.confidence`; `None`→`null` (preserves "not trained") | +| `heartbeat: Option<…>` | `heart_bpm: {…}\|null` | rename `heartbeat`→`heart_bpm` | +| `presence/posture/restlessness` | same names `{value,confidence}\|null` | `posture.value`=`reading.label` (class), else numeric | +| `anomaly: Option<…>` | `anomaly: {value,confidence,threshold}` | inject `threshold`=`MixtureOfSpecialists.veto_threshold` (0.5) | +| `vetoed` / `stale` | `vetoed` / `stale` | pass through (drives the §4.5/§6 banners) | +| *(absent)* | `room_id`, `seeds[]` | injected by the gateway from the **room registry** | + +A **room registry** (config or derived from `GET /api/cal/v1/calibration/baselines`) maps each `room_id` → bank name + serving SEED ids, so `GET /api/homecore/rooms` returns one adapted record per room. `Option::None` → JSON `null` keeps the null-vs-withheld distinction (§6 invariant 3) intact end-to-end. + +### 11.4 SEED registry & device-API proxy + +The gateway holds a **SEED registry** (`device_id` → base URL + bearer token + zone), populated by pairing (§4.10) and persisted server-side. `GET /api/homecore/seeds[/:id]` fans out to each SEED's on-device API and shapes the result to the §4.2 card/detail model. Expected SEED-side endpoints (the contract the SEED firmware must satisfy — a subset of its 98 endpoints): health; vector-store stats (`vector_count`, `dim`, `epoch`, `knn_latency_ms`, ingest rate); witness (`len`, `last_verify`, `valid`) + `POST verify`; onboard sensors (BME280/PIR/reed/ADS1115/vibration); reflex rules + thresholds; cognitive analysis (fragility, coherence phases); ingest feeders (ESP32 node ids + packet type `0xC5110003`/`0xC5110002` + rate). Offline/unreachable SEEDs surface as `online:false` (drives the §4.1 red tint) rather than failing the whole list. + +### 11.5 Appliance metrics collector (§4.1) + +`GET /api/homecore/appliance`, implemented in the gateway: CPU/RAM/uptime from `/proc`; Hailo load + temperature from the Hailo runtime/sysfs (or `ruvector-hailo-worker` stats); service health by probing `ruview-mcp-brain:9876`, `cognitum-rvf-agent:9004`, `ruvector-hailo-worker:50051`; event-bus rate from the `homecore` broadcast channel + its lag counter (already exposed for §4.1/§4.4). + +### 11.6 COG supervisor (§4.6) + +`GET /api/homecore/cogs`: read each `/var/lib/cognitum/apps/*/manifest.json` ([ADR-100](ADR-100-cog-packaging-specification.md)), the pid file, and verify `binary_sha256` + `binary_signature` (Ed25519) → status/shield. `POST …/cogs/:id/{start,stop,restart}` performs supervised process control; `GET …/cogs/:id/logs` tails `output.log`/`error.log`; `GET/PUT …/cogs/:id/config` reads/writes `config.json`. Hailo-arch COGs join the §11.5 Hailo stats. The Cog Store/App-Registry **browsing** panel was removed per product decision; this is operational management only. + +### 11.7 Witness aggregation + privacy (§4.9) + +`GET /api/homecore/witness` merges two chains chronologically: the `homecore` Ed25519 state-transition chain (exposed by a small `homecore-api` route over its witness log) and each paired SEED's SHA-256 ingest chain (proxied via the registry), paginated server-side. `GET/POST /api/homecore/privacy` reads/sets per-SEED privacy mode via the SEED privacy control plane ([ADR-141](ADR-141-bfld-privacy-control-plane-modes-attestation.md)) — the POST is the high-stakes confirmed toggle (§4.9). `GET /api/homecore/witness/export` packages both chains into the downloadable attestation bundle. + +### 11.8 Event history + automation CRUD (§4.8) + +`homecore-api` adds `GET /api/events?since=…` backed by `homecore-recorder` ([ADR-132](ADR-132-homecore-recorder-history-semantic-search.md)) for history (live updates continue over the existing WS). The automation builder persists through `GET/POST/DELETE /api/homecore/automations`, a thin HTTP wrapper over the `homecore-automation` engine's register/list/remove ([ADR-129](ADR-129-homecore-automation-engine.md)). RuView-specific triggers (RoomState thresholds, SEED reflex events) map onto the engine's trigger types. + +### 11.9 Entity provenance convention (§4.4/§6) + +The first-class provenance badge requires each entity to carry its lineage. Convention: every integration writes `attributes.source` (and, where known, `attributes.seed` / `attributes.cog`) when it sets state; `cog-ha-matter` ([ADR-116](ADR-116-cog-ha-matter-seed.md)) populates these from the ESP32 node → SEED → COG path and HA `via_device`. The gateway/UI resolves node→seed→cog from these attributes (no fabrication; missing lineage renders as "unknown", not invented). + +### 11.10 Auth, credentials, config + +- **Browser → gateway:** one long-lived access token (the §4.10 LLAT), sent as `Authorization: Bearer`; validated by `homecore-api`'s `LongLivedTokenStore`. The dev default (`allow_any_non_empty`) stays for local runs; production provisions `HOMECORE_TOKENS`. +- **Gateway → upstreams:** SEED bearer tokens and the calibration token live **only** server-side (SEED registry + `HOMECORE_CALIBRATION_TOKEN`); never sent to the browser. This is the reason the gateway exists. +- **Config:** `HOMECORE_CALIBRATION_URL`, SEED registry store path, per-proxy timeout (default 2 s), `HOMECORE_UI_DEMO` (dev fixture). No browser CORS needed (same origin); gateway→upstream is server-to-server. + +### 11.11 Front-end changes + +`api.js`: drop the mock fallback from the production path — methods call the §11.2 gateway routes; `this.base` stays same-origin; the mock layer is reachable only under `?demo=1`/`HOMECORE_UI_DEMO`. Every panel renders a **typed empty/error state** (not mock) when its route returns `503/504`. `mock.js` moves to a dev fixture (kept for the offline test harness, excluded from the production bundle). The §10 frontend tests are re-pointed at the gateway contract (and gain contract tests per §11.2 route). + +--- + +## 12. Delivery plan to full functionality + +Staged so each wave is independently shippable behind the gateway, lands real data for a coherent set of panels, and has an explicit acceptance gate. "Class" reuses §11's tags. + +| Wave | Scope | Class | Acceptance gate | +|---|---|---|---| +| **W1 — Gateway foundation** | `/api/homecore/*` scaffold in `homecore-server`; auth passthrough; per-proxy timeout + typed errors; `api.js` base + remove prod mock (`?demo=1` only); panels get typed empty/error states | NEW-GW | Entities + live WS still green; with no upstreams, every other panel shows "upstream unavailable", **never** mock (unless `?demo=1`); Rust + JS suites pass | +| **W2 — Rooms + Calibration** | `/api/cal/*` reverse-proxy; `GET /api/homecore/rooms` with the §11.3 RoomState adapter + room registry; wire §4.5 + the §4.7 wizard to real endpoints; delete the in-browser calibration stub | EXISTS (proxy+adapter) | Against a running `calibrate-serve` (replayed CSI), the wizard drives a real baseline→enroll→train→verify and §4.5 shows real `RoomState` with correct stale/veto/null mapping; contract test on the adapter | +| **W3 — Events + Automations** | `GET /api/events` over `homecore-recorder`; `/api/homecore/automations` over `homecore-automation` | NEW-API | §4.8 history loads from recorder; an automation created in the UI persists and fires via the engine | +| **W4 — COG management** | `/api/homecore/cogs*` supervisor over `/var/lib/cognitum/apps/` (manifest + pid + sig verify + logs + config) | NEW-GW | §4.6 lists real installed COGs; start/stop/restart works; sha256/signature shield reflects real verification; logs tail | +| **W5 — SEED tier** | SEED registry + pairing; `/api/homecore/seeds*` device proxy; witness merge + privacy control; ESP32 provisioning | SEED-DEV | Against a real or emulated SEED API, §4.2/§4.3/§4.9/§4.10 show real vector-store/witness/sensor/reflex/cognition data; SEED tokens stay server-side; offline SEED → red tint, not a failed page | +| **W6 — Appliance + federation + Hailo** | `/api/homecore/appliance` (host metrics + service probes); `/api/homecore/hailo`; `/api/homecore/federation` ([ADR-105](ADR-105-federated-csi-training.md)) | NEW-GW + APPLIANCE | §4.1 health is real; §4.6 Hailo HEF/throughput real; §4.3 federation round/coordinator/Krum real | + +**Definition of done (full functionality):** with W1–W6 merged and the upstream tiers running, loading `/homecore` with **no** `?demo=1` flag shows live data on all ten panels, `api.anyDemo()` is false, and no panel renders fabricated values. Panels whose tier is offline show typed empty/error states. The mock layer is reachable only as the `?demo=1` developer fixture. + +### 12.1 Wave status (this revision) + +| Wave | Status | +|---|---| +| **W1 — Gateway foundation** | ✅ DONE — `gateway.rs`, auth passthrough, typed `503/504`, merged into `build_app`; front-end mock removed from prod path + `?demo=1` fixture; typed error states. **Compiled + 12/12 Rust tests + JS suite green + run live.** | +| **W2 — Rooms + Calibration** | ✅ DONE — `/api/cal/*` reverse-proxy + `GET /api/homecore/rooms` RoomState adapter; front-end calibration stub deleted (now proxies the real API). **Proven live against a calibration upstream** (proxy 200 + adapted shape); null-preservation unit-tested. | +| **W3 — Events + Automations** | ⏳ gateway returns typed `503` (recorder/automation HTTP wrappers pending); front-end handles it gracefully (history note, builder still usable). | +| **W4 — COG management** | ✅ supervisor DONE — lists `/var/lib/cognitum/apps/` manifests + pid liveness (returns `[]` live with no apps dir); start/stop/log/config control is the remaining follow-up. | +| **W5 — SEED tier** | ⏳ gateway returns typed `503` (SEED registry + device proxy pending real/emulated SEED hardware). | +| **W6 — Appliance + federation + Hailo** | ◑ appliance host metrics from `/proc` + port probes DONE (live `/proc` data verified); Hailo stats + federation remain `503` (need the accelerator stat source / coordinator). | + +**Status:** the gateway is **compiled and tested on Rust 1.89** (`cargo test -p homecore-server` = 12/12) and was **run live** (curl proof in §10). The one remaining caveat is intrinsic, not an environment limit: **W3/W5/W6-Hailo/federation depend on services/hardware that are not in this repo** (recorder/automation HTTP wrappers, real SEED nodes, the Hailo stat source), so they return honest typed `503`s and the UI shows error states — exactly as §2.2/§11.2 prescribe. W1/W2/W4/W6-appliance are functional now. + +### 12.2 Security review (PR #1082) + +A high-effort public-PR review of the merged gateway + front-end surfaced the following, all fixed and pinned by tests (`cargo test -p homecore-server` is now **18/18**): + +| # | Severity | Finding | Fix | +|---|---|---|---| +| 1 | **HIGH** | **Path-traversal / confused-deputy SSRF** in the `/api/cal/*` reverse-proxy. The wildcard path was interpolated into the upstream URL while `proxy()` attaches the privileged server-side calibration bearer, so `/api/cal/v1/../../x` (or `..%2f`, `%2e%2e`, leading `/`, `\`, double-encoded `%252e`) could escape the `…/api/` scope **with the token**. | `validate_proxy_path()` decode-then-checks and rejects absolute / backslash / dot-segment / encoded-traversal paths with a typed **400 before the URL is built** (GET **and** POST); legit `v1/...` paths still pass. | +| 2 | Correctness | **CORS + tracing didn't cover gateway routes** — `/api/homecore/*` + `/api/cal/*` were `.merge()`d outside `homecore-api::router()`'s layers. | The audited HC-05 `build_cors_layer()` + `TraceLayer` are now applied to the whole merged app in `main.rs`. | +| 3 | Honesty (§6) | **Fabricated data** — hardcoded `anomaly.threshold: 0.5` in the adapter; dashboard rendered `"null%"`/`"null°C"`; COG Hailo pill hardcoded `"connected"`; `rooms.js` defaulted a null threshold to `0.8`. | Threshold passes through the real upstream value or emits `null` (withheld); dashboard renders `—`; the Hailo pill reflects the real appliance probe; the UI treats a null threshold as withheld. | +| 4 | Robustness | A string `hef` (forwarded verbatim) threw on `.forEach`/`.join`; `frames/target` could be `NaN%`/`Infinity%`; calibration Restart leaked the baseline `setTimeout` poll. | `asArray()` coercion; `target > 0` guard; cancellable poll cleared on Restart / panel teardown. | +| 5 | Perf | Sequential per-bank RoomState fetches; blocking `std::net::TcpStream::connect_timeout` probes on an async handler; `mock.js` statically bundled. | Concurrent `futures::join_all`; async `tokio::net::TcpStream` + `timeout`; demo-only dynamic `import()` of `mock.js`. | + +**Known limitations carried forward (not regressions):** +- **`reqwest` rustls-only is a workspace-wide concern.** `homecore-server` opts into `rustls-tls` only, but cargo feature-unification means any sibling crate enabling the default `native-tls` re-introduces OpenSSL into the final binary. A true "no OpenSSL on the appliance" guarantee requires aligning **every** reqwest-pulling crate on rustls-only — out of scope for this PR; documented at the dependency in `Cargo.toml`. +- **DEV-mode auth.** When `HOMECORE_TOKENS` is unset, the token store falls back to `allow_any_non_empty()` (any non-empty bearer accepted) on `0.0.0.0`. This is pre-existing and intentionally **unchanged** here; the loud boot `warn!` is retained. Provision real tokens (`HOMECORE_TOKENS=…`) before exposing the server to a network. diff --git a/api-docs/adr/ADR-132-homecore-recorder-history-semantic-search.md b/api-docs/adr/ADR-132-homecore-recorder-history-semantic-search.md index 9ec48555..d1b75fe7 100644 --- a/api-docs/adr/ADR-132-homecore-recorder-history-semantic-search.md +++ b/api-docs/adr/ADR-132-homecore-recorder-history-semantic-search.md @@ -120,6 +120,42 @@ tested; P3 is planned. HOMECORE-API (ADR-130, P3); automation conditions on historical state are HOMECORE-automation (ADR-129, P3). +## 3a. Security review (2026-06, post-ADR-154–159 sweep) + +A beyond-SOTA security review of `homecore-recorder` covered SQL injection, retention/purge +correctness, fail-closed write integrity, semantic-store NaN poisoning, and PII exposure. + +**Confirmed clean (with evidence):** + +- **SQL injection — clean.** Every query in `db.rs` uses bound `?` parameters; no user- or + entity-influenceable value is interpolated into SQL via `format!`/concatenation. The only + `format!` builds the `LIKE` *pattern* string, which is itself **bound** as a parameter with + `ESCAPE '\\'` and `% _ \` escaping — so a metacharacter payload is matched literally. Pinned + by `malicious_entity_id_is_stored_literally_not_executed` (a `'; DROP TABLE states; --` state + value leaves the table intact and round-trips verbatim) and + `like_metacharacters_in_query_are_literal_not_wildcards`. +- **NaN-index poisoning — structurally impossible.** Embeddings are SHA-256 → `i32` → + `f32`; an `i32`→`f32` cast is always finite (never NaN/Inf), and an all-zero-digest is + guarded by the `norm > 1e-10` check. Empty-index search, empty-string query, and `k=0` were + probed and all return `Ok(0)` with no panic. (Unlike the calibration/vitals/geo paths, no raw + sensor float ever reaches the index.) +- **Fail-closed writes.** A removal event returns `Ok(None)`; semantic-index failure is logged, + not propagated, so it never blocks the durable SQLite write; `EntityId` parse failure falls + back to a sentinel rather than panicking. + +**Fixed (real bounding bugs):** + +- **Memory-DoS — `get_state_history` was unbounded.** No `LIMIT`, so a wide time window over a + high-frequency entity loaded an unbounded row set into memory. Now capped at + `MAX_HISTORY_ROWS` (1,000,000); sibling search paths were already `k`-bounded. +- **Disk-DoS / documented-but-missing `purge`.** The README advertised `Recorder::purge`, but + no retention path existed → unbounded disk growth. Added a **transactional** `purge(older_than)` + with an **exclusive** cutoff (idempotent, no off-by-one) that deletes old `states`/`events` and + GCs orphaned `state_attributes` blobs (dedup-shared blobs kept until their last referrer is gone). + +`homecore-recorder` tests: 19 → 25 (`--no-default-features`) / 25 → 31 (`--features ruvector`), +0 failed. Python deterministic proof unchanged (recorder is off the signal proof path). + ## 4. Links - Crate: `v2/crates/homecore-recorder/` — `Cargo.toml`, `README.md`, `src/lib.rs`, diff --git a/api-docs/adr/ADR-133-homecore-assist-ruflo.md b/api-docs/adr/ADR-133-homecore-assist-ruflo.md index 6a712e89..fd745de8 100644 --- a/api-docs/adr/ADR-133-homecore-assist-ruflo.md +++ b/api-docs/adr/ADR-133-homecore-assist-ruflo.md @@ -174,3 +174,71 @@ vs. an in-memory array at compile time), which intersects with ADR-084 (RabitQ) | **P1** (this ADR) | `intent`, `recognizer` (regex), `handler` (5 built-ins), `runner` (trait + noop), `pipeline` (end-to-end wiring), 10–15 tests | | **P2** | Real `tokio::process::Child` runner with Windows-safe teardown; `SemanticIntentRecognizer` with ruvector HNSW | | **P3** | STT/TTS bridge, satellite protocol, cloud fallback | + +--- + +## 6. Security review (beyond-SOTA, untrusted-input → action path) + +A focused security review of the Assist pipeline — `utterance → recognizer → +intent → handler → action`, plus `RufloRunner` — treating the utterance as +untrusted input (voice transcripts, the WebSocket `assist` command). This +surface was not covered by the ADR-154–159 sweep. + +### 6.1 Finding fixed — HC-ASSIST-01 (unbounded-utterance DoS, LOW) + +Both `RegexIntentRecognizer::recognize` and the semantic `recognize_scored` +accepted utterances of **unbounded length** and ran `to_lowercase()` (a full +clone) + a per-registered-pattern scan (and, in the semantic path, full +tokenisation + feature-hash embedding) before any bound — an allocation/CPU +amplification on attacker-controlled input. The `regex` crate is **linear-time** +(RE2-style finite automaton, no catastrophic backtracking), so this was a +throughput/memory DoS, not a hang. + +**Fix:** `MAX_UTTERANCE_BYTES = 4096` (far above any real spoken command), +checked at **both** recognizer boundaries *before* any allocation/scan. An +over-length utterance **fails closed** to `Ok(None)` — no intent, no action, +identical to an unrecognised phrase — so it can never be coerced into firing a +handler. Pinned by `over_length_utterance_fails_closed` (an over-length +utterance that *contains* a valid command resolves to `None`, which would have +matched on the old code) and `over_length_utterance_fails_closed_semantic`. + +### 6.2 Dimensions confirmed clean (with evidence) + +- **Command / argument injection — NO SUBPROCESS SURFACE.** The `RufloRunner` + has exactly two impls: `NoopRunner` (no process) and `LocalRunner` (runs the + local recognizer, no process). There is **no** `std::process` / `tokio::process` + / `Command` / process `.spawn()` anywhere in the crate — the trait `spawn` is + only a `started: bool` lifecycle flag — and `RufloRunnerOpts.{script_path,env}` + are **inert data, never consumed**. The live `node ruflo-agent.js` runner is + genuinely data-gated/future (P2). Defence-in-depth: the `entity_id` capture + class `[a-z_][a-z0-9_ .]*` **excludes every shell/SQL metacharacter**, so even + when an injection-shaped utterance resolves (the regex is not exact-anchored), + the captured slot is a clean token — sanitisation by construction. Pins: + `shell_metachars_never_survive_into_a_resolved_slot`, + `runner_opts_are_inert_no_process_spawned`, + `pipeline_injection_shaped_utterance_carries_no_metachars_to_service`. +- **ReDoS — STRUCTURALLY IMPOSSIBLE.** `regex 1.12.3` (no `fancy-regex` in the + dependency tree) is linear-time; a classic `(a+)+$` shape on adversarial input + completes in bounded time. Pin: + `pathological_backtracking_pattern_completes_in_bounded_time`. Patterns are + operator-registered, not user-supplied, in any case. +- **NaN-poisoning — EMBEDDINGS STRUCTURALLY FINITE.** The embedding path takes + only `&str` and produces values via FNV feature-hashing + a guarded L2 + normalise (`norm > 1e-12`); no external float input, no unguarded division, so + a crafted utterance cannot inject NaN/Inf to poison the cosine k-NN. Cosine + against the zero vector is a finite `0.0`; an empty index `max_by` returns + `None` (no panic); the NaN-safe `partial_cmp().unwrap_or(Equal)` is already in + place. Pins: `embeddings_are_structurally_finite`, + `cosine_with_zero_vector_is_finite_not_nan`, + `empty_utterance_against_empty_index_no_panic_no_match`. +- **Intent confusion / fail-closed.** An unrecognised utterance → `not_understood()` + (no service call); a recognised intent with no registered handler → + `not_understood()`; semantic below-threshold / empty-index → regex fallback. + No default high-privilege intent, no fail-open path. +- **Panic-on-input.** No `unwrap`/`expect`/index reachable from a crafted + utterance; the one `exemplars[id]` index uses an `id` from `enumerate()` over + the append-only exemplar `Vec` (no remove API), so it is always in bounds. + +`cargo test -p homecore-assist --no-default-features`: **29→36, 0 failed** (+7); +default/`semantic`: **39→48, 0 failed** (+9). Python deterministic proof +unchanged (homecore-assist is off the signal proof path). diff --git a/api-docs/adr/ADR-137-fusion-engine-quality-scoring-evidence.md b/api-docs/adr/ADR-137-fusion-engine-quality-scoring-evidence.md index fd3c8657..af0028cd 100644 --- a/api-docs/adr/ADR-137-fusion-engine-quality-scoring-evidence.md +++ b/api-docs/adr/ADR-137-fusion-engine-quality-scoring-evidence.md @@ -495,3 +495,34 @@ Rejected. `ViewpointFusionEvent` (viewpoint/fusion.rs lines 183–219) is an int **Integration glue -- not yet on the live path:** emission of `CalibrationIdMismatch` / `DriftProfileConflict` / `PhaseAlignmentFailed` once `calibration_id` propagation and the phase-align convergence signal are threaded onto frames; the BFLD witness record emitted on privacy demotion. **Trust contribution:** sensor *agreement made explicit* -- fusion records the evidence it relied on, and any disagreement automatically tightens the downstream privacy class. + +--- + +## Witness Integrity Review (2026-06-14) — domain-separation fix + +A beyond-SOTA security review of `wifi-densepose-engine` (the composition root +that builds the §2.7 trust witness in `witness_of`) found a real **witness +domain-separation gap**, now fixed. + +**Finding (witness-gap, HIGH).** `witness_of` concatenated `model_version`, +`calibration_version`, and `privacy_decision` boundary-to-boundary, and the +variable-length `evidence` list carried no explicit count. A string straddling a +field boundary therefore collided with a *different* trust decision — +e.g. a per-room adapter id (ADR-150 §3.4, operator-influenceable) that absorbs +the leading bytes of the calibration epoch (`model="…cal:00a"`, `cal="b"`) +produces the **same** witness as `model="…"`, `cal="cal:00ab"`. Two distinct +privacy-relevant input tuples → one witness defeats the "any privacy-relevant +delta → different witness" guarantee this ADR's §2.7 witness exists to provide. + +**Fix.** The witness now (a) prepends a domain tag `ruview.engine.witness.v1`, +(b) writes an explicit 8-byte evidence count, and (c) **length-prefixes every +field** (8-byte LE length ‖ bytes), so field framing is unambiguous regardless +of contents. This is a witness-layout change (all prior witness bytes are +invalidated by design); downstream consumers only assert witness *relationships* +(`assert_ne`/`assert_eq` across runs), not absolute bytes, so nothing breaks. + +Pinned by `witness_distinguishes_model_calibration_boundary` and +`witness_distinguishes_evidence_model_boundary` (both fail on the old +concatenation). Witness **determinism** was reviewed and confirmed clean: no +HashMap iteration and no float formatting feed the hash (floats appear only in +the `SemanticState` statement, which is outside the witness). diff --git a/api-docs/adr/ADR-141-bfld-privacy-control-plane-modes-attestation.md b/api-docs/adr/ADR-141-bfld-privacy-control-plane-modes-attestation.md index 9d3d1e57..40c45cfe 100644 --- a/api-docs/adr/ADR-141-bfld-privacy-control-plane-modes-attestation.md +++ b/api-docs/adr/ADR-141-bfld-privacy-control-plane-modes-attestation.md @@ -599,3 +599,53 @@ Per ADR-028/ADR-010, three rows are added to the witness log: **Integration glue -- not yet on the live path:** wiring the registry into `PrivacyGate` class transitions, the MQTT discovery payload, and a read-only Home Assistant diagnostic entity exposing the active mode + proof hash. **Trust contribution:** the *policy spine* -- privacy posture is a tamper-evident, auditable chain rather than a checkbox; an operator's mode choice actively governs whether identity data may even exist. + +--- + +## Privacy Monotonicity Review (2026-06-14) — confirmed clean + +A beyond-SOTA security review of the governed-trust cycle +(`wifi-densepose-engine::StreamingEngine::process_cycle_calibrated`) examined +the privacy-demotion path this ADR governs. **The monotonicity invariant holds: +demotion only ever makes the emitted class more restrictive, never less.** + +Verification (no behaviour change, the result is a clean bill with evidence): + +- Each cycle computes `effective_class` fresh from the active mode's + `target_class()` (the floor) and applies at most a **single-step** demotion + (`demote_one`, clamped at `Restricted`). There is no cross-cycle state that + could let a permissive class overwrite a restrictive one. +- A forced contradiction (calibration mismatch / array-geometry insufficiency / + mesh partition risk, ADR-032) raises the class byte; a clean cycle emits + exactly the base class. +- Pinned by `forced_contradiction_never_relaxes_class`, a property test over + **all five** `PrivacyMode`s asserting `effective_class.as_u8() >= + base_class.as_u8()` (strictly greater unless already clamped at `Restricted`) + under a forced contradiction, and `== base` on a clean cycle. + +Fail-closed boundaries were also pinned: an empty cycle errors (no degenerate +over-permissive output, `empty_cycle_fails_closed`) and the single-node boundary +is characterized as a valid non-demoting mode (`single_node_cycle_is_well_formed`). + +The related witness domain-separation fix from the same review is recorded in +ADR-137 (the witness folds `effective_class`, so the demotion is auditable). +## Security & Privacy Review (2026-06-14) + +Beyond-SOTA privacy+security review of `wifi-densepose-bfld` (the crate was not in the ADR-154–159 sweep). Two real bugs fixed (each pinned by a fails-on-old test), several dimensions confirmed clean. + +### Findings + +| # | Severity | Site | Issue | Fix | Pinned by | +|---|----------|------|-------|-----|-----------| +| 1 | **privacy-bypass (HIGH)** | `pipeline.rs::process_to_frame` | The documented wire-bytes production path stamped the frame header with the active `PrivacyClass` but serialized the caller's `BfldPayload` **unchanged** via `BfldFrame::from_payload` — never routing through `PrivacyGate::demote`. A frame labeled `Anonymous`(2)/`Restricted`(3) carried the full `compressed_angle_matrix` (identity surface) + amplitude/phase + `csi_delta`. A `NetworkSink` accepts class ≥ `Derived`(1), so the identity surface could cross the node boundary despite the restrictive class byte — the byte lied about content. | Apply `PrivacyGate::demote(frame, active_class)` after construction: a same-class transition that strips the sections the class forbids; `Raw`/`Derived` keep the full payload. | `tests/pipeline_to_frame.rs::process_to_frame_at_anonymous_strips_identity_leaky_sections`, `…_in_privacy_mode_strips_amplitude_and_phase` (both FAILED pre-fix); `…_at_derived_preserves_full_payload` (over-strip guard) | +| 2 | **PII/injection (MEDIUM)** | `mqtt_topics.rs::render_events` | `zone_activity` payload built as `format!("\"{zone}\"")` with no JSON escaping (while `ha_discovery.rs` already escapes). A zone name with `"`/`\` produced malformed/injectable JSON on the HA state topic. | `json_string_literal()` escaper mirroring `ha_discovery::push_str_field`. Value-identical for normal zone names. | `tests/mqtt_topic_routing.rs::zone_payload_escapes_json_metacharacters` (FAILED pre-fix) | + +### Dimensions confirmed clean (with evidence) + +- **Event-field privacy gating** — `BfldEvent::apply_privacy_gating` nulls `identity_risk_score` + `rf_signature_hash` at `Restricted`, and `serde(skip_serializing_if = "Option::is_none")` omits them entirely. `render_events`/`render_discovery_payloads` refuse class < `Anonymous` (stricter than the `sink.rs` `NetworkKind` `MIN_CLASS = Derived` — defense in depth toward less leakage). Covered by `event_privacy_gating.rs`, `mqtt_topic_routing.rs`, `ha_discovery.rs`. +- **Witness/hash framing (the engine `witness_of` bug class)** — CLEAN. `SignatureHasher::compute` prefixes a **fixed 4-byte** `day_epoch` then a **fixed-width canonical-f32** feature block (`IdentityFeatures`: Embedding = `EMBEDDING_DIM*4`, RiskFactors = 16 B). `PrivacyAttestationProof::compute` hashes a fixed 32-byte `prev_hash` + three fixed 1-byte values. No variable-length operator-influenceable string is concatenated into any digest — no length-prefix-framing collision is possible. +- **Fail-closed** — `payload.rs::from_bytes` rejects truncated/overflowing/trailing-byte sections (`checked_add`, bounds checks); `frame.rs::from_bytes` validates magic/version/length/CRC; `PrivacyClass::try_from` rejects unknown bytes; `identity_risk::score` maps NaN/degenerate factors → 0.0 (privacy-conservative). The `from_score(NaN) → Accept` choice is a documented, deliberate publish-aggregate-only fallback (NaN never reaches it from `score()`); risk-driven NaN cannot leak identity because identity gating is class-byte-driven, not risk-driven. + +### Observation (not a bug) + +The ADR-141 control plane (`PrivacyMode`/`PrivacyModeRegistry`) is **not yet wired into the emit path** — the emitter/pipeline enforce the raw `PrivacyClass` directly; the registry is exported + unit-tested but advisory. This matches the "Integration glue — not yet on the live path" status above. The class-byte enforcement (emitter + event + renderers + the now-fixed `process_to_frame`) is the live guarantee. Wiring the registry is the documented next step. diff --git a/api-docs/adr/ADR-151-room-calibration-specialist-training.md b/api-docs/adr/ADR-151-room-calibration-specialist-training.md index 6af4e8cb..ecfd70f8 100644 --- a/api-docs/adr/ADR-151-room-calibration-specialist-training.md +++ b/api-docs/adr/ADR-151-room-calibration-specialist-training.md @@ -253,6 +253,54 @@ Validation per CLAUDE.md: `cargo test --workspace --no-default-features` green; --- +## 6. Review notes + +### 6.1 Correctness + security review (2026-06-14) + +Beyond-SOTA correctness+security review of `wifi-densepose-calibration` (this +ADR's pipeline), un-covered by the ADR-154–159 sweep. + +**Finding (FIXED) — NaN-poisoning of the feature path (numerical / fail-closed).** +`Features::from_series` — the carrier for both live inference and training-anchor +extraction — computed `mean`/`variance`/`motion` over the raw scalar series with +no non-finite guard. A single `NaN`/`±inf` sample (corrupt CSI frame) yielded +`mean=NaN, variance=NaN` and an all-`NaN` prototype embedding. Persisted into a +`PresenceSpecialist::threshold`/`empty_mean` at train time, the `NaN` **silently +disabled presence detection** for the bank's lifetime (every `>` / `|·|` +comparison against `NaN` is false → always reads *absent*, confidence 0), with no +error — and an asymmetry against the rigorously NaN-guarded `geometry_embedding`. +Fixed at the production boundary: non-finite samples are dropped (a corrupt frame +counts as no frame), an all-non-finite series degrades to `Features::ZERO` like +the empty series. Value-identical for all-finite input (full-loop + extract tests +unchanged); pinned by `non_finite_samples_do_not_poison_features` and +`all_non_finite_series_is_zero` (both fail on the old code). + +**Clean dimensions (evidence, no invented issues).** +- *File/path handling:* the crate performs **zero** file/path I/O (no + `std::fs`/`Path`/`File`/`read`/`write` in `src/`; only in-memory `serde_json`). + Path-traversal / unbounded-read / artifact-path handling live entirely in the + `wifi-densepose-cli` consumer (`room.rs`), outside this crate's boundary. +- *Untrusted-load:* `SpecialistBank::from_json` shape-validates via serde + (malformed → `CalibrationError::Serde`); banks are local-first (invariant B), + never network-received. A well-formed bank with adversarial numerics is trusted + as-is — acceptable under the local-first threat model; a validate-on-load + defense-in-depth pass is a possible future hardening, not a present bug. +- *Receipt/hash integrity:* the crate emits no hash/receipt/witness/signature, so + the unframed-concatenation bug class (cf. the engine `witness_of` fix) is + structurally absent. +- *Other numerical paths:* `geometry_embedding` sanitizes every input and sweeps + to finite; presence/restlessness/anomaly divisions are `.max(1e-3)`-guarded; + `autocorr_dominant` guards `r0`, short signals, and empty bands; `train` rejects + empty anchors; anomaly requires ≥2 anchors. + +De-magicked the bare specialist threshold literals (breathing/heartbeat default +min-scores, anomaly outlier-spread multiple + label cutoff) into named documented +consts, value-identical, pinned by const-equality tests. Tests +**58→62 unit + 1 integration, 0 failed**; Python deterministic proof unchanged +(off the signal proof path). + +--- + ## 5. Summary > Big models understand the world. Small ruVector models understand *your room*. diff --git a/api-docs/adr/ADR-154-signal-dsp-beyond-sota.md b/api-docs/adr/ADR-154-signal-dsp-beyond-sota.md index 00466dd9..f779bbff 100644 --- a/api-docs/adr/ADR-154-signal-dsp-beyond-sota.md +++ b/api-docs/adr/ADR-154-signal-dsp-beyond-sota.md @@ -231,6 +231,8 @@ Catalogued so nothing is silently dropped. Priority: **P1** correctness-adjacent > **Horizon-ledger one-liner.** Milestone-0 DONE: dead CIR gate (FIXED+proved), NaN/inf adversarial bypass (FIXED+proved), divide-by-(n−1) window trio (FIXED+proved), calibration dead-branch (FIXED), PSD FFT-planner cache (MEASURED), DTW band (MEASURED). **Milestone-1 DONE (2026-06-13): all four P1 backlog items cleared — circular phase variance #1 (RESOLVED/MEASURED metric, DATA-GATED threshold), Welford n=0 guard #10 (RESOLVED/MEASURED), threshold magic-constants #9 & #13 (RESOLVED-PARTIAL/DATA-GATED — de-magicked + boundary-tested, values unchanged).** **Milestone-2 DONE (2026-06-13): bench-first P2 perf subset + missing boundary tests cleared — spectrogram per-subcarrier FFT re-plan #20 (MEASURED-HOT, 1.40–1.84×, bit-identical); attention/tomography/Kalman #5/#6/#7 (MEASURED-NULL — benched, not hot, left as-is); field_model eigendecompose #8 (MEASUREMENT-ONLY, BLAS un-buildable on this Windows host, number deferred to a BLAS box, NOT fabricated); fft_operator tolerance #14, phase-align convergence-cap #16, csi-ratio epsilon #19 (RESOLVED, tests added).** **Milestone-3 DONE (2026-06-13): the lumped §7.4 row #21–45 P3 backlog cleared, and with it residual P3 items #2/#12/#17/#18 — 22 magic constants de-magicked into named EMPIRICAL-DEFAULT consts (each pinned == prior literal) + 6 boundary/characterization tests across 11 modules; ~4 doc-only; not-real findings (unreachable attractor_drift div0, non-existent gesture thresholds, proof-path features.rs) reported + skipped, no churn; no operating value changed; workspace 3,275/0, Python proof bit-exact `f8e76f21…`.** **§7.4 deferred backlog is now FULLY CLEARED across M0–M3 — nothing silently dropped.** +> **Sibling-crate sweep extension (2026-06-14) — `wifi-densepose-geo` + `wifi-densepose-pointcloud`.** The ADR-154-class numerical-robustness sweep (non-finite-input-poisons-persistent-state + divide-by-zero / asin-domain / degenerate-geometry) was extended to two crates *outside* this ADR's signal scope. **Two real `geo` bugs FIXED, each fails-on-old-pinned:** `terrain.rs::parse_hgt` usize-underflow panic on empty/sub-2x2 SRTM data (`1.0/(side-1)` → panic in debug / inf `cell_size_deg` poisoning `ElevationGrid::get` in release — a truncated download / 404 HTML body reaches it; now `bail!`s when `side < 2`); `coord.rs::haversine` `asin(>1)→NaN` for near-antipodal points (`h` rounds to `1.0+4e-16`; clamped to `[0,1]`). The ±90° pole `cos(lat)=0` ENU singularity is pinned no-panic without changing the transform. **`pointcloud` is confirmed-robust (no manufactured finding):** its only persistent auto-accumulating state (`occupancy` EMA + vitals) is fed solely by the integer-rssi/`sqrt`/`atan2` parser (always finite) and is provably self-healing even under an adversarial NaN/inf `CsiFrame` (`motion_score=(NaN/100).min(1.0)→1.0`; breathing `→0→clamp(5,40)→5.0`) — pinned by `nonfinite_frame_does_not_poison_persistent_state` + degenerate-voxel-fusion no-panic tests. `geo` 9→15 lib / 8 integration; `pointcloud` 18→22; 0 failed; workspace green; Python proof bit-exact `f8e76f21…`. See CHANGELOG `[Unreleased] → Fixed`. + --- ## 8. Consequences diff --git a/api-docs/adr/ADR-161-homecore-server-layer-security.md b/api-docs/adr/ADR-161-homecore-server-layer-security.md index 4404c608..6fea0862 100644 --- a/api-docs/adr/ADR-161-homecore-server-layer-security.md +++ b/api-docs/adr/ADR-161-homecore-server-layer-security.md @@ -265,3 +265,74 @@ Result at time of writing (all 0 failed): perform (B5). - Files kept under the 500-line guideline (`engine.rs` 462; behavioral tests moved to `tests/engine_behaviors.rs`). + +## Addendum — `homecore-api` follow-up security review (beyond-SOTA pass) + +A later network-facing review of `homecore-api` (the remote REST + WS attack +surface) — independent of the ADR-154–159 sweep — found and fixed two real +issues the original M7 pass (which focused on the WS auth bypass HC-WS-01, the +reply-theater HC-WS-02, and the bin token provisioning HC-WS-08) did not catch. +Both are LOW severity and reported at true severity. + +### HC-API-AUTH-01 — `GET /api/` was unauthenticated (FIXED) + +`rest::api_root` took no headers and unconditionally returned +`200 {"message":"API running."}`, while every sibling route gates on +`BearerAuth::from_headers`. HA's `APIStatusView` inherits `requires_auth = True`, +so `/api/` must return **401** for a missing/wrong bearer. HA clients use the +status route as a token-validation probe; a 200 told a bad-token client its +token was valid and let an unauthenticated party confirm a live endpoint. +LOW severity (the body is a static string; no entity/state data leaks). + +**Fix:** `api_root(headers, State)` now validates the bearer like `get_config`. +**Pinned by** (fail-on-old, `tests/server_bin_auth.rs`): +`api_root_rejects_missing_bearer`, `api_root_rejects_wrong_bearer` (both 200→401), +guarded by `api_root_accepts_correct_bearer` (still 200 with a valid token). + +### HC-WS-LAG-01 — `subscribe_events` killed the stream on a broadcast lag (FIXED) + +The per-subscription task matched `Err(_) => break` on both broadcast +`recv()` arms. `RecvError::Lagged(n)` (a slow consumer falling +>`EVENT_CHANNEL_CAPACITY` = 4,096 events behind) is **recoverable** — the bus +doc says "Lagged receivers must re-sync" and HA keeps the subscription alive +across a lag. The old code treated the first lag as fatal, so after an event +burst the client's stream went permanently silent with no error frame — a +self-inflicted event-delivery DoS under load. + +**Fix:** `Lagged(_) => continue` (skip the dropped window, re-sync), +`Closed => break`, on both the system and domain arms of the `select!`. +**Pinned by** `subscription_survives_broadcast_lag` (`tests/ws_handshake.rs`): +subscribes to a filtered event type, floods 6,000 unrelated events past the +4,096 capacity to force a `Lagged`, then asserts a subsequent subscribed event +is still delivered (old code: 5s-timeout panic). + +### Dimensions confirmed clean (with evidence) + +- **AuthN/AuthZ** — all 7 other REST handlers gate on `BearerAuth::from_headers` + → `LongLivedTokenStore::is_valid` before any work; the WS handshake validates + the `auth` token against the same store before the command loop, and + privileged commands are unreachable pre-`auth_ok`. Token compare is + `HashSet::contains` (content-independent timing — not the byte-`==` oracle of + ADR-157 §B4), so no timing-oracle finding. No route skips the gate; no + result-ignored check; no default/empty token accepted. +- **Path traversal** — no route maps user input to a filesystem path (state is an + in-memory `DashMap`); `:entity_id` passes through `EntityId::parse`, a strict + `[a-z0-9_]+\.[a-z0-9_]+` ASCII allowlist that rejects `..`, `/`, `\`, and + absolute paths. No traversal surface. +- **Injection** — no SQL, no shell/subprocess, no `format!`-into-response; + service/state bodies are typed `serde_json::Value` handed to the in-process + registry (HA-equivalent). +- **Info-leak** — `ApiError` maps to fixed status + a typed `{message}`; + `ServiceError::HandlerFailed(String)` is integration-controlled (HA surfaces + the handler error too), never framework internals/paths/stack-traces — no + ADR-080-class leak. +- **CORS** — explicit allowlist with `allow_credentials(false)` (HC-05), + not `permissive()`. +- **De-magic** — no bare security-relevant literals in the crate worth + extracting (`EVENT_CHANNEL_CAPACITY` is already named in `homecore`; CORS + dev-default ports are documented). + +**Tests:** `homecore-api --no-default-features` **25 → 29** (+2 api-root auth, ++1 api-root accept-guard, +1 WS lag-survival), 0 failed. Workspace green. +Python deterministic proof unchanged (homecore-api is off the signal proof +path). diff --git a/api-docs/adr/ADR-165-homecore-migrate-from-home-assistant.md b/api-docs/adr/ADR-165-homecore-migrate-from-home-assistant.md index 9c841764..4f68b231 100644 --- a/api-docs/adr/ADR-165-homecore-migrate-from-home-assistant.md +++ b/api-docs/adr/ADR-165-homecore-migrate-from-home-assistant.md @@ -78,6 +78,23 @@ converts the entity registry; full conversion of the remaining artifacts is defe - `MigrateError` carries context (`path`, line/field) for I/O, JSON, YAML, missing-field, unsupported-schema-version, and entity-id parse failures (`src/lib.rs`). +- **Secret-leak hardening (security review, 2026-06).** `secrets.yaml` parse failures must + NOT use the generic `MigrateError::YamlParse { source }` variant: `serde_yaml`'s message + for a typed-tag coercion error (e.g. `port: !!int `) embeds the offending scalar + verbatim (`invalid value: string ""`), and that error propagates through + the `InspectSecrets` CLI path to stderr — leaking a secret value despite the CLI's + deliberate `` design. `read_secrets` now maps such failures to a dedicated + redacting variant `MigrateError::SecretsParse { path, line, column }` that carries only the + file path and a coarse location (`serde_yaml::Error::location()`), never the scalar content. + Pinned by `secrets::tests::malformed_secrets_error_never_contains_secret_value` (asserts the + rendered error **and its full `#[source]` chain** never contain the secret value). + **Review dimensions confirmed clean with evidence:** source is never mutated (no + `fs::write`/`remove`/`create` anywhere — P1 reads source, writes nothing); paths are + user-supplied dirs joined with fixed filenames (no `..`/absolute traversal beyond the + user's own privileges); malformed/typed/truncated `.storage` JSON and YAML **error, never + panic** (every production `unwrap`/`expect` is test-only); unknown schema `minor_version` + hard-errors fail-closed; no SQL/shell/path injection surface (the tool emits diagnostics + only, persists nothing in P1). ### 2.5 Deferred to P2+ (NOT built — honestly labelled) @@ -89,7 +106,9 @@ converts the entity registry; full conversion of the remaining artifacts is defe ### 2.6 Test evidence (as shipped) -- 19 tests (`cargo test -p homecore-migrate`), per the crate README badge. +- 21 tests (`cargo test -p homecore-migrate`) — 19 as originally shipped plus 2 added by the + 2026-06 security review (`secrets::tests::malformed_secrets_error_never_contains_secret_value`, + `malformed_secrets_error_reports_location`). ## 3. Consequences diff --git a/api-docs/adr/ADR-172-cli-core-csi-deserialiser-security-review.md b/api-docs/adr/ADR-172-cli-core-csi-deserialiser-security-review.md new file mode 100644 index 00000000..001a4f07 --- /dev/null +++ b/api-docs/adr/ADR-172-cli-core-csi-deserialiser-security-review.md @@ -0,0 +1,117 @@ +# ADR-172: `wifi-densepose-cli` + `wifi-densepose-core` CSI-Deserialiser Security Review + +| Field | Value | +|-------|-------| +| **Status** | Accepted — clean-with-evidence, 4 regression pins added | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **CSI-DESERIALISER-HARDENING** | +| **Supersedes / amends** | none (records review; references ADR-127 §9 for the `core` portion, ADR-136 for the pre-existing DoS ACs) | + +## Context + +The beyond-SOTA security sweep (branch `feat/v2-beyond-sota-sweep`) reviewed each +`v2/` crate for real, reproducible defects. Two crates had no prior dedicated +security ADR: + +- **`wifi-densepose-core`** — the dependency root for all 12 downstream crates + (types, traits, error types, CSI frame primitives). A defect here is a + force-multiplier: every consumer inherits it. +- **`wifi-densepose-cli`** — the user-facing entrypoint + (`calibrate`/`calibrate-serve`/`enroll`/`train-room`/`room-watch` + MAT-gated), + which parses untrusted UDP CSI packets and operator-supplied paths. + +A **specific hypothesis** motivated the core review. Three earlier reviews in +this campaign found a systemic **NaN-state-poisoning bug class** in crates that +depend on core (`wifi-densepose-calibration`, `-vitals`, `-geo`): a non-finite +(NaN/Inf) input latched into persistent filter/accumulator state (IIR `y1/y2`, +running mean, Welford/von-Mises accumulator, voxel grid) → silent **permanent** +feature failure. The load-bearing question for this review: **does that bug class +originate in a shared `wifi-densepose-core` primitive** (making the right fix a +single root fix), or was it independently re-implemented in each downstream +crate (making the three existing local fixes complete)? + +## Decision + +Record the review outcome and lock in the existing DoS guards with regression +tests. **No production code is changed** — both crates were already hardened +(ADR-136 acceptance criteria + `sanitize_room_id`); the gap was *untested* +guards, which a future refactor could silently remove. + +### Load-bearing question — VERDICT: **NO** (the NaN class does not live in core) + +`wifi-densepose-core` exposes **no stateful accumulator of any kind** — no +Welford/running-mean, no von-Mises/circular-mean, no IIR/biquad filter state, no +voxel grid. + +- **MEASURED:** `grep` over `core/src` for + `welford|von_mises|biquad|y1|y2|running_mean|accumulat|voxel|self.*+=` matched + only the `InvalidState` *error* enum variant, "reset state" doc comments, and a + test-only LCG — **zero** stateful logic. The only float math in core is + construction-time projection (`CsiFrame::new` → amplitude/phase via `mapv`) and + pure stateless `utils` functions; nothing persists across frames. +- **Corroboration:** `wifi-densepose-calibration::Features::from_series` + (`extract.rs:103–133`) already filters non-finite samples → `Features::ZERO`. + The downstream fixes are independently re-implemented, confirming each crate + rolls its own accumulator and each local fix is correct and complete. **A fix + in core would be a no-op (there is nothing to fix).** + +Consequence: the NaN-state-poisoning class is a *downstream-local* pattern, not a +core-rooted defect. No hidden fourth instance exists in the shared primitive. + +### Findings (all pins — guards already present, now tested) + +| # | Location | Guard (pre-existing) | Regression pin | Evidence (MEASURED) | +|---|----------|----------------------|----------------|---------------------| +| 1 | `core` `types.rs:801` `from_canonical_bytes` | `saturating_mul` shape-vs-length check before `Vec::with_capacity(rows*cols)` | `canonical_decode_oversized_shape_is_bounded_not_allocated` | With guard removed: **panics `capacity overflow` at `types.rs:801`**; with guard: passes | +| 2 | `core` `types.rs` decoder | typed `CanonicalDecodeError`, never panics | `canonical_decode_never_panics_on_arbitrary_bytes` (fuzz sweep) | panic-free on arbitrary bytes | +| 3 | `cli` `calibrate.rs:276–291` | length check `buf.len() < 20 + n_pairs*2` before `Array2::zeros(n_antennas*n_subcarriers)` | `test_parse_csi_packet_oversized_claim_is_rejected_not_allocated` | 255×65535 claim in a 2 KB packet → `None` (no allocation) | +| 4 | `cli` `calibrate.rs` parser | `None`-returning on malformed input | `test_parse_csi_packet_never_panics_on_arbitrary_bytes` (fuzz sweep) | panic-free on arbitrary UDP bytes | + +### Dimensions confirmed clean (with evidence) + +1. **Panic-on-adversarial-input = 0** — `from_canonical_bytes` returns a typed + error for every malformed class; `parse_csi_packet` returns `None`. Both + fuzz-swept panic-free. +2. **NaN handling** — `Confidence::new` rejects NaN + (`!(0.0..=1.0).contains(&NaN)` ⇒ `Err`); `compute_bounding_box` / + `to_flat_array` are NaN-tolerant (f32 min/max ignore NaN). +3. **Empty-frame safety** — `amplitude_variance` / `mean_amplitude` are + panic-free on an empty `Array2` (ndarray 0.17 returns finite / `None`). +4. **Unbounded-memory DoS** — bounded in both deserialisers (findings 1 & 3). +5. **Path traversal** — `calibrate-serve` defends every client-supplied + `room_id`/`bank`/`baseline` via `sanitize_room_id` (`[A-Za-z0-9_-]`, 64-char + cap) with existing tests; bearer-auth gate + non-loopback-bind warning present. + `mat export` writes to an operator-supplied `PathBuf` (acceptable CLI behavior). +6. **Secrets** — `--token` is read from `CALIBRATE_TOKEN` env, never embedded. + +## Validation + +- `cargo test -p wifi-densepose-core` → **35 → 37** lib passed, 0 failed (+3 doctests) +- `cargo test -p wifi-densepose-cli --no-default-features` → **24 → 26** passed, 0 failed +- `cargo test --workspace --no-default-features` → **exit 0**, 0 failed +- `python archive/v1/data/proof/verify.py` → **VERDICT: PASS**, hash + `f8e76f21a0f9852b70b6d9dd5318239f6b20cbcb4cdd995863263cecdc446f7a` **unchanged** + (core/cli are off the signal proof path — confirms no pipeline alteration) + +## Consequences + +### Positive +- Two CSI deserialisers (the untrusted-input boundary of both the library root + and the network-facing CLI) now have their DoS guards pinned against + regression — a future refactor that drops a length check fails CI. +- The NaN-state-poisoning class is settled as downstream-local; reviewers no + longer need to suspect a shared-root defect, and the three prior local fixes + are confirmed complete. + +### Negative +- None. Test-only change; no behavior or API change. + +### Neutral +- The `core` portion is also noted in ADR-127 §9 (shared security-review log); + this ADR is the canonical record for the `wifi-densepose-cli` review. + +## Links +- ADR-127 — HOMECORE state machine (shared security-review log, §9) +- ADR-136 — pre-existing CSI deserialiser DoS acceptance criteria +- ADR-151 — per-room calibration (`calibrate`/`calibrate-serve` surfaces) diff --git a/api-docs/adr/ADR-173-metric-locked-pck-mpjpe-accuracy-harness.md b/api-docs/adr/ADR-173-metric-locked-pck-mpjpe-accuracy-harness.md new file mode 100644 index 00000000..8de81dd1 --- /dev/null +++ b/api-docs/adr/ADR-173-metric-locked-pck-mpjpe-accuracy-harness.md @@ -0,0 +1,123 @@ +# ADR-173: Metric-Locked PCK/MPJPE Accuracy Harness + +| Field | Value | +|-------|-------| +| **Status** | Accepted — implemented, deterministically tested | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **METRIC-LOCK** | +| **Amends** | ADR-155 (generalizes the torso-only `metrics_core::pck_canonical` to a selectable normalization) | +| **Motivated by** | `docs/research/sota-nn-train-benchmark-brief.md` (PR #1090) | + +## Context + +The beyond-SOTA SOTA-research brief (PR #1090) identified the single biggest +threat to any "beyond-SOTA" accuracy claim this project makes: **metric +ambiguity**. Three PCK@20 numbers circulate, computed under three *different and +unstated* normalizations, so they cannot be compared: + +- **96.09–96.61%** — WiFlow-STD reproduction, **image/bounding-box-normalized** PCK (the looser convention). +- **81.63%** — an internal MM-Fi number reported as **"torso-PCK"** (tighter). +- **61.1%** — GraphPose-Fi (arXiv 2511.19105), **standard torso-diameter** PCK on the MM-Fi random split (the academic frontier). + +The project has been burned by this twice: a previously-published 92.9% was +retracted because it used **absolute-pixel** normalization, not torso. Until +there is *one canonical, documented, tested* PCK definition — and every reported +number carries the definition it was computed under — no accuracy comparison is +credible, and the "prove everything" bar cannot be met for the benchmark half of +the work. + +This is measurement infrastructure, not an accuracy claim. The deliverable's job +is to make the metric **unambiguous and reproducible**, so future numbers are +comparable and an unlabeled PCK is structurally impossible. + +## Decision + +Add a metric-locked accuracy harness as a new module +`v2/crates/wifi-densepose-train/src/accuracy.rs` (404 non-test lines; inline +deterministic tests bring the file to 708), re-exported at the crate root. It +**extends, not duplicates** — it reuses `metrics_core`'s geometric primitives +(`bounding_box_diagonal`, canonical hip indices `CANON_LEFT_HIP/RIGHT_HIP`), so +there remains exactly one implementation of each geometric reference; the +existing ADR-155 `pck_canonical` (torso-only) is unchanged and this generalizes +it. + +### Public API + +- `enum PckNormalization { TorsoDiameter, BoundingBoxDiagonal, AbsolutePixels(f32) }` + — the three conventions the three historical numbers used, now **explicit and + selectable**. `.label()` / `.tolerance(...)`. +- `pck_at(pred, gt, vis, k, norm) -> (correct, total, pck)` — PCK@k = + fraction of *visible* keypoints whose predicted-vs-GT distance ≤ the tolerance, + where tolerance = `k%` of the chosen normalizer (or an absolute threshold for + `AbsolutePixels`). +- `mpjpe(pred, gt, vis) -> f32` — mean per-joint position error (2D/3D, coordinate + units; mm for mm inputs). Re-exported crate-root as `pck_mpjpe` to avoid + colliding with the existing `eval::mpjpe`. +- `struct PoseAccuracy { pck_at: BTreeMap, mpjpe, normalization, n_keypoints, n_frames }` + — **a reported number always carries its `normalization`**; an unlabeled PCK is + structurally impossible to produce through this surface. +- `struct PoseFrame { pred, gt, visibility }` + `accuracy_report(frames, ks, norm) -> PoseAccuracy` + (micro-averaged over keypoints). + +### Correctness is proven by hand-computed deterministic tests (no GPU, no data) + +The tests construct synthetic keypoint sets whose PCK/MPJPE can be computed by +hand, and assert the harness matches. Highlights (all pass): + +| Test | Construction | Expected | +|------|--------------|----------| +| perfect_prediction | pred==gt | PCK=1.0 (all 3 norms), MPJPE=0 | +| all_just_outside | every error just past τ@20 | PCK=0.0 | +| half_in_half_out | 2 exact, 2 just outside | PCK=0.5 | +| **three_normalizations (KEY PROOF)** | identical pred; nose err .06, shoulder .10, hips exact | torso=**0.50**, bbox=**1.00**, abs(.08)=**0.75** | +| mpjpe_2d / mpjpe_3d | (3,4)→5 / (1,2,2)→3 | 2.5 / 3.0 | +| mpjpe_excludes_invisible | invisible joint err 100 ignored | 5.0 | +| zero_torso_unscoreable | coincident hips | `(0,0,0.0)`, **not** false-perfect | +| no_visible_keypoints | vis=∅ | `(0,0,0.0)` | +| nan_coords | one NaN pred coord | counted wrong, **no panic** | +| empty report | no frames | 0.0, **not** NaN | +| bbox≥torso ordering | same frames | bbox-PCK ≥ torso-PCK | + +### The key proof (the ambiguity is real and quantified) + +Identical predictions, three declared normalizations → **0.50 / 1.00 / 0.75**. +Mechanism: the bbox diagonal `√(0.20² + 0.80²) = 0.825` is ~4× the hip-span torso +`0.20`, so τ@20 is 0.165 (bbox) vs 0.040 (torso) — the looser image-normalized +convention passes joints the strict torso convention rejects. This is *exactly* +why 96% / 81.6% / 61% cannot be lined up without declaring the enum, demonstrated +in-code. + +## Validation + +- `cargo test -p wifi-densepose-train --no-default-features` → lib **191 → 206** + (+15), `test_metrics` **12 → 14** (+2), doc-tests 8 — **0 failed**. +- `cargo test --workspace --no-default-features` → **exit 0**, 0 failed. +- `python archive/v1/data/proof/verify.py` → **VERDICT: PASS**, hash + `f8e76f21a0f9852b70b6d9dd5318239f6b20cbcb4cdd995863263cecdc446f7a` **unchanged** + (off the signal proof path — confirms no pipeline alteration). + +## Consequences + +### Positive +- The three historical PCK numbers can now be **recomputed under one declared + definition** and compared honestly. The retracted-number class of error + (silent normalization mismatch) is structurally prevented going forward. +- Establishes the measurement substrate for the beyond-SOTA target: GraphPose-Fi + cross-environment **PCK@20 = 12.9%** (standard torso PCK) is now a number this + harness can produce comparably. + +### Negative +- None functional. The harness is additive; no existing metric path changed. + +### Neutral +- Producing actual model numbers under this harness requires the trained models + + datasets (MM-Fi) and, for cross-domain splits, is the next sub-deliverable of + the benchmark/optimization milestone — out of scope here (this ADR is the + *instrument*, not the *reading*). + +## Links +- ADR-155 — metric core (`pck_canonical`, torso-only) — generalized here +- ADR-152 — WiFi-Pose SOTA 2026 intake / WiFlow-STD benchmark +- `docs/research/sota-nn-train-benchmark-brief.md` — the motivating gap analysis +- GraphPose-Fi — arXiv 2511.19105 (verified cross-env PCK@20 = 12.9% anchor) diff --git a/api-docs/adr/ADR-174-ci-bench-regression-compile-verify-gate.md b/api-docs/adr/ADR-174-ci-bench-regression-compile-verify-gate.md new file mode 100644 index 00000000..7aed7fe9 --- /dev/null +++ b/api-docs/adr/ADR-174-ci-bench-regression-compile-verify-gate.md @@ -0,0 +1,110 @@ +# ADR-174: CI Bench-Regression Gate (Compile-Verify) + +| Field | Value | +|-------|-------| +| **Status** | Accepted — implemented, caught one real bit-rotted bench | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **BENCH-GATE** | +| **Milestone** | benchmark/optimization re-balance — sub-deliverable 8.3 | +| **Motivated by** | `docs/research/sota-nn-train-benchmark-brief.md` (target 3: criterion benches as CI regression baselines) | + +## Context + +The v2/ workspace ships **26 criterion benches across 18 crates** (e.g. +`nvsim/pipeline_throughput`, `wifi-densepose-ruvector/{ann,sketch,fusion}_bench`, +`wifi-densepose-signal/{signal,dsp_perf,features,calibration,cir,…}_bench`, +`wifi-densepose-mat/detection_bench`, `wifi-densepose-nn/{inference,native_conv}_bench`, +`wifi-densepose-engine/engine_cycle`, …). Because **benches are not part of +`cargo test`**, nothing in CI compiled them — so they bit-rot silently the moment +a public API they call changes, and the rot is invisible until someone manually +runs `cargo bench` months later. + +The SOTA brief named "wire existing criterion benches into CI as regression +baselines" as a concrete benchmark-hygiene target. The honest difficulty: true +*timing*-regression gating on shared GitHub runners is unreliable — wall-clock +varies 2–3× run-to-run (a captured 10-sample run showed `float_l2/512` ranging +307–444 ns), so a hard threshold or a cross-runner `criterion --baseline` compare +(baseline and PR land on different physical machines) would manufacture false +regressions. A gate that cries wolf gets disabled. + +## Decision + +Add `.github/workflows/bench-regression.yml` with **two jobs of explicitly +different authority** — and do NOT pretend to gate on timing. + +### `bench-compile` — HARD GATE (real regression detection) +`cargo bench --workspace --no-default-features --no-run` compiles + links every +default-feature bench (no measurement → fully deterministic), plus a +`--features cir` compile of the gated `cir_bench`. Benches aren't in `cargo test`, +so this is the genuine guard: **the build fails the moment a bench stops +compiling.** + +### `bench-fast-run` — INFORMATIONAL (`continue-on-error: true`, never gates) +Runs a curated pure-CPU subset (`nvsim/pipeline_throughput`, +`ruvector/{sketch,fusion}_bench`) in criterion quick-mode (1 s warm-up / 2 s +measure / 10 samples), targeted per-`--bench`, and uploads logs as an artifact. +Every number it produces is **informational only** — explicitly stated in the +workflow header. + +### What is NOT done, and why (honest scope) +No timing-regression gate, no committed baseline JSON. The workflow header +documents the exact condition under which true timing-gating becomes honest: a +frequency-pinned **self-hosted** runner with a generous (>2×) floor. A +cross-runner baseline would be dishonest, so none is committed. + +### Proof it matters (MEASURED) +Running the new gate on the current tree immediately caught +`wifi-densepose-mat/detection_bench` failing to compile: +`error[E0063]: missing field last_rssi in initializer of SensorPosition` — the +struct gained a field; the bench was never updated. **Fixed** in the same change +(`last_rssi: None`, the simulated-zone convention) and re-verified +(`cargo bench -p wifi-densepose-mat --no-default-features --bench detection_bench --no-run` +→ `Finished`). The gate paid for itself on its first run. + +### Exclusions (documented in-workflow) +- `ruvector/crv_bench` — its crates.io dep `ruvector-crv 0.1.1` fails to build on + stable (upstream `E0308` in `stage_iii.rs`); excluded with a re-add condition. +- `onnx_bench` / `mqtt_throughput` — feature-gated (ort / mqtt), left to their + crates' own workflows. `wasm-edge/process_frame_bench` — workspace-excluded. + +Conventions mirror existing workflows: `submodules: recursive` (the workspace +path-deps `vendor/rufield`), Swatinem/rust-cache `workspaces: v2`, Tauri/GTK apt +deps (a `--workspace` bench link pulls the whole graph), path-filtered triggers. + +## Validation + +- **Bit-rot caught + fixed** (above), re-verified `--no-run`. +- **MEASURED locally** (`--no-default-features`, Windows): nvsim, ruvector + (sketch/fusion/ann), signal/cir_bench, mat/detection_bench (post-fix), + vitals, ruview-swarm/swarm_bench all compile; fast subset runs (`nvsim + pipeline_run/d1/256` ≈ 55 µs; `ruvector sketch_hamming` ≈ 3–7 ns vs `float_l2` + ≈ 63–371 ns). +- `cargo test -p wifi-densepose-mat --no-default-features` → 166/6/2 passed, 0 failed. +- `python archive/v1/data/proof/verify.py` → **VERDICT: PASS**, hash + `f8e76f21…46f7a` unchanged. +- **Honest limitation:** the full `--workspace --no-run` could not be + end-to-end validated on this Windows box (`desktop` needs GTK, `candle-core` + fails on MSVC, `swarm_bench` LTO-links OOM under parallel pressure — all + Windows-env artifacts; each affected bench compiles standalone here). **The + first green Linux CI run on the PR is the authoritative proof of the + `--workspace` step.** + +## Consequences + +### Positive +- Bench bit-rot is now a hard CI failure, not a silent surprise — the 26 benches + stay compilable as the APIs they exercise evolve. +- The benchmark-infrastructure half of the DoD (step 5) is satisfied honestly, + setting up the next sub-deliverable (QAT-int8 measurement) to be + regression-protected. + +### Negative / Neutral +- No automated timing-regression detection (deliberate — see scope). Revisit only + with a frequency-pinned self-hosted runner. +- One bench (`crv_bench`) excluded pending an upstream dep fix. + +## Links +- ADR-173 — metric-locked accuracy harness (sub-deliverable 8.1) +- `docs/research/sota-nn-train-benchmark-brief.md` — motivating target +- ADR-134 (CIR), ADR-135 (calibration), ADR-154 (signal DSP benches) — benched paths diff --git a/api-docs/adr/ADR-175-int8-quantization-half-pose-model-measured.md b/api-docs/adr/ADR-175-int8-quantization-half-pose-model-measured.md new file mode 100644 index 00000000..cc63e33a --- /dev/null +++ b/api-docs/adr/ADR-175-int8-quantization-half-pose-model-measured.md @@ -0,0 +1,172 @@ +# ADR-175: int8 Quantization of the WiFlow-STD "half" Pose Model — MEASURED accuracy/size trade-off + +| Field | Value | +|-------|-------| +| **Status** | Accepted — MEASURED, reproducible (honest negative) | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **EDGE-INT8** | +| **Sub-deliverable** | 8.2 of the benchmark/optimization milestone | +| **Metric lock** | ADR-173 (one declared PCK normalization for every reported number) | +| **Motivated by** | `docs/research/sota-nn-train-benchmark-brief.md` (§edge int8) | + +## Context + +The SOTA brief characterized the int8 edge story for the WiFlow-STD pose net as +"fully characterized" for PTQ on the **published 2.23M** model (static QDQ +conv-only = the sweet spot; dynamic int8 ≈ no-op on this all-conv net), and named +**QAT-int8 on the strictly-dominating 843,834-param "half" model** as "the one +untested edge lever." This ADR is the reading of that lever — a MEASURED +fp32-vs-int8 trade-off for the half model, not a claim. + +The half model (`half_best.pth`, 843,834 params) is the efficiency-sweep winner +from ADR-152 (`run_sweep.py` VARIANTS[0]: `tcn=[270,220,170,120]`, +`conv=[4,8,16,32]`, `attn_groups=4`). Its fp32 accuracy was recorded in the sweep; +this ADR re-measures it under the locked normalization and quantizes it. + +**The whole point of this deliverable is reproducibility.** Every number below was +produced by running `v2/crates/wifi-densepose-train/scripts/quantize_half_int8.py` +on host `ruvultra` (RTX 5080, torch 2.11.0+cu128) against the real checkpoint and +the real seed-42 test split. The script + the exact command + the recorded stdout +**is** the proof artifact. Nothing here is estimated. + +## Decision + +Quantize the half model to int8 with **both** levers and report both honestly: + +1. **QAT (primary target)** — FX graph-mode quantization-aware training, fbgemm + backend, 3 epochs of fake-quant fine-tuning from `half_best.pth` (AdamW lr 2e-5, + the existing `PoseLoss`), then `convert_fx` to a true int8 graph. +2. **PTQ static QDQ (the brief's "sweet spot", measured as the honest fallback)** — + FX graph-mode static PTQ, fbgemm, calibrated on 64 train batches. + +### Locked normalization (ADR-173) + +**Torso-diameter PCK** — neck (keypoint idx 2) → pelvis (idx 12) distance — the +standard MM-Fi/GraphPose-Fi convention. This is exactly the default +`use_torso_norm=True` path of the upstream harness's `utils/metrics.calculate_pck`. +The **same** `calculate_pck`/`calculate_mpjpe` that produced the sweep's fp32 +numbers scores **both** fp32 and int8 here, so the comparison is metric-locked: no +normalization is mixed, and the fp32 baseline reproduces the sweep's recorded +`half` test numbers bit-for-bit (PCK@20 clean = 96.62%), confirming the harness is +the same one. + +### Device note (why int8 is CPU) + +PyTorch int8 quantized kernels execute on CPU (fbgemm/x86), not CUDA. So int8 eval +is CPU. To keep the accuracy delta device-matched (not confounding int8-vs-fp32 +with CPU-vs-GPU), the script measures an **fp32-CPU** baseline too. fp32-CPU and +fp32-GPU agree to 4 decimals (PCK@20 clean 0.96623 vs 0.96623), so CPU/GPU +introduces no drift — the int8 deltas below are pure quantization effect. + +## MEASURED results (clean test subset = 52,560 NaN-free windows; torso-PCK) + +Source: stdout of the run below + `~/wiflow-std-bench/sweep/int8/int8_results.json`. + +| model | quant | size (MB) | PCK@20 | PCK@50 | MPJPE | Δ PCK@20 | Δ PCK@50 | size win | +|-------|-------|-----------|--------|--------|-------|----------|----------|----------| +| **fp32** (cpu) | — | **3.351** | **96.62%** | **99.47%** | **0.008981** | — | — | 1.00× | +| int8 PTQ static | PTQ | 1.046 | 40.98% | 94.98% | 0.038262 | **−55.64 pp** | −4.49 pp | 3.20× smaller | +| int8 QAT (3 ep) | **QAT** | 1.043 | 67.48% | 98.69% | 0.026548 | **−29.15 pp** | −0.78 pp | 3.21× smaller | + +Full-test-set (54,000 windows incl. NaN-zero-filled files 487–499) tracks the +clean subset: fp32 96.10% / int8-PTQ 41.11% / int8-QAT 67.48% PCK@20 — same shape, +recorded in the JSON. + +### Verdict + +**int8 is NOT a win for this model at the tight PCK@20 edge target — honest no.** + +- **PTQ static collapses** (−55.64 pp PCK@20). Naive static QDQ destroys the half + model. The "sweet spot" characterization from the brief does not transfer from + the 2.23M model to this 843k model at the strict torso-PCK@20 threshold. +- **QAT recovers a large share of the relative gap** (PTQ 40.98% → QAT 67.48%) but + still **loses 29.15 pp** at PCK@20 for a 3.21× size reduction. At the loose + PCK@50 threshold QAT is nearly lossless (−0.78 pp), i.e. coarse-localization + survives int8 but fine-localization does not. +- The size win is real and consistent (3.2× smaller, 3.351 MB → ~1.04 MB), but + **3.2× compression at −29 pp PCK@20 is a bad trade** when the half model already + fits comfortably in edge flash at fp32. Recommendation: **keep fp32 (or fp16) + for the half model on the edge**; do not ship this int8 variant as-is. + +### Observed fake-quant → int8 conversion gap (disclosed, not hidden) + +During QAT the **fake-quant** model's val PCK@20 reached 83.45% (epoch 3), but the +**converted int8** model scores 67.48% on test. A ~16 pp drop on `convert_fx` is a +real effect — the fbgemm int8 kernels are not bit-identical to the fake-quant +simulation (per-tensor activation quant + the axial-attention `einsum`/softmax path +quantize worse than the straight-through estimate predicts). This gap is the honest +reason QAT did not close the loss, and it is exactly the kind of number that would +be invisible if one only reported the fake-quant proxy. We report the **converted +int8** number as the deliverable, not the fake-quant proxy. + +## Reproduction + +```bash +ssh ruvultra 'cd ~/wiflow-std-bench && source venv/bin/activate && \ + python ~/quantize_half_int8.py --mode both --qat-epochs 3 2>&1' +``` + +- Script (committed): `v2/crates/wifi-densepose-train/scripts/quantize_half_int8.py` + (scp'd to `~/quantize_half_int8.py` on ruvultra for the run). +- Inputs (on ruvultra, unmodified): `~/wiflow-std-bench/sweep/half_best.pth`, + `~/wiflow-std-bench/preprocessed_csi_data/` (seed-42 file-level 70/15/15 split), + upstream `models`/`dataset`/`utils/metrics`/`losses` (DY2434/WiFlow @ 06899d29, + Apache-2.0), and `sweep/model_compact.py` (the half-model definition). +- Outputs (written, non-destructive): `~/wiflow-std-bench/sweep/int8/` — + `half_int8_qat.pth`, `half_int8_ptq_static.pth`, `int8_results.json`, + `int8_run.log`. **No existing file under `~/wiflow-std-bench` was modified.** +- Run metadata: host `ruvultra`, GPU RTX 5080, torch `2.11.0+cu128`, fbgemm engine, + `date_utc 2026-06-15T12:35:06Z`, QAT ≈ 97 s/epoch. + +## What is MEASURED vs CLAIMED + +- **MEASURED:** every PCK/MPJPE/size number in the table; the fp32 baseline (which + reproduces the recorded sweep `half` numbers); the PTQ collapse; the QAT partial + recovery; the fake-quant→int8 conversion gap; the 3.2× size reduction. +- **CLAIMED / not done here:** ONNX/TFLite export; on-real-edge (ESP32/Pi/Hailo) + latency or energy (int8 here is measured on x86 fbgemm, the dev box, **not** an + edge SoC — the size number transfers, a latency number does **not**); a + per-layer mixed-precision search that might keep the attention block in fp32; QAT + beyond 3 epochs or with learned-quant-range schedules. Those are the obvious next + levers if int8 is revisited; none is asserted as a result. + +## Honest scope / limitations + +- **Single eval split** — one seed-42 file-level test partition; no cross-room / + cross-environment generalization split (the GraphPose-Fi frontier from ADR-173 is + a separate, harder split and is not what is measured here). +- **In-domain only** — these are in-distribution test numbers; they say nothing + about the cross-environment robustness gap. +- **x86 int8, not edge-SoC int8** — accuracy and size transfer to an edge int8 + runtime; the runtime/latency does not (different kernels, different SoC). No + latency claim is made. +- **QAT lightly tuned** — 3 epochs, single LR, default fbgemm qconfig. A longer / + better-tuned QAT might narrow the −29 pp, but on the evidence here int8 does not + reach fp32 at PCK@20, and that is the reportable result today. + +## Consequences + +### Positive +- The "one untested edge lever" (QAT-int8 on the half model) is now MEASURED. The + edge int8 question for the half model is answered with reproducible numbers: at + the strict PCK@20 target it loses, and we can say so with a committed script. +- Establishes a reusable, metric-locked quantization+eval harness + (`quantize_half_int8.py`) for any future int8 attempt on these compact variants. + +### Negative +- None to the codebase (additive script + ADR + CHANGELOG only; no production Rust + or signal-pipeline change; Python deterministic proof hash + `f8e76f21a0f9852b70b6d9dd5318239f6b20cbcb4cdd995863263cecdc446f7a` unchanged). + +### Neutral +- The negative verdict means the half model stays fp32/fp16 on the edge for now. + int8 for these compact pose nets is parked pending the next-lever work above. + +## Links +- ADR-173 — metric-locked PCK/MPJPE harness (the locked normalization used here) +- ADR-152 — WiFi-Pose SOTA 2026 intake / WiFlow-STD benchmark / efficiency sweep + (produced `half_best.pth`) +- `docs/research/sota-nn-train-benchmark-brief.md` — §edge int8 (the "one untested + lever" this ADR measures) +- Script: `v2/crates/wifi-densepose-train/scripts/quantize_half_int8.py` diff --git a/api-docs/adr/ADR-176-ruview-swarm-nan-fail-open-safety-review.md b/api-docs/adr/ADR-176-ruview-swarm-nan-fail-open-safety-review.md new file mode 100644 index 00000000..b4151587 --- /dev/null +++ b/api-docs/adr/ADR-176-ruview-swarm-nan-fail-open-safety-review.md @@ -0,0 +1,103 @@ +# ADR-176: `ruview-swarm` NaN-Fail-Open Safety Review + +| Field | Value | +|-------|-------| +| **Status** | Accepted — 4 real safety bugs fixed + pinned; 2 issues documented for follow-up | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **SWARM-FAILCLOSED** | +| **Reviews** | ADR-148 (`ruview-swarm` drone swarm control plane) | +| **Milestone** | #9 (ungated-crate security sweep) — crate 1 of 4 | + +## Context + +`ruview-swarm` (ADR-148) is the drone swarm control plane — hierarchical-mesh +topology, Raft consensus, MARL, CSI sensing payload, MAVLink/PX4 command +dispatch. It is the highest-stakes of the four never-reviewed v2 crates: a defect +here can produce an **unsafe physical drone command**. It had no prior security +ADR. + +### Trust-boundary map +Untrusted input enters via `SwarmOrchestrator::receive_peer_state` / +`receive_peer_detection`, which accept full `DroneState` / `CsiDetection` serde +structs with **f64/f32 fields and no finite-check**, and via +`SwarmConfig`/`FhssConfig`/`Geofence` deserialization. The MAVLink wire formats in +`mavlink_messages.rs` are **integer-encoded** (i32 mm / u8) and provably cannot +carry NaN — so the NaN class is reachable through the **serde struct path, not the +MAVLink decode path**. Commands flow out to a `FlightController` (PX4/ArduPilot). + +The unifying bug class found: **IEEE-754 NaN/Inf silently defeating a safety +comparison** (`NaN < threshold` evaluates to `false`), causing safety logic to +**fail OPEN**. This is distinct from — but rhymes with — the NaN-state-poisoning +class found earlier in calibration/vitals/geo (there, NaN latched into persistent +state; here, NaN slips through a one-shot guard). Both are "non-finite input +defeats logic," and the fix discipline is the same: **reject non-finite at the +trust boundary, fail CLOSED.** + +## Decision + +Fix the four reachable fail-open bugs by making each safety predicate +non-finite-aware and fail-closed, each pinned by a fails-on-old test. Document +two further genuine issues that need larger, riskier changes rather than churning +them in a security pass. + +### Findings fixed (all MEASURED fails-on-old) + +| # | Severity | File:line | Issue | Fix | Pin (old behavior) | +|---|----------|-----------|-------|-----|--------------------| +| F1a | **HIGH** | `failsafe/mod.rs:51` | `nearest_neighbor_dist < collision_dist_m` fails open on a NaN peer position → **collision avoidance silently disabled** | `!is_finite() ||` → `EmergencyDiverge` | `test_nan_neighbor_distance_fails_closed_to_diverge` (old → `Nominal`) | +| F1b | **HIGH** | `failsafe/mod.rs:75` | NaN `battery_pct` bypasses every battery check → drone stays Nominal on unknown battery | `!is_finite() ||` → `ReturnToHome` | `test_nan_battery_fails_closed_to_rth` (old → `Nominal`) | +| F2 | **MEDIUM** | `security/geofence.rs:33` | NaN `z` altitude skips the altitude-breach check and point-in-polygon returns `Safe` → silent geofence bypass | leading non-finite coord → `HardBreach` | `test_nan_altitude_fails_closed` (old → `Safe`) | +| F3 | **MEDIUM/DoS** | `security/antijamming.rs:65,71,102` | empty deserialized `channels_mhz` → `% 0` **panic** in `next_hop`/`current_channel_mhz`/`evasive_hop`/`tick`, crashing the radio task | `len == 0` early-return (`0.0` sentinel) | `test_empty_channels_does_not_panic` (old → panic `divisor of zero`) | +| F4 | **LOW** | `sensing/multiview.rs:70` | NaN `victim_position` passes the `is_some()` filter and propagates into the fused "confirmed victim" location dispatched to the swarm | require finite confidence + position (drop) | `test_nan_victim_position_dropped_from_fusion` (old → non-finite fused position) | + +### Dimensions confirmed clean (with evidence) +- **MAVLink decode panic-safety** — `SwarmNodeState::decode(&[u8;20])` `try_into().unwrap()`s are over fixed const ranges of a fixed-size array → provably infallible; no arbitrary-length `&[u8]` decode path exists. +- **UWB/GPS anti-spoofing NaN-safe** — `(gps_dist - uwb_dist).abs() <= tol` already fails CLOSED on a NaN range (counts as inconsistent → spoof rejected); covered by `test_spoofed_gps_invalid`. +- **Bounded grid / no allocate-from-length-field** — `ProbabilityGrid` bounds-checks `cx/cy`; `pos_to_cell` uses saturating `as u32` (no UB). +- **Mesh `nearest_k` NaN-safe sort** — `partial_cmp(..).unwrap_or(Equal)` cannot panic on NaN. +- **No hardcoded secrets** — `MavlinkSigner` key is constructor-injected `[u8;32]`; grep-confirmed nothing embedded. + +### Documented, not fixed (genuine — deferred to avoid churn/regression risk) + +1. **Raft `AppendEntries` lacks the Log-Matching consistency check** + (`topology/raft.rs:187`). A follower appends a leader's entries when + `term >= current_term` **without validating `prev_log_index`/`prev_log_term`**, + so a malformed/byzantine leader can corrupt a follower's log — a genuine + consensus-safety gap. A correct fix reworks the log-append plus the + caller-side vote-tally contract (the existing `handle_message` delegates + tallying to the caller) — a larger change with test-rewrite risk, so it is + recorded here rather than rushed in a security pass. +2. **`MavlinkSigner::verify` uses a non-constant-time tag `==` and has no + replay/timestamp-window rejection** (`security/mavlink_signing.rs:64`). The + module doc already flags the replay limitation as a demo/test simplification. + Hardening (constant-time compare + monotonic timestamp window) is a focused + follow-up. + +These two are the recommended scope of the next `ruview-swarm` hardening pass. + +## Validation + +- `cargo test -p ruview-swarm --no-default-features` → **117 → 123** passed, 0 failed (+6 pins). +- All 6 new tests MEASURED fails-on-old (2× `Nominal`, `Safe`, panic `divisor of zero`, non-finite fused position); pass on the fix. +- `cargo test --workspace --no-default-features` → **exit 0**, 0 failed. +- `python archive/v1/data/proof/verify.py` → **VERDICT: PASS**, hash + `f8e76f21…46f7a` unchanged (ruview-swarm off the signal proof path). + +## Consequences + +### Positive +- Four reachable fail-open paths in a *physical-safety* control plane (collision + avoidance, battery RTH, geofence, anti-jamming radio task) now fail CLOSED on + hostile/degenerate input, each regression-pinned. +- Extends the "non-finite input defeats logic" defense from the state-poisoning + variant (calibration/vitals/geo) to the fail-open-comparison variant. + +### Negative / Neutral +- Two genuine issues (Raft log-matching, MAVLink signer) remain open by choice — + see Documented-not-fixed; they define the next hardening pass. + +## Links +- ADR-148 — `ruview-swarm` drone swarm control system +- ADR-172 — core/cli review (where the NaN bug-class root question was settled NO) +- ADR-127 — homecore review (sibling NaN/concurrency hardening) diff --git a/api-docs/adr/ADR-177-nvsim-degenerate-input-hardening.md b/api-docs/adr/ADR-177-nvsim-degenerate-input-hardening.md new file mode 100644 index 00000000..cf65f907 --- /dev/null +++ b/api-docs/adr/ADR-177-nvsim-degenerate-input-hardening.md @@ -0,0 +1,92 @@ +# ADR-177: `nvsim` Degenerate-Input Hardening (NV-Diamond Simulator) + +| Field | Value | +|-------|-------| +| **Status** | Accepted — 2 real MEDIUM bugs fixed + pinned; determinism preserved | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **NVSIM-FAILCLOSED** | +| **Reviews** | ADR-089 (`nvsim` NV-diamond magnetometer pipeline simulator) | +| **Milestone** | #9 (ungated-crate security sweep) — crate 2 of 4 | + +## Context + +`nvsim` (ADR-089) is a standalone, **WASM-ready** deterministic NV-diamond +magnetometer pipeline simulator — a forward-only leaf: +`scene → source → propagation → NV ensemble → digitiser → MagFrame + SHA-256 +witness`. It has no network surface, so the real attack surface is **degenerate +physical-parameter input** crossing the external boundary — specifically the +WASM `config_json` / `scene_json` entry points. + +Two properties matter for this crate that don't for others: it is billed +**deterministic** (a published cross-machine witness must reproduce bit-exactly), +and under `panic=abort` WASM any panic **aborts the whole module**. So a +config-induced panic is a denial-of-service, and a silent numeric corruption +defeats the simulator's entire purpose. + +## Decision + +Fix the two reachable degenerate-input bugs at their funnel points, each pinned +by a fails-on-old test, **without perturbing the deterministic happy path** (the +guards fire only on non-finite / degenerate input; the published witness is +unchanged). + +### Findings fixed (both MEASURED-reproduced) + +| # | Severity | Location | Issue | Fix | +|---|----------|----------|-------|-----| +| NVSIM-DT-01 | MEDIUM (DoS) | `pipeline.rs:58,95` | `dt = config.dt_s.unwrap_or(1.0 / f_s_hz)`; an external `f_s_hz == 0.0` → `dt = +Inf` → `(dt*1e6) as u64` saturates to `u64::MAX` → `(sample as u64) * dt_us` **panics `attempt to multiply with overflow`** at `sample ≥ 2` (debug/WASM-abort; garbage `t_us` in release). MEASURED: panic at `pipeline.rs:95:30`. | Sanitise `dt` (non-finite/non-positive → 1 µs fallback), cap the `u64` cast at `u64::MAX`, `saturating_mul` the timestamp — no config can overflow it. | +| NVSIM-NAN-01 | MEDIUM (silent corruption) | funnel `digitiser.rs::adc_quantise` (root: near-field clamp bypass in `source.rs`) | A non-finite scene param (NaN/Inf dipole position, Inf moment, NaN loop radius) **bypasses the near-field clamp** (`NaN < R_MIN_M == false` → the `1/r³` path runs → NaN field), and at the ADC `NaN as i32 == 0` (Rust saturating cast) emits a frame `b_pt=[0,0,0]` with **`ADC_SATURATED` CLEAR** — indistinguishable from a legitimate zero-field reading. MEASURED: `b=[NaN,NaN,NaN] sat=false` → `b_pt=[0,0,0] flags=0b0000`. | `adc_quantise`: any non-finite input → code `0` **with the saturation flag raised**; the pipeline's existing `adc_sat` OR-reduction propagates `ADC_SATURATED` onto the frame, making the corruption visible downstream. | + +This is the same **NaN-fail-open / NaN-poisoning** family seen across +calibration/vitals/geo and ruview-swarm — non-finite input defeating a guard — +but bounded here to a single frame (no cross-timestep accumulator). + +### Dimensions confirmed clean (with evidence) + +1. **Determinism integrity — clean.** One RNG only: `ChaCha20Rng::seed_from_u64(seed)`, + fully caller-seeded (grep: one `seed_from_u64`, **zero** `thread_rng`/`getrandom`/ + `SystemTime`/`Instant`/`HashMap`); `Cargo.toml` pins `rand`/`rand_chacha` + `default-features=false` (no OS entropy). Box–Muller draws + `gen_range(f64::EPSILON..=1.0)` (avoids `ln(0)=-Inf` by construction). Frame + bytes fixed LE; source summation order fixed by `Vec` order. **The published + cross-machine witness `cc8de9b0…93b4` (`proof_witness_publishes_a_known_value`) + passes UNCHANGED after both fixes** — the happy path is byte-identical; guards + touch only degenerate inputs. *Attested caveat (not a finding): libm + `cos`/`ln`/`sqrt` could differ x86↔wasm; the witness is documented as + x86_64-captured.* +2. **Panic-free deserialisation — clean.** `MagFrame::from_bytes` validates + len/magic/version, then per-field `buf[a..b].try_into().expect(...)` are over + fixed sub-ranges of an already-length-checked 60-byte buffer (provably + infallible). No `unsafe`, no `panic!`/`unreachable!` in production; every other + `unwrap`/`expect` is `#[cfg(test)]`. +3. **Div-by-zero / numerical landmines — clean.** `dipole_field`/`current_loop_field` + clamp `r_norm < R_MIN_M` before `1/r³`,`1/r²` (finite inputs); `shot_noise_floor` + guards `denom <= 0`; `vec3_normalise` guards `n < 1e-20`. The only hole was the + NaN *bypass* of the clamp — closed at the ADC funnel (NVSIM-NAN-01). + +## Validation + +- `cargo test -p nvsim --no-default-features` → **50 → 53** passed, 0 failed (+3 pins: + `degenerate_zero_sample_rate_does_not_panic`, + `non_finite_scene_input_flags_frame_instead_of_silently_zeroing`, + `adc_quantise_flags_non_finite_as_saturated`). +- `cargo test --workspace --no-default-features` → **exit 0**, 0 failed. +- `python archive/v1/data/proof/verify.py` → **VERDICT: PASS**, hash + `f8e76f21…46f7a` unchanged (nvsim off the signal proof path). +- nvsim's own cross-machine witness `cc8de9b0…93b4` reproduces unchanged. + +## Consequences + +### Positive +- A config-induced WASM-abort DoS and a silent NaN→fake-zero-field corruption are + closed at their funnel points, each regression-pinned, with the deterministic + witness proven intact. + +### Negative / Neutral +- None. Guards affect only degenerate inputs; happy-path output is byte-identical. + +## Links +- ADR-089 — `nvsim` NV-diamond magnetometer simulator +- ADR-176 — `ruview-swarm` (sibling NaN-fail-open review) +- ADR-172 — core/cli (where the NaN-bug-class root was settled NO) diff --git a/api-docs/adr/ADR-178-desktop-ipc-injection-and-capability-least-privilege.md b/api-docs/adr/ADR-178-desktop-ipc-injection-and-capability-least-privilege.md new file mode 100644 index 00000000..d57e3286 --- /dev/null +++ b/api-docs/adr/ADR-178-desktop-ipc-injection-and-capability-least-privilege.md @@ -0,0 +1,87 @@ +# ADR-178: `wifi-densepose-desktop` IPC Injection Fix + Capability Least-Privilege + +| Field | Value | +|-------|-------| +| **Status** | Accepted — 2 real MODERATE bugs fixed + pinned (MEASURED on Windows) | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **DESK-LOCKDOWN** | +| **Reviews** | `wifi-densepose-desktop` (Tauri v2 desktop app) | +| **Milestone** | #9 (ungated-crate security sweep) — crate 3 of 4 | + +## Context + +`wifi-densepose-desktop` is the Tauri v2 desktop app (ESP32 discovery, firmware +flashing, OTA, provisioning, server control). The real attack surface is the +**Tauri IPC boundary** — `#[tauri::command]` handlers that take arguments from the +webview/JS — and the **capability/allowlist scope**. The crate **builds and tests +on Windows** (Tauri 2.10.3, webview2 path, no GTK), so both findings are MEASURED, +not source-analysis-only. + +## Decision + +Fix the two real findings; attest the rest of the surface clean with evidence. + +### Findings fixed (both MEASURED) + +| # | Severity | Location | Issue | Fix | +|---|----------|----------|-------|-----| +| WDP-DESK-01 | MODERATE | `src/commands/discovery.rs:438` (`configure_esp32_wifi`) | Webview-supplied `ssid`/`password` are concatenated into newline-terminated serial commands (`wifi_config {} {}\r\n`, `set ssid {}\r\n`) with **no validation** → a `\r\n` in either field **injects an arbitrary follow-up firmware command** (`reboot`, `erase_nvs`) across the IPC trust boundary. | `validate_wifi_credentials()` — WPA2 length bounds (SSID 1–32, password 8–63) **+ reject all control chars** (`char::is_control()`), called fail-closed before any serial write. | +| WDP-DESK-02 | MODERATE | `capabilities/default.json:7-8` | `shell:allow-execute` + `shell:allow-open` granted to the webview but **unused** (Rust spawns via `std::process::Command`; the UI uses only `dialog.open`). A webview compromise (a UI-dependency XSS) → arbitrary **unscoped host command execution**. | Removed both `shell:` permissions (kept `core:default` + the two in-use `dialog:` perms); regenerated `gen/schemas/capabilities.json` now asserts `["core:default","dialog:allow-open","dialog:allow-save"]`. | + +Both are MODERATE (not HIGH): each requires a webview compromise or a malicious +local caller to weaponize. The unifying lesson is **least privilege at the IPC +boundary** — validate every webview-supplied argument that reaches a serial/FS/ +process sink, and grant only the capabilities actually exercised. + +### Tauri-command + capability audit (every handler) + +All 30+ command handlers were mapped. Only `configure_esp32_wifi` lacked input +validation on a string that reached a command sink (WDP-DESK-01). Every +subprocess uses `Command::new(prog).args([...])` (argv vector — no shell-string +interpolation), so `port`/`source`/`chip`/`baud` cannot inject a second command +even unvalidated. `tauri.conf.json` ships **no** `fs`/`http` plugin and **no** +`"all":true`/`"$HOME/**"` scope; after WDP-DESK-02 the allowlist is minimal. + +### Dimensions confirmed clean (with evidence) + +1. **Directory traversal / arbitrary file** — path args (`firmware_path`/`wasm_path`) + are blobs the local user selects via the native `dialog.open` picker; settings + I/O is a fixed filename under `app_data_dir`. No attacker-named path sink. +2. **Shell-string injection** — every subprocess is an argv vector; grep found no + shell-string interpolation anywhere. +3. **SSRF-to-secret** — `node_ip`-built URLs target the local ESP32 mesh and return + only device status JSON; no credential returned to the webview. +4. **Panic-on-input** — handlers use `.map_err(|e| e.to_string())?`; the one + `expect` is guarded by an `is_none()` early-return; provision/discovery + deserializers bounds-check every slice index (NVS size capped ≤ 4096). +5. **Hardcoded secrets** — `ota_psk` is a per-call `Option`, never embedded; + grep for embedded keys/tokens over `src/` is empty. +6. **Shell plugin genuinely unused** — `tauri_plugin_shell` is `init()`-ed but its + `Command`/`open` API is never invoked from Rust or the TS UI (which imports only + `@tauri-apps/plugin-dialog`) — confirming WDP-DESK-02 is safe to remove. + +## Validation + +- `cargo check -p wifi-densepose-desktop --no-default-features` → `Finished` (Windows, MEASURED). +- `cargo test -p wifi-densepose-desktop --no-default-features` → lib **18 → 21** (+3 validator pins: + `test_validate_wifi_credentials_rejects_injection` / `_rejects_out_of_range` / `_accepts_valid`), + integration 21/21, **0 failed**. +- Capability narrowing MEASURED: regenerated `capabilities.json` permission set verified. +- `python archive/v1/data/proof/verify.py` → **VERDICT: PASS**, hash `f8e76f21…46f7a` + unchanged (desktop off the signal proof path). + +## Consequences + +### Positive +- An IPC serial-command-injection path and an over-broad shell capability are + closed in the desktop app, each pinned / verified, with the rest of the + 30-command IPC surface attested clean. + +### Negative / Neutral +- None. The removed shell capability was unused; the validator rejects only + malformed/hostile credentials. + +## Links +- ADR-176 / ADR-177 — sibling Milestone-#9 reviews (ruview-swarm, nvsim) +- ADR-172 — core/cli review diff --git a/api-docs/adr/ADR-179-occworld-candle-checkpoint-load-hardening.md b/api-docs/adr/ADR-179-occworld-candle-checkpoint-load-hardening.md new file mode 100644 index 00000000..cc2924c6 --- /dev/null +++ b/api-docs/adr/ADR-179-occworld-candle-checkpoint-load-hardening.md @@ -0,0 +1,81 @@ +# ADR-179: `wifi-densepose-occworld-candle` Checkpoint-Load Hardening + +| Field | Value | +|-------|-------| +| **Status** | Accepted — 1 HIGH + 2 LOW bugs fixed + pinned (MEASURED on Windows) | +| **Date** | 2026-06-15 | +| **Deciders** | ruv | +| **Codename** | **OCCWORLD-DTYPE** | +| **Reviews** | `wifi-densepose-occworld-candle` (Candle occupancy-world model) | +| **Milestone** | #9 (ungated-crate security sweep) — crate 4 of 4 — **CLOSES the milestone** | + +## Context + +`wifi-densepose-occworld-candle` is a Candle-based occupancy-world model +(VQ-VAE + transformer over occupancy tokens). The real risk surface for an ML +crate is degenerate-input / malformed-weights handling: a `#[forbid(unsafe_code)]` +crate can still **panic** (a DoS, and under WASM an abort) when a tensor op hits an +inconsistent shape. The crate **builds and tests on Windows**, so all findings are +MEASURED. + +## Decision + +Fix the three reachable bugs, each pinned by a fails-on-old test; attest the rest +clean with evidence. + +### Findings fixed (all MEASURED) + +| # | Severity | Location | Issue | Fix | +|---|----------|----------|-------|-----| +| 1 | **HIGH** | `model.rs:95` (`Dtype::I32 => Some(DType::I64)`) | **Crash on any int32-tensor checkpoint.** An I32 byte buffer (4 B/elem) is handed to `from_raw_buffer(.., I64, shape, ..)`; candle derives `elem_count = data.len()/8`, **halving** the count while keeping the original shape → a tensor that claims 2× its storage. Reading it **panics** with a slice-OOB (`range end index 6 out of range for slice of length 3`) inside candle-core. A checkpoint with any int32 tensor (index/buffer tensors are common in PyTorch exports) → **DoS on load**. | Map `I32 → DType::I32`, `I16 → DType::I16` (both first-class candle dtypes). Pinned by `int32_tensor_loads_with_consistent_shape_and_values` (panics on old, passes on new). | +| 2 | LOW | `inference.rs::predict` | Frame/batch dims weren't validated (only H/W/D were): `f_in > num_frames*2` over-indexes the temporal embedding → a cryptic candle `InvalidIndex` *error* (not a panic — candle bounds-checks); zero frame/batch feeds a zero-element tensor. | Boundary guard rejects zero / over-capacity frame+batch with a clear `ShapeMismatch`. 5 pins. | +| 3 | LOW | `vqvae.rs:141` (`z.elem_count() / last`) | **Divide-by-zero panic** in public `VQCodebook::encode` on a rank-0 / empty-last-dim tensor (`last == 0`). | Fail-closed guard returns a clear error. Pinned by `encode_rejects_scalar_without_panicking`. | + +The HIGH finding is the notable one: the crate's own dtype mapping **defeated** +the upstream `safetensors::validate()` byte-length guarantee by misdeclaring the +dtype — the one place malformed/widened weights could reach a panicking candle op. + +### Dimensions confirmed clean (with evidence) + +- **Panic surface** — grep for `unwrap()/expect()/panic!/unreachable!` across `src/` + → **zero in production paths**; all ops use `?`/`map_err`; the `last().unwrap_or(&0)` + is now guarded. `as` casts operate only on config-bounded/internal values. +- **NaN-state-poisoning (the named class) — N/A.** The engine is **stateless between + `predict` calls** (no persistent world-model buffer to latch into), and input is + `u8` class indices (non-finite input structurally impossible). NaN weights flow to + `argmax` (deterministic, bounded to a valid class index) — no panic, no persistence. +- **Unbounded alloc / shape-data mismatch from malformed weights** — defended upstream + by `safetensors::validate()` (overflow-checked `nelements*dtype.size()` vs declared + byte range + contiguous-offset + buffer-length checks), rejected before reaching + candle. Finding #1 was the one place the crate defeated that guarantee. +- **Model/path loading** — `load`/`load_safetensors` check `path.exists()` → typed + `CheckpointNotFound`; corrupt bytes → `CheckpointParse` (pinned). No path-traversal + surface (caller-supplied path, opened read-only, never joined with untrusted segments). +- **Secrets** — grep clean (only `token_h`/`token_w` config fields match `token`). +- **Determinism** — the crate's central honesty claim, verified by the pre-existing + `tests/predict_honesty.rs` (3 tests, still pass). +- `unsafe_code = "forbid"` in the manifest. + +## Validation + +- `cargo test -p wifi-densepose-occworld-candle --no-default-features` → **31/31** + (lib 17, checkpoint_loading 4, input_validation 5, predict_honesty 3, doctests 2), + 0 failed. +- `cargo test --workspace --no-default-features` → 0 failed across every crate (a lone + `wifi-densepose-desktop --test api_integration` "Access is denied (os error 5)" was a + Windows file-lock/AV flake — re-ran isolated 21/21, unrelated). +- `python archive/v1/data/proof/verify.py` → **VERDICT: PASS**, hash `f8e76f21…46f7a` + unchanged (occworld off the signal proof path). + +## Consequences + +### Positive +- A checkpoint-load DoS (the int32 dtype-widening panic) and two degenerate-input + panics are closed in the world-model crate, each pinned. **Milestone #9 (all 4 + ungated crates) is complete.** + +### Negative / Neutral +- None. Guards reject only malformed/degenerate inputs. + +## Links +- ADR-176 / ADR-177 / ADR-178 — sibling Milestone-#9 reviews (ruview-swarm, nvsim, desktop) diff --git a/api-docs/adr/ADR-182-npx-ruview-harness-via-metaharness.md b/api-docs/adr/ADR-182-npx-ruview-harness-via-metaharness.md new file mode 100644 index 00000000..a434cd60 --- /dev/null +++ b/api-docs/adr/ADR-182-npx-ruview-harness-via-metaharness.md @@ -0,0 +1,279 @@ +# ADR-182: `npx ruview` — A RuView Agent Harness Minted via MetaHarness + +| Field | Value | +|-------|-------| +| **Status** | Accepted — **P1+P2 implemented & validated** (`harness/ruview/`, 17/17 tests, MCP handshake + `ruview.verify` PASS against the real repo, packs to 16.7 kB / 21 files) · P3 publish-ready (name decision pending) · P4 (router + provenance) designed | +| **Date** | 2026-06-17 | +| **Deciders** | ruv | +| **Codename** | **RUVIEW-HARNESS** | +| **Builds on** | MetaHarness (`metaharness@0.1.15`, `@metaharness/kernel`, `@metaharness/host-*`, `@metaharness/router`), the `ruview-*` Claude Code subagents (`ruview-onboarding-guide`, `ruview-config-engineer`, `ruview-training-engineer`), the `wifi-densepose` CLI (`calibrate`/`enroll`/`train-room`/`room-watch`), the sensing-server, ADR-028 (witness verification), ADR-095/096 (rvCSI runtime), ADR-260/262 (RuField bridge) | +| **Supersedes** | none | + +## Context + +RuView (WiFi-DensePose) is a deep stack — 15 Rust crates, an ESP32 firmware line, +a sensing-server, a CLI, ~180 ADRs, a calibration pipeline, training recipes, and a +hard cultural rule that **every claim must be independently reproducible** (the +"prove everything" ethos, after the project was accused of AI-slop). The barrier to +entry is correspondingly steep: a newcomer who wants to "set up WiFi sensing" must +discover the right firmware variant, provision an ESP32 over a Windows-only Python +subprocess, point it at the sensing-server, run `calibrate` → `enroll` → +`train-room`, and know which numbers are MEASURED vs CLAIMED. We already encode this +knowledge as **Claude Code subagents** (`ruview-onboarding-guide`, +`ruview-config-engineer`, `ruview-training-engineer`) — but those only exist inside +*this* repo's `.claude/agents/`, only on Claude Code, and only for someone who has +already cloned the monorepo. + +Separately, this session shipped **MetaHarness** (`metaharness@0.1.15`): a tool that +*"mints a custom AI agent harness from any repo"*, runnable on **9 hosts** +(claude-code, codex, pi-dev, hermes, openclaw, rvm, copilot, opencode, +github-actions) over a wasm-primary / NAPI-RS-fallback **kernel**, with a +**cost-optimal model router** (`@metaharness/router`, the productized DRACO Phase-2 +k-NN finding) and ed25519/SLSA/SBOM provenance baked in. Crucially, MetaHarness +**already ships a `vertical:ruview` template** in its template list. That template +is generic scaffolding; it is not wired to RuView's actual tools, agents, or the +"prove everything" guardrails. + +The gap: **there is no single, host-portable, provenance-signed entry point that +gives any user an AI agent that actually knows how to operate RuView.** A user +should be able to run one command — + +```bash +npx ruview +``` + +— in an empty directory (or alongside an ESP32) and get an agent harness that can +onboard them, configure firmware, drive a live capture, train a room model, and +**refuse to overstate accuracy** — on whichever coding host they already use. + +## Decision + +**Mint a first-class RuView agent harness from this repo using MetaHarness, harden +its `vertical:ruview` template into a RuView-specific harness with a real MCP tool +surface and the project's honesty guardrails, and publish it as `npx ruview`.** + +`npx ruview` is *not* a new runtime. It is a **thin, versioned distribution** of a +MetaHarness harness: the kernel + host adapters + a RuView "genome" (skills, agents, +MCP tools, guardrails) generated from and pinned against this monorepo. The harness +is the product; `npx ruview` is the front door. + +### Why mint-from-repo instead of hand-writing a harness + +MetaHarness's value here is exactly the work we would otherwise hand-roll across 9 +hosts: host-specific config (`.claude/settings.json` MCP + hooks for claude-code, +the codex/copilot/opencode equivalents), the kernel that abstracts wasm-vs-native, +the cost router, and the provenance chain. We write the **RuView knowledge once** as +host-neutral genome assets; MetaHarness projects them onto each host adapter. This +also keeps the harness regenerable: when the CLI or an ADR changes, re-mint and +re-pin rather than maintaining 9 divergent copies. + +### What the harness contains (the RuView genome) + +1. **Skills / playbooks** (host-neutral markdown, projected to each host's skill + format): + - `onboard` — zero-to-sensing path picker (Docker demo / repo build / live + ESP32), the physics caveats, the hardware table. Port of + `ruview-onboarding-guide`. + - `provision-node` — ESP-IDF v5.4 Windows-subprocess build/flash/provision flow + (the exact MSYSTEM-stripped invocation from `CLAUDE.local.md`), firmware + variant selection (8MB display / 4MB no-display / C6), NVS + WiFi + channel / + MAC-filter overrides (ADR-060). + - `calibrate-room` — `baseline → enroll → extract → train` via the + `wifi-densepose` CLI (`calibrate`/`calibrate-serve`/`enroll`/`train-room`/ + `room-watch`, ADR-151). + - `train-pose` — camera-supervised + camera-free training, the MEASURED-vs-CLAIMED + discipline, the mean-pose baseline check (ADR-079, ADR-152, ADR-181). + - `verify` — run the witness bundle + Python proof (`verify.py` → VERDICT: PASS), + ADR-028. + - Ports of `ruview-config-engineer` and `ruview-training-engineer`. + +2. **MCP tool surface** (`@metaharness/kernel`-hosted MCP server, one schema per + capability — see "MCP tools" below). This is what makes the harness *operate* + RuView, not just talk about it. + +3. **Guardrails** (the differentiator): the harness's system prompt and a + pre-output hook enforce the "prove everything" rule — accuracy numbers must be + tagged MEASURED (with a reproducer) or CLAIMED; the agent must run the mean-pose + baseline before quoting PCK; firmware fixes are never presented as + hardware-validated without a real boot log (the exact discipline this session + followed for `v0.8.1-esp32`). + +4. **Host adapters** — claude-code first (P1), then codex / opencode / copilot / + pi-dev / hermes / rvm / github-actions (P3+), each via the published + `@metaharness/host-*` package. + +5. **Router** — `@metaharness/router` routes each step to the cheapest adequate + model (e.g. a var-rename or a log-grep → Haiku; calibration-math reasoning or a + security review → Sonnet/Opus), mirroring the repo's 3-tier routing (ADR-026). + +### MCP tools (the operational surface) + +| Tool | Wraps | Purpose | +|------|-------|---------| +| `ruview.onboard` | docs + agent | Pick a setup path, print the next concrete command | +| `ruview.node.flash` | ESP-IDF subprocess (ADR `CLAUDE.local.md`) | Build + flash a firmware variant to a COM port | +| `ruview.node.provision` | `provision.py` | Set SSID/password/target-ip/channel/MAC-filter over serial | +| `ruview.node.monitor` | pyserial | Stream boot log; assert CSI is flowing (MGMT+DATA) | +| `ruview.server.up` | sensing-server | Start the Axum sensing-server (`:3000`/`:5005`/`:8765`) | +| `ruview.calibrate` | `wifi-densepose calibrate`/`enroll`/`train-room` | Run the ADR-151 room pipeline | +| `ruview.room.watch` | `wifi-densepose room-watch` | Live presence/vitals from a trained room | +| `ruview.verify` | `scripts/generate-witness-bundle.sh` + `verify.py` | Produce/verify the witness bundle (must be N/N PASS) | +| `ruview.claim.check` | static lint | Scan output for untagged accuracy claims; flag MEASURED-vs-CLAIMED | + +Each tool returns structured JSON and is fail-closed: a tool that cannot prove its +result (e.g. `ruview.node.monitor` sees no CSI callbacks) returns an honest negative, +never a fabricated success — consistent with the RuField `map_privacy` fail-closed +posture (ADR-262 §3.3). + +### The mint + pin flow (how the harness is produced) + +```bash +# P1 — mint from this repo, claude-code host, RuView vertical +npx metaharness ruview --template vertical:ruview --host claude-code \ + --from-existing . --description "RuView WiFi-sensing operator agent" \ + --target ./harness/ruview + +# readiness + fit/cost/safety scorecards (ADR-041) — gate before publish +npx metaharness genome . # 7-section repo readiness +npx metaharness score . --json # fit / cost / safety +npx metaharness analyze . # recommended harness plan (no-exec) +``` + +The minted harness is committed under `harness/ruview/` and **pinned** (kernel + +host-adapter + router versions locked) so `npx ruview` is reproducible. Re-minting on +a CLI/ADR change is a reviewed PR, not an implicit regeneration. + +### Distribution: `npx ruview` + +A small published package whose `bin` boots the pinned harness via the kernel: + +- **Preferred name:** `ruview` (currently **free** on npm — verified 2026-06-17). +- **Risk:** npm's typosquat filter may reject `ruview` as too close to `review` / + `preview` (this session hit exactly that on `ruvn`→`levn`/`raven` and + `worldgraph`→`world-graph`). **Fallback:** publish scoped `@ruvnet/ruview` (also + free) and/or `npx ruvnet/ruview` straight from GitHub. Decide at publish time; + do not unpublish to rename (the 24-h name-lock lesson from `worldgraphs`). +- `bin: { "ruview": "bin/cli.js" }` — note **`bin/cli.js`, not `./bin/cli.js`** (npm + strips the `./` form; this broke `ruvn@0.1.0` this session). +- `npx ruview` with no args → `onboard` skill (interactive path picker). + `npx ruview [...]` → run a specific skill. `npx ruview --host codex` → + install the harness into an existing repo for that host. + +## Architecture + +``` + npx ruview (thin bin — boots the pinned harness) + │ + @metaharness/kernel (wasm primary · NAPI-RS native fallback) + ├── host adapter ── claude-code | codex | opencode | copilot | pi-dev | hermes | rvm | github-actions + ├── @metaharness/router (k-NN cost-optimal model routing — DRACO P2 / ADR-026) + └── RuView genome (pinned) + ├── skills onboard · provision-node · calibrate-room · train-pose · verify + ├── mcp tools ruview.node.* · ruview.calibrate · ruview.room.watch · ruview.verify · ruview.claim.check + └── guardrails MEASURED-vs-CLAIMED · mean-pose baseline · no-unvalidated-firmware-claims + │ + RuView assets (the real system the agent drives) + ├── wifi-densepose CLI calibrate / enroll / train-room / room-watch + ├── sensing-server :3000 / :5005 / :8765 + ├── ESP-IDF subprocess build / flash / provision / monitor (COM8/COM9/COM12) + └── witness bundle + verify.py +``` + +Provenance: the harness ships an **ed25519 witness + SBOM (SPDX) + SLSA** chain +(MetaHarness already does this for minted harnesses), so a recipient can verify the +RuView harness was built from a specific monorepo commit — the agentic analogue of +the firmware witness bundle (ADR-028). + +## Phases + +- **P1 — Mint & pin (claude-code).** `npx metaharness ruview --template + vertical:ruview --from-existing . --host claude-code`. Port the three `ruview-*` + subagents into host-neutral genome skills. Commit under `harness/ruview/`, pin + versions. Acceptance: `npx metaharness score .` ≥ threshold; the harness can run + `onboard` and `verify` end-to-end locally. +- **P2 — MCP tool surface.** Implement the `ruview.*` MCP tools over the kernel + (start with `onboard`, `verify`, `claim.check`, `node.monitor` — the read-only / + proving tools), then the mutating ones (`node.flash`, `provision`, `calibrate`). + Acceptance: `ruview.verify` returns the witness bundle PASS as structured JSON; + `ruview.claim.check` flags a seeded untagged "100% accuracy" string. +- **P3 — Publish `npx ruview` + multi-host.** Publish the bin package (name decision + per Distribution). Add codex / opencode / copilot / pi-dev / hermes / rvm / + github-actions adapters. Acceptance: `npx ruview` cold-starts on ≥3 hosts and runs + `onboard`; provenance verifies. +- **P4 — Router + guardrail hardening.** Wire `@metaharness/router`; calibrate the + 3-tier routing on a RuView task set. Make the MEASURED-vs-CLAIMED guardrail a hard + pre-output gate. Acceptance: a benchmark of RuView tasks shows cost reduction vs + all-Opus with no quality regression; the guardrail blocks an untagged accuracy + claim in a red-team prompt. + +## Consequences + +**Positive** +- One reproducible, signed entry point (`npx ruview`) that operates RuView on the + host the user already has — onboarding goes from "clone a 15-crate monorepo" to a + single `npx`. +- The "prove everything" ethos becomes **executable**, not just documentation: the + harness *enforces* MEASURED-vs-CLAIMED and the mean-pose baseline. +- Knowledge written once (host-neutral genome) instead of 9× per host; regenerable + from the repo as the system evolves. +- Dogfoods MetaHarness on a hard real vertical, surfacing bugs back to + `agent-harness-generator` (this session already filed #9–#13 there). + +**Negative / risks** +- **Drift:** a pinned harness goes stale as the CLI/ADRs move; mitigated by a + re-mint-on-change PR ritual and a CI check that the genome's referenced + CLI flags still exist. +- **Surface area:** mutating MCP tools (`node.flash`, `provision`) touch hardware and + the network — must be permission-gated and fail-closed; the firmware-flash tool + must never claim hardware validation without a captured boot log. +- **Name/typosquat:** `ruview` may be rejected at publish; scoped fallback decided in + P3. Do not unpublish-to-rename. +- **Host parity:** not all 9 hosts support MCP + hooks equally; the guardrail gate + may degrade to advisory on weaker hosts — must be disclosed in the badge, not + hidden (same honesty principle as ADR-181's backend badge). +- **Windows-coupled tooling:** the ESP-IDF flow is Windows-subprocess-specific + today; the `node.*` tools are gated to that environment until a cross-platform + path exists. + +## Alternatives considered + +1. **Keep the `ruview-*` subagents repo-local (status quo).** Zero new surface, but + stays Claude-Code-only and clone-gated; no portable front door. Rejected — it's + the gap this ADR exists to close. +2. **Hand-write a bespoke `npx ruview` harness (no MetaHarness).** Full control, but + re-implements the kernel, 9 host adapters, the router, and the provenance chain + we already ship — months of duplicated work and 9 divergent configs to maintain. + Rejected. +3. **Use the generic `vertical:ruview` template as-is.** It's scaffolding with no + real tools or guardrails — it would *talk about* RuView without being able to + *operate* it or enforce honesty. Rejected as insufficient; P2 is precisely the + hardening that makes it real. +4. **Ship only an MCP server (no harness/host adapters).** Covers tools but not the + skills, routing, guardrails, or multi-host projection — a strictly smaller subset + of this design. Folded in as the P2 layer rather than the whole. + +## Open questions + +- Final published name: bare `ruview` vs scoped `@ruvnet/ruview` vs GitHub-only + `npx ruvnet/ruview` — resolve against the typosquat filter at P3. +- Does the harness bundle the `wifi-densepose` binary, shell out to a user-installed + one, or offer both? (Leaning: shell out; print install guidance if absent.) +- Where do the `node.*` hardware tools live for non-Windows users — defer, or wrap + the rvCSI runtime (ADR-095/096) which is cross-platform Rust? +- Should `ruview.verify` gate `npx ruview` self-tests in CI (harness can't publish if + the witness bundle regresses)? +- Relationship to the RuField MFS harness surface (ADR-260/262) — one harness with a + RuField skill, or a sibling `npx rufield`? + +## References + +- MetaHarness: `metaharness@0.1.15` (`npx metaharness`, templates incl. + `vertical:ruview`; hosts: claude-code/codex/pi-dev/hermes/openclaw/rvm/copilot/ + opencode/github-actions), `@metaharness/kernel`, `@metaharness/router`, + `@metaharness/host-*`, repo `github.com/ruvnet/agent-harness-generator`. +- RuView subagents: `ruview-onboarding-guide`, `ruview-config-engineer`, + `ruview-training-engineer` (`.claude/agents/`). +- ADR-026 (3-tier model routing), ADR-028 (witness verification), ADR-041 + (MetaHarness scorecards), ADR-060 (channel / MAC-filter overrides), ADR-079 + (camera ground-truth training), ADR-095/096 (rvCSI runtime), ADR-151 (per-room + calibration), ADR-152/181 (WiFlow / browser pose), ADR-260/262 (RuField bridge). diff --git a/api-docs/adr/ADR-183-onboard-led-gamma-stimulus-csi-colormap.md b/api-docs/adr/ADR-183-onboard-led-gamma-stimulus-csi-colormap.md new file mode 100644 index 00000000..329be5e1 --- /dev/null +++ b/api-docs/adr/ADR-183-onboard-led-gamma-stimulus-csi-colormap.md @@ -0,0 +1,98 @@ +# ADR-183: Onboard LED as a 40 Hz Gamma Stimulus, Colour-Mapped from Live CSI via `ruv-neural-viz` + +| Field | Value | +|-------|-------| +| **Status** | Accepted — implemented & hardware-confirmed on ESP32-S3 N16R8 (COM8) | +| **Date** | 2026-06-17 | +| **Deciders** | ruv | +| **Codename** | **GAMMA-VIZ** | +| **Builds on** | `ruv-neural-viz::ColorMap` (now `no_std` — ruvnet/ruv-neural#3 / RuView#1126), the ESP32 edge `motion_energy` metric (`edge_processing.c`), PR #962 (WS2812 on GPIO 48) | + +## Context + +Two threads converged. (1) `ruv-neural-viz::ColorMap` — the viridis/cool-warm +palette the rUv-Neural stack uses to render brain-topology graphs — was `std`-only, +so it couldn't run on the ESP32. (2) The onboard WS2812 on the S3 CSI node was dead +weight: the firmware only cleared it on boot (and on the wrong pin for N16R8 — GPIO +38 vs the actual 48, see #962). + +The ask: make the LED do something real and honest, using the project's own visual +capability — not a decorative blink. The natural fit is a **40 Hz gamma stimulus** +(the GENUS gamma-entrainment frequency from Alzheimer's light-therapy research) +whose **colour is driven by live sensed motion**, so the node's front panel is both +a known bio-stimulus waveform and a truthful readout of what the CSI is detecting. + +## Decision + +### Part A — make `ColorMap` `no_std` + +`colormap.rs` is self-contained (no cross-crate deps), so expose it on `no_std` +targets. The only blockers were two `std`-only `f64` ops: + +- `f64::round` / `f64::abs` → replaced with `core`+`alloc`-safe helpers `fround` + (round via `f64 as i64` truncation — a `core` cast, no `libm`) and `fabs`. +- `Vec`/`String`/`format!` → from `alloc`. + +The graph-bound modules (`animation`/`ascii`/`export`/`layout`) and their heavy deps +move behind a default `std` feature; `--no-default-features` builds the crate `no_std` +and exposes only `colormap`. Output is **byte-identical** (8/8 colormap tests pass with +the same RGB values), so this is a pure portability change. + +### Part B — the LED stimulus (firmware) + +`firmware/esp32-csi-node/main/main.c`, on boot: + +- WS2812 on **GPIO 48** (N16R8 / DevKitC-1 v1.1; GPIO 8 on C6). +- An `esp_timer` periodic at **12 500 µs toggles a square wave → 40 Hz, 50 % duty** + (full-on / full-off — a *perceptible* gamma flicker, not a colour drift). +- **ON-phase colour = live CSI motion.** Each ON phase reads `edge_get_vitals().motion_energy`, + normalises it (`/ LED_MOTION_FULLSCALE`, clamped `[0,1]`), and indexes a **60-step + viridis LUT generated from `ColorMap::viridis().map()`** — still = dark purple, + strong motion = yellow. + +The LUT is baked from the real crate (Part A makes the same `ColorMap` embeddable +for a future direct FFI path once the ESP Rust toolchain is in CI). The colours are +therefore provably `ruv-neural-viz`'s, and the motion is provably real. + +## Honesty (what it is and is not) + +- **40 Hz is a real square-wave stimulus** (12.5 ms on / 12.5 ms off), not a label on + a colour sweep. It is *not* tied to any measured 40 Hz brain rhythm — it is an + *output* stimulus at the gamma frequency, not a readout of neural gamma. +- **Colour is a real CSI readout** — `motion_energy` is the on-device phase-variance + motion metric the node already computes; no fabrication. At rest the LED sits at the + purple (low) end and flickers there. +- No therapeutic claim is made. 40 Hz GENUS entrainment is cited as the *origin of the + frequency choice*, not as a validated medical effect of this device. + +## Consequences + +**Positive** +- The LED is now an honest front-panel: gamma-frequency flicker + a live motion readout. +- `ColorMap` is embeddable (`no_std`), unblocking on-device use of the rUv-Neural + palette beyond this LED. +- Confirms #962's GPIO-48 fix visually (the LED lights on N16R8). + +**Negative / risks** +- Changes the *default* firmware behaviour: the onboard LED animates instead of staying + off. Now **gated by `CONFIG_LED_GAMMA_VIZ`** (default `y`); set it `n` for a dark, + lower-power boot (the LED is just cleared) — no source change needed. +- A 40 Hz flicker can be an issue for photosensitive users; document on the enclosure + and disable `CONFIG_LED_GAMMA_VIZ` in those deployments. +- The saturation point is now `CONFIG_LED_MOTION_FULLSCALE_MILLI` (default 250 = 0.25), + operator-tunable; still not auto-calibrated per-environment. +- The colour uses a baked LUT, not the live Rust `ColorMap` (FFI path deferred — needs + the ESP Rust/xtensa toolchain, not yet in CI). + +## Validation + +- `ruv-neural-viz`: `cargo build` (std) ✓, `cargo test colormap` 8/8 ✓ (identical RGB), + `cargo build --no-default-features` compiles `no_std` ✓. +- Firmware: built (1.13 MB), flashed to ESP32-S3 N16R8 (COM8). Boot log: + `Onboard WS2812: 40 Hz gamma flicker (GENUS), colour=CSI motion via ruv-neural-viz, GPIO 48`; + CSI continues (27–38 pps), `motion=0.00` at rest → purple flicker as designed. +- Full on-device (xtensa) Rust build of `ColorMap` not run — ESP Rust toolchain absent. + +## References +- ruvnet/ruv-neural#3 (ColorMap no_std), RuView#1126 (submodule bump), #962 (GPIO 48). +- Singer/Tsai GENUS 40 Hz gamma entrainment (origin of the frequency, not a device claim). diff --git a/api-docs/adr/ADR-262-rufield-ruview-integration.md b/api-docs/adr/ADR-262-rufield-ruview-integration.md index 2484dc65..64702aaf 100644 --- a/api-docs/adr/ADR-262-rufield-ruview-integration.md +++ b/api-docs/adr/ADR-262-rufield-ruview-integration.md @@ -2,7 +2,7 @@ | Field | Value | |-------|-------| -| **Status** | Proposed — P1 implemented | +| **Status** | Proposed — **P1 + P3 implemented** (live `/api/field` + `/ws/field`; P3 signs with a **dedicated dev/sensing key**, deferring the §8 Q1 `cog-ha-matter` key-ownership decision to P2) | | **Date** | 2026-06-14 | | **Deciders** | ruv | | **Codebase target** | New thin bridge crate `wifi-densepose-rufield` (v2 workspace member); taps `wifi-densepose-sensing-server` emit path + `wifi-densepose-engine` `TrustedOutput`; depends on `vendor/rufield/crates/rufield-*` via path (the `vendor/rvcsi` pattern) | @@ -30,7 +30,15 @@ This project has been publicly accused of "AI slop." This ADR answers with **evi - **Privacy (§3.3 crux)** — `map_privacy()` maps by information content, **fail-closed**: `Raw → P0`, `Derived → P4` (or `P5` if identity-bound — **never P1**), `Anonymous → P2`, `Restricted → P2`; a `demoted` cycle floors egress to ≥ P2. - **Gates that pass** (`tests/p1_gates.rs`, 15 tests / 0 failed = 5 unit + 9 integration + 1 doc): round-trip (snapshot → `FieldEvent` → serde → equal); `is_fusable` (verified ed25519 receipt); `RuFieldFusion::ingest` accept + `infer()` runs; **privacy-safety** (`gate_privacy_safety_derived_never_maps_to_low_privacy` — `Derived → P4/P5`, never P1; full §3.3 table; fail-closed demotion); determinism (same snapshot + same signer seed → byte-identical event). -**Deferred:** the §3.3 *provenance carrier* recommendation (reuse the `cog-ha-matter` SHA-256+Ed25519 chain + embed the BLAKE3 engine witness) is **not** in P1 — P1 takes a dedicated `Signer` param (the §8 open question 1 key-ownership decision is unresolved). P2's BLAKE3-embed, P3 (live `/ws/field` surfacing — the bridge is **not** wired into the running server yet), and P4 (multi-modality) remain future work. **No accuracy is claimed** (§0 / §6) — P1 is tested plumbing + a safe privacy mapping. +**P3 (§4) is implemented** as the live RuField surface in `wifi-densepose-sensing-server` (the bridge is now wired into the running server): + +- **Tap** — at the ESP32 governed-trust cycle (`main.rs` `observe_cycle` ~`:5886` / `SensingUpdate` build ~`:5938`), a new `emit_rufield_event` joins the cycle's `SensingUpdate` (features / classification / signal_field) with the engine's recorded `effective_class` / `demoted` trust state into a `wifi_densepose_rufield::SensingSnapshot`, then `snapshot_to_field_event(&snap, &signer)`. Existing endpoints (`/ws/sensing` etc.) are **unchanged** — purely additive. +- **Surface** — `GET /api/field` (latest signed `FieldEvent`s + signer pubkey + a `dev_signing_key` flag) and `GET /ws/field` (broadcast stream, mirroring `/ws/sensing`), both mounted on the HTTP port and `/ws/field` also on the WS port. A small bounded ring buffer (`FIELD_RING_CAPACITY = 64`) holds recent **network-surfaced** events. New handler code lives in `src/rufield_surface.rs`, not in the 8k-line `main.rs`. +- **Signer (defers the P2 key decision)** — a **dedicated standalone `Signer`** held in server state, seeded from `WDP_RUFIELD_SIGNING_SEED` (64-hex or ≥32-byte value), else a deterministic dev default with a logged `WARN`. Reusing the `cog-ha-matter` Ed25519 key (§8 Q1) is the **deferred P2** decision — P3 uses a standalone sensing key so it does not pre-empt that call. +- **Egress privacy (fail-closed)** — `network_egress_allowed` is *stricter* than `DefaultPrivacyGuard` for an unattended live surface: only **P1/P2** leave the box; P0 (raw) and P3/P4/P5 (identity/biometric/aggregate above the default P2 ceiling) are held edge-local. A `Derived` cycle maps to P4/P5 and is therefore **never** surfaced. No-presence cycles emit nothing (no phantom events). +- **Gates that pass** (`tests/rufield_surface_test.rs`, 4 integration via `tower::oneshot` + 4 module unit, 0 failed): a well-formed **signed** event (`Modality::WifiCsi`, P2 not P1, `is_fusable` ed25519-verified, real timestamp); **empty cycle → no phantom**; **privacy-safety** — an injected `Derived` trust never surfaces on `/api/field`; a mixed stream surfaces only egress-safe events. + +**Deferred:** the §3.3 *provenance carrier* recommendation (reuse the `cog-ha-matter` SHA-256+Ed25519 chain + embed the BLAKE3 engine witness) is **not** in P1/P3 — both take a dedicated `Signer` (the §8 open question 1 key-ownership decision is unresolved; P3 uses a standalone dev/sensing key precisely so it does not pre-empt P2). P2's `cog-ha-matter` key reuse + BLAKE3-embed, and P4 (multi-modality), remain future work. **No accuracy is claimed** (§0 / §6) — P1/P3 are tested plumbing on a live endpoint + a safe privacy mapping; the live surface is single-link CSI with its existing caveats (no validated room-coordinate accuracy — `field_localize`). --- diff --git a/api-docs/adr/ADR-263-rtl8720f-2-4ghz-fmcw-radar-platform.md b/api-docs/adr/ADR-263-rtl8720f-2-4ghz-fmcw-radar-platform.md new file mode 100644 index 00000000..f6d984ec --- /dev/null +++ b/api-docs/adr/ADR-263-rtl8720f-2-4ghz-fmcw-radar-platform.md @@ -0,0 +1,171 @@ +# ADR-263: Adopt RTL8720F 2.4 GHz FMCW radar as an optional RuView sensing platform + +- **Status**: proposed +- **Date**: 2026-07-18 +- **Deciders**: ruv +- **Tags**: realtek, rtl8720f, ameba, fmcw, radar, cfr, csi, hardware +- **Relates to**: ADR-018, ADR-063, ADR-064, ADR-095, ADR-097, ADR-260, ADR-262 + +## Context + +Realtek's `RTL8720F-2.4G-Radar-Advantages_EN.pptx` describes an RTL8720F mode that shares the +2.4 GHz radio between Wi-Fi, Bluetooth, and an active FMCW radar. It offers two data products that +are useful to RuView: + +1. **CFR (Channel Frequency Report)**, described by Realtek as the same concept as Wi-Fi CSI. +2. **Near and far Range-FFT reports**, preserving near-field content while extending observation to + approximately 5–6 m. + +The proposed radio uses one transmit and one receive antenna, 20/40/70 MHz sweeps, configurable +8/16/32/64 microsecond chirp symbols, a maximum 2.56 ms FMCW packet, and a configurable frame +interval above 15 ms. The deck recommends 40 MHz outside Japan and 20 MHz in Japan. It also +describes EDCCA/CTS channel access, Wi-Fi/BT/radar time division, interference reporting, and +priority arbitration in the driver. + +This is not a drop-in replacement for ESP32 CSI: + +- it is **active monostatic FMCW**, while the ESP32 path observes Wi-Fi packet CSI; +- one Tx/one Rx has no angle-of-arrival or native multi-target separation; +- the stated 40 MHz range resolution is about 3.15 m, despite a finer 0.59 m Range-FFT report step; +- the presentation is a capability description, not an SDK contract. It contains no header names, + function signatures, callback ABI, binary layouts, toolchain version, licensing terms, or public + RTL8720F board package. + +Realtek's public Ameba RTOS repository is the base. Release v1.2.1 includes the CSI API and fixes a +CSI application-buffer semaphore issue, but does not expose the radar application surface. Open +upstream PR #1336 (2026-07-18 snapshot) adds RTL8720F project artifacts, `AT+RAD`, `AT+RADDBG`, and +the public configuration call `wifi_radar_config(struct rtw_radar_action_parm *)`. Its public +parameter struct confirms mode, channel, 70/40/20 MHz bandwidth selector, trigger period, and +enable/config actions. Report reception still crosses non-public/placeholder HAL symbols such as +`wifi_hal_radar_recv_data(frame_num, frame_type, data)`, so the report layout and buffer lifetime +remain vendor-gated. Therefore the integration stays split at that boundary. + +## Decision + +RuView will support RTL8720F radar as an **optional, capability-negotiated source**, without +replacing the ESP32 firmware or treating radar CFR as byte-compatible with ADR-018 CSI. + +The integration has three layers: + +1. **Realtek device firmware**: a small application built in the vendor-supported Ameba SDK calls + the radar API, owns coexistence configuration, and emits versioned reports. This code lives under + `firmware/rtl8720f-radar/` only after the redistributable SDK/API is available. +2. **Transport-neutral wire contract**: CFR and Range-FFT reports are framed independently from the + vendor ABI and sent over UDP, USB CDC, or UART. ADR-264 defines this boundary. +3. **Rust host adapter**: `wifi-densepose-hardware` parses reports from bytes and converts CFR into + the existing CSI-domain representation, while Range-FFT remains a radar modality and feeds the + RuField/RuView cross-modality bridge from ADR-260/262. + +The two report types remain semantically distinct: + +| RTL8720F output | RuView representation | Permitted use | +|---|---|---| +| CFR | `CsiFrame` through a Realtek calibration adapter | CSI feature extraction after validation | +| Range-FFT near/far | `RadarFrame` / RuField `mmwave_radar`-class event with a 2.4 GHz descriptor | range, motion, presence, fusion | +| Vendor AI presence probability | derived observation with model/version provenance | advisory input, never ground truth | +| Interference report | quality/provenance metadata | reject, down-weight, or mark contaminated frames | + +The modality registry should eventually distinguish `fmcw_radar_2_4ghz` from `mmwave_radar`; until +that RuField schema revision is accepted, the adapter must attach `carrier_hz = 2.4e9` and must not +claim millimetre-wave provenance. + +## Delivery phases and gates + +### P0 — Vendor enablement + +Obtain the PR #1336-or-newer RTL8720F SDK package, radar API headers/libraries, a supported evaluation +board, flashing/debug instructions, report definitions, and written redistribution terms. + +**Gate:** compile and run Realtek's unmodified radar example and capture CFR plus near/far +Range-FFT output. Until this passes, device firmware is `VENDOR_BLOCKED`, not implemented. + +### P1 — Host-first contract + +Implement ADR-264 types, parsers, fixtures, fuzz tests, and replay support without linking vendor +code. Use the Rust `Rtl8720fSimulator` as the only pre-hardware live source. It emits deterministic +CFR, near/far Range-FFT, interference, and capabilities frames through the same ADR-264 encoder and +parser used by hardware. Every simulated frame sets `RadarFlags::SYNTHETIC`; simulation results are +never reported as device measurements. + +**Gate:** malformed inputs never panic; encode/decode round trips; unknown versions and report +types fail closed. + +### P2 — RTL8720F firmware adapter + +Wrap only the minimum vendor API surface: initialization, profile configuration, start/stop, +callback acquisition, interference status, and report serialization. Keep vendor types out of the +wire protocol. + +**Gate:** 30-minute simultaneous Wi-Fi telemetry and radar capture with no watchdog reset, bounded +loss, monotonic sequence numbers, and explicit coexistence/interference statistics. + +### P3 — Calibration and signal validation + +Calibrate CFR phase/amplitude, Range-FFT bin spacing, static leakage, and clock drift. Compare +reported range against measured targets at multiple distances and bandwidths. + +**Gate:** publish measured error distributions. Do not infer accuracy from report-bin spacing and +do not advertise multi-person pose or vital signs from the vendor deck. + +### P4 — Fusion and productization + +Feed calibrated CFR through the CSI path and Range-FFT through RuField, retaining source, mode, +bandwidth, calibration, firmware, and interference provenance. + +**Gate:** ablation shows whether the radar stream improves a named RuView metric over ESP32 CSI +alone. If it does not, ship it only as an independent presence/range sensor. + +## Consequences + +### Positive + +- One low-cost radio can provide active radar and CSI-like CFR while retaining Wi-Fi connectivity. +- Range-FFT adds an independent physical measurement for presence/range fusion. +- The vendor SDK is isolated from the Rust sensing core and from the stable on-wire contract. +- Capability negotiation permits future Realtek parts without another application-level fork. + +### Negative + +- The first implementation is blocked on access to the actual RTL8720F radar SDK/API and hardware. +- Active 2.4 GHz transmission changes coexistence, privacy, power, and regional compliance concerns. +- 1T1R and limited sweep bandwidth cannot provide the spatial resolution of multi-antenna mmWave. +- A second embedded toolchain and firmware release process must be maintained. + +### Neutral + +- ESP32 remains the default CSI node. +- Existing consumers receive normalized frames and do not link against Realtek code. +- Vendor AI output is optional metadata; RuView retains responsibility for its own validation. + +## Rejected alternatives + +1. **Map Range-FFT directly to `CsiFrame`.** Rejected because range bins and channel-frequency + samples have different axes and physical meaning. +2. **Link the Realtek SDK into the Rust server.** Rejected because it couples host builds to a + proprietary embedded ABI and toolchain. +3. **Wait to define any interface until hardware arrives.** Rejected because the host protocol, + parser safety, replay, and provenance can be developed and reviewed independently. +4. **Replace ESP32 nodes.** Rejected because the modes are complementary and availability differs. + +## Open vendor questions + +- Exact RTL8720F part/board identifier and production availability. +- SDK repository/tag, compiler, RTOS, binary blobs, license, and redistribution permissions. +- Radar initialization/configuration/callback API signatures and threading/ISR constraints. +- CFR and near/far Range-FFT element type, complex ordering, scaling, endianness, and timestamps. +- Whether CFR is calibrated complex data and whether phase remains coherent across frames. +- Maximum report rates, buffer ownership, DMA/cache constraints, and Wi-Fi throughput impact. +- Region/channel enforcement and whether 70 MHz operation is allowed by the supplied firmware. +- Secure boot, signed OTA, unique device identity, and firmware attestation support. + +## Sources + +- Realtek Semiconductor, `RTL8720F-2.4G-Radar-Advantages_EN.pptx`, slides 3 and 10–19, + supplied 2026-07-18. This is product material, not measured RuView validation. +- [Ameba-AIoT/ameba-rtos releases](https://github.com/Ameba-AIoT/ameba-rtos/releases), reviewed + 2026-07-18; v1.2.1 is the current QC release and includes a CSI buffer-semaphore fix. +- [Ameba-AIoT/ameba-rtos PR #1336](https://github.com/Ameba-AIoT/ameba-rtos/pull/1336), reviewed + 2026-07-18; exposes RTL8720F build assets, `wifi_radar_config`, and radar AT commands while report + internals remain in binary/private layers. +- ADR-063 (mmWave sensor fusion), ADR-095/097 (source normalization), and ADR-260/262 (RuField + multimodal event model and live bridge). diff --git a/api-docs/adr/ADR-263-ruview-npm-harness-deep-review.md b/api-docs/adr/ADR-263-ruview-npm-harness-deep-review.md new file mode 100644 index 00000000..83d90480 --- /dev/null +++ b/api-docs/adr/ADR-263-ruview-npm-harness-deep-review.md @@ -0,0 +1,191 @@ +# ADR-263: `@ruvnet/ruview` npm Harness — Deep Review + Optimization Strategy + +| Field | Value | +|-------|-------| +| **Status** | Accepted — **implemented** (O1–O9, `@ruvnet/ruview@0.2.0`): fail-closed `claim-check`, async MCP dispatch (ping answered mid-`verify`, pinned by e2e test), zero-dependency install, bounded output tails, argv-passed monitor port, package.json-sourced version, prepack skill sync, memoized `which()`, underscore-canonical tools with dotted aliases, word-boundary guardrail matching. 30/30 tests (MEASURED, `node --test test/*.test.mjs`); CI gate in ADR-265's `npm-packages.yml` | +| **Date** | 2026-07-02 | +| **Deciders** | ruv | +| **Codename** | **RUVIEW-NPM-REVIEW-1** | +| **Supersedes / amends** | none (records review of the ADR-182 P1+P2 artifact; feeds ADR-265 distribution strategy) | + +## Context + +ADR-182 minted and published **`@ruvnet/ruview@0.1.0`** (`harness/ruview/`) — the +`npx ruview` operator harness: a dependency-free ESM CLI + minimal MCP stdio server +exposing six `ruview.*` tools (onboard / claim_check / verify / node_monitor / +calibrate / node_flash), five skill playbooks, and the executable +MEASURED-vs-CLAIMED guardrail (`src/guardrails.js`). The package is live on npm +(0.1.0, 49.5 kB unpacked / 21 files — MEASURED, `npm view @ruvnet/ruview` + +`npm pack --dry-run`) and is the recommended MCP registration path +(`npx -y @ruvnet/ruview mcp start` in the bundled `.claude/settings.json`). + +This ADR is the first dedicated deep review of that npm artifact: correctness, +fail-open/fail-closed posture, performance (cold start + request handling), +packaging hygiene, and security of the subprocess surface. All 17 bundled tests +pass on Node 22 (MEASURED, `node --test test/*.test.mjs`, 17/17, ~108 ms). + +## Findings + +Severity reflects impact on the package's stated contract: *fail-closed operator +tools + an honesty guardrail that must never fail open*. + +### F1 (HIGH, fail-open): `claim-check` passes silently on empty input + +`bin/cli.js` `claim-check` with **neither `--text` nor `--file`** sends +`text: undefined` → `claimCheck(String(args.text ?? ''))` → `''` → `ok: true`, +**exit 0**. A CI hook wired as `npx ruview claim-check --text "$BODY"` where +`$BODY` expands empty therefore reports PASS. This is the single tool whose whole +purpose is to fail closed; empty input must be an error, not a pass. +Reproducer: `node bin/cli.js claim-check` → `{"ok": true}`, exit 0. + +### F2 (HIGH, head-of-line blocking): MCP server is fully synchronous + +`src/mcp-server.js` dispatches `tools/call` inside the readline `line` handler, +and every heavyweight handler in `src/tools.js` uses **`spawnSync`** +(`ruview.verify` up to 180 s, `ruview.calibrate` up to 300–600 s, +`ruview.node_monitor` up to `seconds+10`). While one call runs, the event loop is +blocked: `ping`, `tools/list`, and concurrent `tools/call` requests are not even +read from stdin. Hosts that health-check with `ping` during a long `calibrate` +will conclude the server is dead and kill it mid-run. + +### F3 (MEDIUM, cold start): optionalDependencies triple the `npx` install for a path that never uses them + +`package.json` declares `optionalDependencies` on `@metaharness/kernel` and +`@metaharness/host-claude-code`. npm installs optional deps **by default**, so +every cold `npx -y @ruvnet/ruview mcp start` fetches 3 extra packages (kernel + +host + transitive `@ruvector/emergent-time`). MEASURED (npm 10.9.7, this +container): default install = **4 packages, 620 kB, 71 files**; with +`--omit=optional` = **1 package, 172 kB, 22 files**. The operator-tool and MCP +paths never import these — only `doctor`/`install` do, and both already +dynamic-import inside `try/catch` and degrade gracefully when absent +(`kernel/host: not installed (ok…)`). The optional deps buy nothing on the hot +path and cost 3 registry round-trips + ~450 kB on every cold start. + +### F4 (MEDIUM, silent truncation): `spawnSync` default `maxBuffer` (1 MiB) + +`run()` in `src/tools.js` never sets `maxBuffer`. `cargo run -p +wifi-densepose-cli` (the `calibrate` fallback path) and a chatty `verify.py` can +exceed 1 MiB of stdout, at which point the child is killed with `ENOBUFS` and the +tool reports a spawn error that looks like a proof/calibration failure. The +handlers only ever consume the last 8 kB/1.5 kB; buffering should be bounded but +generous (e.g. `maxBuffer: 16 MiB`) or streamed with a tail ring. + +### F5 (MEDIUM, injection surface): `node_monitor` interpolates the port into Python source + +The handler builds a `python -c` script by string interpolation: +`` `ser=serial.Serial(${JSON.stringify(port)},115200,…)` `` and +`` `while time.time()-t<${dur}:` ``. `JSON.stringify` produces a *JavaScript* +string literal; Python string-literal semantics differ at the edges (`\uXXXX` is +shared, but e.g. JS emits raw U+2028/U+2029 unescaped pre-ES2019 rules aside, and +any future non-JSON-safe field added the same way would be executable). `port` +arrives from the MCP caller (an agent), so this is an agent-controlled string +concatenated into an interpreter invocation. `dur` is `Number()`-guarded; `port` +should be passed out-of-band (`sys.argv`/env), never spliced into source. + +### F6 (LOW, drift): server version hardcoded + +`SERVER_INFO = { name: 'ruview', version: '0.1.0' }` in `src/mcp-server.js` +duplicates `package.json.version` (the CLI's `--version` already reads +package.json at runtime). First release bump will drift the MCP handshake +version. + +### F7 (LOW, duplication): every skill ships twice + +`skills/*.md` and `.claude/skills/*/SKILL.md` are byte-identical (same sha256 in +`.harness/manifest.json`). ~8 kB of the 49.5 kB unpacked payload is duplicate +content, and — worse than size — two copies must be kept in sync by hand. + +### F8 (LOW, perf + portability): `which()` is uncached and shells out + +`which()` runs up to twice per tool call (`python` then `python3`), each a +blocking `spawnSync`; the POSIX branch spawns a shell (`shell: true`). Results +are stable for the process lifetime and should be memoized; the lookup can be +done dep-free with a PATH scan instead of a shell. + +### F9 (LOW, interop): dot-named tools + minimal protocol surface + +Tool names (`ruview.onboard`, `ruview.claim_check`, …) contain dots. MCP itself +does not restrict names, but downstream host APIs commonly enforce +`^[a-zA-Z0-9_-]{1,64}$` for tool names; hosts must then sanitize or reject. +The server also answers `resources/list` / `prompts/list` with `-32601` (it does +not advertise those capabilities, so this is spec-legal, but empty-list stubs are +cheaper than every host's error path). Protocol version is pinned to +`2024-11-05` with no negotiation fallback. None of this breaks Claude Code today; +it narrows portability, which is the harness's whole pitch (9 hosts, ADR-182). + +### F10 (LOW, CI gap): the published package has zero CI + +No workflow under `.github/workflows/` runs `harness/ruview` tests (checked: +no workflow references `harness/ruview`, `ruview-mcp`, or `ruview-cli`), and +`ci.yml` pins `NODE_VERSION: '18'` while the package declares +`engines.node >= 20`. Note also `node --test test/` (directory form) fails on +Node 22 while the documented glob form passes — CI should pin the working +invocation. Consolidated CI/publish strategy is ADR-265. + +### F11 (MEDIUM, guardrail precision): `METRIC_TERMS` substring matching false-positives on ordinary prose + +Found by dogfooding this review: `claimCheck` matches metric terms with +`lower.includes(t)`, so the two-character terms `'map'` and `'f1'` fire inside +ordinary words and labels — "source **map**s", "the **map**s can never +resolve", finding IDs like "**F1** (HIGH…)". MEASURED reproducer: running +`npx ruview claim-check --file` over this ADR and ADR-264 yields 4 and 16 +medium findings respectively, the majority of which are `map`/`F1` +false positives on lines carrying no accuracy claim. A guardrail that cries +wolf trains people to ignore it — precision is part of its fail-closed +contract. Short/ambiguous terms need word-boundary matching (`\bmap\b`, +`\bf1\b`, likewise `auc`, `iou`), and section-heading label patterns +(`F\d+`, `O\d+`) should not count as metric mentions. + +## Decision + +Adopt the following optimization strategy, in priority order. Each item is +independently shippable; F-numbers map to findings. + +- **O1 (F1):** `claim-check` with no `--text`/`--file` (or empty text after read) + exits 2 with a usage error. Add a regression test pinning exit ≠ 0. +- **O2 (F2):** make the MCP dispatch async: convert `run()`/`which()` to + promise-based `spawn`, make `tools/call` handlers `async`, and keep reading + stdin while calls run (respond to `ping`/`tools/list` concurrently; serialize + only same-tool hardware operations). Acceptance: `ping` round-trips < 50 ms + while a synthetic 30 s `calibrate` is in flight. +- **O3 (F3):** drop the two `optionalDependencies`; `doctor`/`install` already + degrade and should print the exact `npm i @metaharness/kernel + @metaharness/host-claude-code` hint on the miss path. Acceptance: cold + `npm i @ruvnet/ruview` installs exactly 1 package (MEASURED baseline above). +- **O4 (F4):** set `maxBuffer: 16 * 1024 * 1024` in `run()` (or stream + tail). +- **O5 (F5):** pass `port` to the monitor script via `sys.argv` + (`python -c script -- `), never by source interpolation. +- **O6 (F6):** read the MCP `serverInfo.version` from `package.json` once at + startup (same pattern the CLI already uses). +- **O7 (F7):** make `skills/*.md` the single source and generate + `.claude/skills/*/SKILL.md` in a `prepack` script (or vice versa); manifest + hashes then pin one canonical set. +- **O8 (F8, F9):** memoize `which()`; add underscore aliases for the dot-named + tools (accept both in `tools/call`, advertise the underscore form) and add + empty `resources/list` / `prompts/list` stubs. +- **O9 (F11):** switch `METRIC_TERMS` matching to word-boundary regexes for + short terms (`map`, `f1`, `auc`, `iou`) and skip label tokens matching + `\b[FO]\d+\b`. Acceptance: `claim-check --file` over ADR-263/264/265 reports + only the genuinely tagged-or-taggable percentage lines, and the existing 17 + guardrail tests still pass plus new false-positive pins ("source maps", + "F1 (HIGH)" → no finding). + +Non-goals: no new runtime dependencies (the zero-dep MCP server is a feature, +not an accident — keep it), no build step, no change to the fail-closed tool +contracts. + +## Consequences + +- The honesty guardrail becomes fail-closed end-to-end (its current empty-input + pass is the exact failure mode the guardrail exists to prevent). +- `npx` cold start drops ~450 kB / 3 packages (MEASURED baseline in F3) with no + feature loss; `doctor` output already communicates the optional-dep story. +- Long-running `verify`/`calibrate` no longer starve the MCP channel — the + harness survives host health checks during real calibration runs. +- Two-copy skill drift becomes impossible at pack time. +- Costs: async conversion touches every handler signature in `src/tools.js` + (mechanical, ~6 handlers); alias tools add a small compatibility table. +- Verification for the implementing PR: bundled tests extended for O1/O2/O5 + (target ≥ 20 tests), `npm pack --dry-run` file-count asserted, and the F3 + install measurement re-run and quoted MEASURED in the PR body — which must + itself pass `npx ruview claim-check`. diff --git a/api-docs/adr/ADR-264-rtl8720f-radar-wire-protocol.md b/api-docs/adr/ADR-264-rtl8720f-radar-wire-protocol.md new file mode 100644 index 00000000..99a82cca --- /dev/null +++ b/api-docs/adr/ADR-264-rtl8720f-radar-wire-protocol.md @@ -0,0 +1,148 @@ +# ADR-264: Versioned wire protocol for RTL8720F CFR and Range-FFT reports + +- **Status**: proposed +- **Date**: 2026-07-18 +- **Deciders**: ruv +- **Tags**: realtek, rtl8720f, protocol, cfr, range-fft, udp, serial +- **Depends on**: ADR-263 +- **Relates to**: ADR-018, ADR-095, ADR-097, ADR-099, ADR-260 + +## Context + +ADR-263 adopts RTL8720F radar behind an anti-corruption boundary. The Realtek presentation names +CFR, near Range-FFT, far Range-FFT, and interference reports, but does not specify their binary ABI. +RuView needs a stable, testable contract that can be implemented before the vendor SDK arrives and +that will not expose vendor structs, pointer layouts, padding, or callback lifetime rules over the +network. + +ADR-018 already defines ESP32 CSI framing. Reusing its magic or pretending that Realtek radar is an +ESP32 packet would make source detection ambiguous and erase radar-specific calibration metadata. + +## Decision + +Define a new little-endian `RtlRadarFrameV1` envelope with its own magic and explicit payload type. +This is a RuView protocol, not a claim about Realtek's native memory layout. + +### Envelope + +All integer fields are little-endian. Floating-point payloads use IEEE-754 binary32. No C struct is +sent by `memcpy`; firmware serializes each field explicitly. + +| Offset | Size | Field | Meaning | +|---:|---:|---|---| +| 0 | 4 | magic | ASCII `RTR1` (`0x31525452`) | +| 4 | 1 | version | `1` | +| 5 | 1 | report_type | 1 CFR, 2 range-near, 3 range-far, 4 interference, 5 capabilities | +| 6 | 2 | header_len | complete header size, initially 56 | +| 8 | 4 | frame_len | header + payload + CRC | +| 12 | 4 | sequence | wraps modulo 2^32 | +| 16 | 8 | timestamp_us | monotonic device time at acquisition | +| 24 | 8 | device_id | stable pseudonymous identifier, not a MAC address | +| 32 | 4 | center_freq_khz | RF centre frequency | +| 36 | 2 | bandwidth_mhz | 20, 40, or 70 | +| 38 | 2 | flags | calibration/interference/saturation/time-sync flags | +| 40 | 2 | element_count | complex samples or range bins | +| 42 | 1 | element_format | 0 bytes/TLV, 1 complex-i16, 2 complex-f32, 3 power-u16, 4 power-f32 | +| 43 | 1 | antenna_count | expected to be 1 for the deck's 1T1R configuration | +| 44 | 4 | scale | quantized-to-physical multiplier; `1.0` for float payloads | +| 48 | 4 | bin_spacing | Hz for CFR, metres for Range-FFT | +| 52 | 4 | calibration_id | device calibration revision/hash prefix | +| 56 | variable | payload | determined by type, count, and format | +| final-4 | 4 | crc32 | IEEE CRC-32 over header and payload | + +If vendor evidence shows that 56 bytes is too costly, a later protocol version may introduce a +compact header. V1 favors auditable provenance over premature byte savings. + +### Payload semantics + +- **CFR** contains ordered complex channel-frequency samples. The adapter must know the frequency + origin/order and must not fabricate missing phase. Uncalibrated frames carry the uncalibrated flag + and cannot enter phase-sensitive processing. +- **Range-near/range-far** contains ordered range bins. Near and far are separate report types so + filtering and leakage behavior are never hidden from consumers. +- **Interference** contains a versioned TLV set for channel-busy, detected-during-chirp, estimated + interference power, and packet jitter. Unknown TLVs are skipped by length. +- **Capabilities** is emitted at boot and on request. It declares supported report types, bandwidths, + chirp lengths, maximum elements/report, maximum frame rate, firmware version, and SDK identifier. + +### Transport + +The identical envelope is supported over: + +- UDP datagrams for normal RuView ingestion; +- USB CDC or UART with COBS framing and a zero-byte delimiter; +- file replay as a length-prefixed sequence of envelopes. + +One envelope must fit one UDP datagram. Fragmentation is not part of V1; firmware rejects a profile +whose maximum report exceeds the configured MTU and reports the required size through capabilities. + +### Parser and trust rules + +The host parser: + +1. validates magic, version, lengths, enum values, element count/format multiplication, and CRC + before allocating or decoding the payload; +2. caps frames at 64 KiB and elements at a configured hardware maximum; +3. rejects non-finite float metadata/payload values; +4. tracks sequence gaps and timestamp regressions per device; +5. preserves unknown flags but never interprets them as trusted; +6. attaches transport source, firmware/SDK version, calibration ID, and interference state to + provenance; +7. labels fixture/generated frames as synthetic. + +No vendor-provided presence probability bypasses RuView privacy, provenance, or quality gates. + +## Consequences + +### Positive + +- Firmware, transport, parser, replay, and fusion can evolve independently. +- Fuzzing and golden fixtures require no Realtek SDK or board. +- CFR and Range-FFT retain correct axes and calibration provenance. +- A boot-time capabilities frame makes SDK/API drift observable. + +### Negative + +- Serialization adds CPU and bandwidth overhead compared with dumping a vendor buffer. +- V1 fields may need revision after the actual API and report limits are disclosed. +- UDP provides integrity/error detection, not authenticity or confidentiality. + +### Neutral + +- Authentication can be layered with ADR-032 device identity or a signed RuField receipt without + changing report semantics. +- ESP32 ADR-018 framing remains unchanged. + +## Implementation plan + +1. Add `rtl8720f` types/parser module to `wifi-densepose-hardware` behind no vendor dependency. +2. Add golden CFR, near/far Range-FFT, interference, and capabilities fixtures. +3. Add property/fuzz tests for length arithmetic, enum handling, CRC, and float validation. +4. Add a replay CLI that prints normalized metadata without running inference. +5. Once SDK access exists, implement the embedded serializer and verify captured frames against the + host golden decoder. +6. Revise this proposed ADR with measured element counts, rates, and API names before acceptance. + +Host-side steps 1–3 are implemented in `wifi-densepose-hardware::rtl8720f`: typed report and +element enums, semantic type/format validation, bounded length arithmetic, CRC verification, +finite-float checks, encode/decode round trips, corruption/truncation tests, and deterministic +arbitrary-input panic checks. Cross-language vectors remain blocked on the vendor SDK callback ABI. +Bit 15 of `flags` is reserved by RuView as `SYNTHETIC`; the Rust simulator always sets it and real +firmware must never set it. The simulator is deterministic by seed and exercises the production +encoder/parser rather than a parallel mock representation. + +## Acceptance criteria + +- Rust encode/decode round-trip for every report type. +- Cross-language golden vector produced by the RTL8720F firmware. +- Zero parser panics over the fuzz corpus and arbitrary byte input. +- Detection of single-bit corruption, truncation, count overflow, timestamp regression, and gaps. +- Captured CFR frequency order and Range-FFT bin spacing verified against vendor documentation and a + measured target. + +## Sources + +- Realtek Semiconductor, `RTL8720F-2.4G-Radar-Advantages_EN.pptx`, slides 11–19, supplied + 2026-07-18. +- ADR-018 (ESP32 framing), ADR-095/097 (hardware normalization), ADR-260 (multimodal event model), + and ADR-263 (platform decision). diff --git a/api-docs/adr/ADR-264-rvagent-mcp-and-cli-npm-deep-review.md b/api-docs/adr/ADR-264-rvagent-mcp-and-cli-npm-deep-review.md new file mode 100644 index 00000000..f29eaf94 --- /dev/null +++ b/api-docs/adr/ADR-264-rvagent-mcp-and-cli-npm-deep-review.md @@ -0,0 +1,169 @@ +# ADR-264: `@ruvnet/rvagent` MCP Server + `@ruv/ruview-cli` — Deep Review + Optimization Strategy + +| Field | Value | +|-------|-------| +| **Status** | Accepted — **implemented** (O1–O9, `@ruvnet/rvagent@0.2.0`): `exports` fixed (types-first, no phantom `.cjs`), map-free tarball (127,704 B unpacked / 46 files / 0 maps — MEASURED, `npm pack --dry-run`, from 188 kB), Streamable HTTP **wired** behind `RVAGENT_HTTP_PORT` with per-session transports + 1 MiB body cap + port-aware origin gate, underscore tool names with dotted router aliases, single Zod validation gate with generated JSON Schemas, fd-leak fixed + persisted job records + bounded log tails, probing `detectCogBinary`, package.json-sourced version, `ruview-cli` bin renamed. 99/99 jest tests (MEASURED); both transports smoke-tested live | +| **Date** | 2026-07-02 | +| **Deciders** | ruv | +| **Codename** | **RUVIEW-NPM-REVIEW-2** | +| **Supersedes / amends** | none (reviews the ADR-104/ADR-124 artifacts; feeds ADR-265 distribution strategy) | + +## Context + +Two TypeScript npm packages expose RuView sensing to agents and shells: + +- **`@ruvnet/rvagent@0.1.0`** (`tools/ruview-mcp/`) — SENSE-BRIDGE, the MCP + server over the sensing-server HTTP API + cog binaries: 12 tools + (csi/pose/count/registry/train/job + ADR-124 BFLD/presence/vitals). Published + (188 kB unpacked — MEASURED, `npm view @ruvnet/rvagent`). Deps: + `@modelcontextprotocol/sdk` + `zod`. +- **`@ruv/ruview-cli@0.0.1`** (`tools/ruview-cli/`) — `private: true` yargs CLI + mirroring the same capabilities; intentionally duplicates `http.ts`/`cog.ts`/ + `config.ts` (~150 lines) to stay standalone. + +This ADR records a deep review of both: packaging correctness (verified against +the **published** tarball, not just the source tree), protocol/interop, resource +lifecycle, and the honesty of the package's own self-description — the same +MEASURED-vs-CLAIMED bar the project applies to accuracy numbers. + +## Findings + +### F1 (HIGH, broken export): `require` condition points at a file that does not exist + +`package.json` `exports["."].require = "./dist/index.cjs"`, but the build is +plain `tsc` (ESM only) and **the published 0.1.0 tarball contains no +`index.cjs`** (verified by listing the registry tarball). Any CJS consumer doing +`require('@ruvnet/rvagent')` resolves to a nonexistent file → +`ERR_MODULE_NOT_FOUND`. Additionally the `types` condition is listed **after** +`import`/`require`; TypeScript requires `types` first or it may be ignored under +`moduleResolution: bundler/node16`. + +### F2 (MEDIUM, tarball bloat): a third of the published package is dead source maps + +The 0.1.0 tarball ships **44 `.map` files = 62,698 B** against 78,209 B of +actual `.js` (MEASURED, extracted registry tarball). `src/` is not published, so +every `sourceMappingURL` points at `../src/*.ts` that consumers do not have — +the maps can never resolve. Also `files` lists `CHANGELOG.md`, which does not +exist in `tools/ruview-mcp/` (npm silently skips it), so the advertised file set +is partly fictional. + +### F3 (MEDIUM, honesty): the package description claims a transport it does not start + +The description reads "**dual-transport MCP server (stdio + Streamable HTTP)**", +but `main()` in `src/index.ts` wires **stdio only**. `http-transport.ts` is a +complete, tested scaffold that nothing imports at runtime — there is no flag, +env var, or subcommand that starts it. By this project's own rule this is a +CLAIMED capability presented as shipped. Either wire it (`--http` / +`RVAGENT_HTTP_PORT` gate) or de-claim the description until it is. + +### F4 (MEDIUM, interop + inconsistency): two tool-naming conventions, one of them dot-based + +Six tools use `ruview_snake_case`; six (ADR-124 additions) use +`ruview.dotted.names`. Same interop caveat as ADR-263 F9 (host tool-name +regexes commonly `^[a-zA-Z0-9_-]{1,64}$`), plus the split convention makes the +tool surface look like two products. Standardize on underscores and accept the +dotted forms as aliases for one deprecation cycle. + +### F5 (MEDIUM, double work + drift): every tool input is validated twice from two hand-maintained schemas + +`CallToolRequestSchema` handler runs `TOOL_INPUT_SCHEMAS[name].safeParse(args)`, +then each tool handler runs its own `schema.parse(args)` again — two full Zod +passes per call. Separately, the `inputSchema` JSON advertised via `tools/list` +is **hand-written** and duplicates the Zod schema field-by-field (defaults, +min/max, descriptions) — schema drift between what is advertised and what is +enforced is a matter of time. Parse once at the gate, pass the typed result to +handlers, and generate the advertised JSON Schema from the Zod source +(`zod-to-json-schema` at build time, or Zod 4's native `z.toJSONSchema` when the +SDK's peer range allows). + +### F6 (MEDIUM, resource lifecycle): `train_count` leaks 2 fds per job; job registry is process-local + +`trainCount` opens `logFdOut`/`logFdErr` with `openSync` and never closes them +in the parent — the spawned cargo child inherits duplicates, but the parent's +descriptors stay open for the MCP server's lifetime: 2 leaked fds per training +job. `jobRegistry` is an in-memory `Map`, so `ruview_job_status` after a server +restart reports "not found" for a training run that is still burning GPU (the +source comments acknowledge this; the fix — persist `~/.ruview/jobs/.json`, +already the documented layout — is small). Also `jobStatus` re-`import`s +`node:fs` on every poll and reads the entire log to return 20 lines. + +### F7 (MEDIUM, security/robustness of the HTTP scaffold): unbounded body + one shared session transport + +`http-transport.ts` buffers the request body with no size cap (memory DoS the +moment it is wired to a socket), reuses a **single** +`StreamableHTTPServerTransport` with `sessionIdGenerator` for all clients (the +SDK's stateful mode expects one transport per session — a second client's +`initialize` collides), and the Origin allowlist is exact-match +(`http://localhost` will not match a real browser origin `http://localhost:5173`). +Must be fixed **before** F3 wires it in; bearer-token + 127.0.0.1 defaults are +already right. + +### F8 (LOW, dead/misleading code): `detectCogBinary` always returns the bare name + +It builds a 4-candidate appliance-path array and then returns +`candidates[candidates.length - 1]` — i.e. always `name` — without checking +existence. The candidates are dead weight that reads as if path detection +happens. Either probe with `existsSync` or delete the array. + +### F9 (LOW, drift + hygiene): hardcoded versions, unused/mismatched devDeps, bin-name collision + +`PACKAGE_VERSION = "0.1.0"` (index.ts) duplicates package.json; +`@types/express` is unused (`http-transport` uses `node:http`); `@types/jest@30` +against `jest@29`; `ruview-cli` hardcodes `.version("0.0.1")`. And +`@ruv/ruview-cli` claims the **`ruview`** bin name, which collides with +`@ruvnet/ruview`'s bin (ADR-182) if both are ever installed globally — +ADR-263/265 give the `ruview` name to the harness; the CLI must rename or fold. + +## Decision + +- **O1 (F1):** fix `exports`: drop the `require` condition (ESM-only is fine for + a bin-first package) or add a real CJS build; put `types` first. Add a CI + smoke test that does `npm pack` + `node -e "import('')"`. +- **O2 (F2):** publish without maps: `declarationMap: false`, `sourceMap: false` + in a `tsconfig.build.json` used by `prepack` (or add `!dist/**/*.map` to + `files`). Remove the phantom `CHANGELOG.md` entry or create the file. + Acceptance: unpacked size ≤ ~125 kB (from 188 kB — MEASURED, `npm pack --dry-run`). +- **O3 (F3, F7):** wire the HTTP transport behind an explicit opt-in + (`RVAGENT_HTTP_PORT` or `--http`), after F7 fixes: per-session transport map + keyed by `mcp-session-id`, 1 MiB body cap, origin matching that honors ports + (compare `URL.origin` prefixes or document exact origins). Until then, change + the description to "stdio MCP server (Streamable HTTP scaffold, unwired)". +- **O4 (F4):** rename dotted tools to underscore (`ruview_bfld_last_scan`, …), + keep dotted aliases in the call router for one release, note it in the README. +- **O5 (F5):** single validation gate: the registry maps name → Zod schema → + typed handler; advertised `inputSchema` generated from Zod at build time. +- **O6 (F6):** close parent fds after spawn (`closeSync` post-`spawn` — the + child holds its own copies), persist job records to + `/.json`, and read log tails with a bounded read. +- **O7 (F8):** make `detectCogBinary` actually probe (`existsSync` over the + candidates) — it is the entire reason the function exists. +- **O8 (F9):** single-source versions from package.json; drop `@types/express`; + align `@types/jest` with jest 29 (or move to `node:test` like the harness and + drop the jest toolchain entirely — it is the heaviest devDep in both + packages). +- **O9 (F9, scope):** fold `@ruv/ruview-cli` into `rvagent` as a second bin + (`rvagent-cli`) sharing `http/cog/config`, or keep it private-forever and say + so in its README. Its `ruview` bin name is surrendered to `@ruvnet/ruview` + either way. + +## Consequences + +- CJS consumers stop hitting a guaranteed-broken export path (F1 is the only + finding that fails for every consumer of that entry point deterministically). +- The published artifact shrinks ~33% (MEASURED, F2 tarball listing: 62,698 B + of maps in a 188 kB unpacked payload) and stops advertising files/transports + it does not contain — the package description itself passes the project's + claim-check bar. +- One schema source ends advertised-vs-enforced drift and halves per-call + validation cost; naming unification makes the 12-tool surface read as one + product and survive strict host tool-name validation. +- Long-lived MCP servers stop accumulating fds during training campaigns, and + job polling survives restarts. +- Costs: the alias cycle (O4) briefly doubles the advertised tool count unless + aliases are router-only (recommended: router-only, advertise underscore names + exclusively); folding the CLI (O9) retires a package name already in use in + scripts, so it needs a deprecation note. +- Verification for the implementing PR: `npm pack --dry-run` asserted file list + (no `.map`, no phantom entries), pack-size budget in CI (ADR-265), jest/`node + --test` suite green, and a tarball-install smoke test for both `import` and + the `rvagent` bin. diff --git a/api-docs/adr/ADR-265-ruview-npm-distribution-strategy.md b/api-docs/adr/ADR-265-ruview-npm-distribution-strategy.md new file mode 100644 index 00000000..945626be --- /dev/null +++ b/api-docs/adr/ADR-265-ruview-npm-distribution-strategy.md @@ -0,0 +1,124 @@ +# ADR-265: RuView npm Distribution Strategy — CI Gate, Provenance, Version Single-Sourcing, Namespace + +| Field | Value | +|-------|-------| +| **Status** | Accepted — **D1–D4 implemented**: `.github/workflows/npm-packages.yml` (matrix gate: tests, version-literal grep, pack-content/size gate, tarball-install smoke test, README claim-check), `.github/workflows/ruview-npm-release.yml` (publish-from-CI with `npm publish --provenance`), version single-sourcing (all three packages read package.json), `ruview` bin owned by `@ruvnet/ruview` (`@ruv/ruview-cli` bin renamed `ruview-cli`), `ci.yml` NODE_VERSION 18→20. D5 (no workspace) stands as recorded | +| **Date** | 2026-07-02 | +| **Deciders** | ruv | +| **Codename** | **RUVIEW-NPM-DIST** | +| **Supersedes / amends** | none (cross-cutting layer above ADR-263 and ADR-264; complements ADR-182 P3/P4) | + +## Context + +The monorepo now ships (or stages) **three Node packages** with no shared +distribution engineering: + +| Package | Dir | Published | Bin(s) | Tests in CI | +|---------|-----|-----------|--------|-------------| +| `@ruvnet/ruview` | `harness/ruview/` | 0.1.0 (live) | `ruview` | **none** | +| `@ruvnet/rvagent` | `tools/ruview-mcp/` | 0.1.0 (live) | `rvagent`, `ruview-mcp` | **none** | +| `@ruv/ruview-cli` | `tools/ruview-cli/` | private | `ruview` (collides) | **none** | + +Cross-cutting facts established during the ADR-263/264 reviews: + +- **Zero CI coverage.** No workflow under `.github/workflows/` references any of + the three directories. Two of the packages are *live on the registry* and were + published from a laptop state CI never saw. Meanwhile the Rust side has a + 1,031+-test gate and a witness-bundle culture (ADR-028) — the npm surface is + the only shipped artifact class with no verification gate at all. +- **`ci.yml` pins `NODE_VERSION: '18'`** while all three packages declare + `engines.node >= 20`. +- **Version triplication.** Each package hardcodes its version in source at + least once beyond package.json (harness `SERVER_INFO`, rvagent + `PACKAGE_VERSION`, cli `.version("0.0.1")`). +- **Bin-name collision.** Two packages claim the `ruview` bin. +- **No provenance.** Neither published package carries npm provenance + attestations, in a project whose differentiator is signed, reproducible + evidence (ADR-028 witness bundles, ADR-182 P4 ed25519/SLSA design). +- **No pack-content gate.** ADR-264 F1/F2 (broken `require` target, 33% dead map weight — MEASURED, tarball listing — and a phantom + `CHANGELOG.md` in `files`) are exactly the defect class an + `npm pack --dry-run` assertion catches in seconds. + +## Decision + +Adopt one distribution layer for all Node packages. Per-package code fixes live +in ADR-263/264; this ADR fixes the machinery around them. + +### D1 — One `npm-packages.yml` CI workflow (the gate) + +Matrix over `[harness/ruview, tools/ruview-mcp, tools/ruview-cli]` × +Node `[20, 22]`: + +1. `npm ci` where a lockfile is committed (the TS packages); the harness + installs with `npm install` — repo policy gitignores lockfiles under + `harness/`, and the package is dependency-free after ADR-263 O3 so there is + nothing to pin. +2. `npm test` (harness: `node --test test/*.test.mjs` — pin the glob form, + the directory form fails on Node 22; TS packages: build + jest or `node:test` + per ADR-264 O8). +3. **Pack gate:** `npm pack --dry-run --json` asserted against a checked-in + expected file list + a max unpacked-size budget per package (harness ≤ 60 kB; + rvagent ≤ 130 kB post ADR-264 O2). Any new/missing/renamed shipped file is a + reviewed diff, not a surprise. +4. **Tarball smoke test:** install the packed tarball into a temp dir; run + `ruview --version`, `ruview doctor`, `rvagent` `--help`-equivalent, and a + Node `import()` of each declared export condition — this is the test that + would have caught ADR-264 F1 (`require` → nonexistent `dist/index.cjs`). +5. Bump `ci.yml` `NODE_VERSION` to `'20'` (independent of the matrix above). + +### D2 — Publish only from CI, with provenance + +Manual `npm publish` from laptops stops. A tag-triggered workflow +(`ruview-npm-release.yml`, mirroring the firmware release discipline) runs the +D1 gate, then `npm publish --provenance --access public` under the GitHub OIDC +token. Consequence: every published version is attested to a public commit + +workflow run — the npm-side analogue of the ADR-028 witness bundle. The +`prepublishOnly` script in each package runs the pack gate locally as a +belt-and-braces (publishing outside CI fails loudly, not silently). + +### D3 — Version single-sourcing + +Rule: **package.json is the only place a version string lives.** Runtime code +reads it (`createRequire(import.meta.url)('./package.json').version` or a +build-time define for the TS packages). CI greps for `\d+\.\d+\.\d+` literals in +`src/` of each package and fails on match (allowlist: test fixtures). This +retires ADR-263 F6 and ADR-264 F9 permanently instead of per-incident. + +### D4 — Namespace and bin ownership + +- `@ruvnet/ruview` **owns the `ruview` bin** (it is the published front door, + ADR-182). `@ruv/ruview-cli` renames its bin or folds into `rvagent` + (ADR-264 O9) — decided here so neither package ADR relitigates it. +- New Node packages in this repo use the `@ruvnet/` scope (the `@ruv/` scope + holds `rvcsi` legacies; do not grow it). +- Every package README + description must pass + `npx ruview claim-check` — enforced in the D1 gate. The guardrail package + linting its sibling packages' claims is the cheapest dogfooding we have + (ADR-264 F3 is the standing example of why). + +### D5 — Shared-code policy (bounded) + +Do **not** introduce an npm workspace or a shared runtime package yet: three +packages, two of which may merge (ADR-264 O9), do not justify workspace +machinery, and the harness's zero-dep property is load-bearing. Revisit if a +fourth package appears or if the `http/cog/config` duplication survives the +ADR-264 O9 fold. Record the duplication as intentional in each file header (the +CLI already does this). + +## Consequences + +- The npm artifacts get the same class of gate the Rust workspace has had since + ADR-028: no publish without tests, no shipped file set without an asserted + manifest, no version without provenance. The two defects that reached the + registry (broken `require` condition, dead maps) become CI-impossible. +- Cold-path costs stay near zero: the D1 matrix is 6 fast jobs (the harness + suite runs in ~108 ms MEASURED; TS builds dominate at a few tens of seconds). +- Publishing gains one constraint (must go through CI) and loses one failure + mode (laptop-state publishes) — the right trade for a project whose brand is + reproducible evidence. +- D3's grep gate is blunt but cheap; if it over-fires, scope it to + `version`-adjacent identifiers before weakening it. +- Follow-ups tracked elsewhere: per-package code fixes (ADR-263 O1–O8, ADR-264 + O1–O9); ADR-182 P4 (metaharness router + ed25519 provenance chain) remains + the deeper provenance story that D2's npm attestations complement, not + replace. diff --git a/api-docs/adr/ADR-266-mediatek-filogic-csi-platform.md b/api-docs/adr/ADR-266-mediatek-filogic-csi-platform.md new file mode 100644 index 00000000..7ad6e4d0 --- /dev/null +++ b/api-docs/adr/ADR-266-mediatek-filogic-csi-platform.md @@ -0,0 +1,75 @@ +# ADR-266: MediaTek Filogic CSI Platform + +- **Status**: accepted +- **Date**: 2026-07-18 +- **Deciders**: RuView maintainers +- **Tags**: mediatek, filogic, mt76, csi, openwrt, rust + +## Context + +RuView needs a high-antenna-count, router-class Wi-Fi sensing path beyond ESP32. +MediaTek Filogic platforms are attractive because the upstream BSD-3-Clause +`mt76` driver supports MT7915/MT792x/MT7996 families and OpenWrt supports +MT7981/MT7986/MT7988 systems. The OpenWrt One (MT7981B + MT7976C) additionally +publishes schematics, platform datasheets, register documentation, serial, and +JTAG access. The BPI-R3 (MT7986 + MT7975N/P) offers dual-band 4x4 radios. + +The current upstream `mt76` tree has testmode, debugfs, RX descriptors, and MCU +event plumbing, but no supported public interface for exporting per-packet +complex channel estimates. Public MediaTek SDK material likewise does not expose +an equivalent to Espressif's CSI callback. PHY computation of channel estimates +does not imply that firmware transfers those estimates to host memory. + +Existing RuView documents that describe MT7661 CSI-over-UDP or released +MediaTek CSI tools are unverified architectural hypotheses, not supported +hardware claims. + +## Decision + +1. Use the OpenWrt One as the primary future hardware/upstreaming target and the + BPI-R3 as the secondary 4x4 validation target. +2. Build a Rust-first simulator and host transport before hardware arrives. +3. Keep the transport independent of private firmware structures. A future + `mt76` adapter must translate a documented kernel/firmware report into it. +4. Prefer Generic Netlink for capability/control messages and relayfs or a + bounded character-device stream if sustained CSI volume exceeds Netlink's + practical throughput. +5. Do not redistribute vendor firmware, private headers, or SDK components. +6. Label simulator frames end-to-end and never present them as physical capture. +7. Do not claim MediaTek hardware CSI support until complex CSI from a physical + device passes calibration, sequence, timestamp, and repeatability tests. + +## Consequences + +### Positive + +- Development and integration testing can start without fabricating a vendor ABI. +- OpenWrt One provides a repairable, upstream-friendly hardware target. +- The same RuView ingestion path can accept simulator, replay, and future driver data. +- Rust bounds checking isolates untrusted kernel/network input from inference code. + +### Negative + +- The simulator cannot prove firmware export availability or sensing accuracy. +- A firmware change or MediaTek cooperation may be required before physical CSI exists. +- Router-class builds and driver iteration are slower than MCU firmware development. + +### Neutral + +- NeuroPilot may later accelerate inference but is unrelated to CSI capture. +- Wi-Fi 7/MLO support remains a later phase after a single-link contract is stable. + +## Hardware gates + +- Identify a firmware/host report containing complex channel estimates. +- Document dimensions, quantization, chain ordering, subcarrier indexing, lifetime, + timestamps, sequence behavior, calibration, maximum size, and report rate. +- Validate OpenWrt One first, then BPI-R3 4x4, before considering MT7996/MLO. + +## Links + +- [ADR-123: BFLD capture path](ADR-123-bfld-capture-path-nexmon-and-esp32.md) +- [ADR-264: RTL8720F radar wire protocol](ADR-264-rtl8720f-radar-wire-protocol.md) +- [upstream mt76](https://github.com/openwrt/mt76) +- [OpenWrt One](https://openwrt.org/toh/openwrt/one) +- [MediaTek OpenWrt feed](https://git01.mediatek.com/openwrt/feeds/mtk-openwrt-feeds/) diff --git a/api-docs/adr/ADR-267-mediatek-mimo-csi-wire-protocol.md b/api-docs/adr/ADR-267-mediatek-mimo-csi-wire-protocol.md new file mode 100644 index 00000000..f7512389 --- /dev/null +++ b/api-docs/adr/ADR-267-mediatek-mimo-csi-wire-protocol.md @@ -0,0 +1,61 @@ +# ADR-267: MediaTek MIMO CSI Wire Protocol + +- **Status**: accepted +- **Date**: 2026-07-18 +- **Deciders**: RuView maintainers +- **Tags**: mediatek, csi, protocol, rust, udp, replay + +## Context + +The MediaTek simulator, captured regression fixtures, and a future `mt76` agent +need one safe host-side representation. Copying an undocumented firmware layout +would couple RuView to a private ABI and make malformed kernel/network data risky. +MIMO CSI also requires explicit Tx/Rx/subcarrier dimensions and per-Rx-chain RSSI. + +## Decision + +Define `MTC1` version 1 as a little-endian, self-delimiting envelope: + +- 72-byte fixed header with magic, version, report kind, total length, sequence, + monotonic timestamp, device ID, chipset profile, frequency, bandwidth, flags, + Tx/Rx dimensions, numeric format, PPDU type, subcarrier count, noise floor, + scale, subcarrier spacing, calibration ID, and payload length. +- CSI payload begins with one signed RSSI byte per Rx chain, followed by + `tx_count * rx_count * subcarrier_count` complex values in Tx-major, + Rx-major, subcarrier-major order. +- Supported numeric formats are complex signed i16 and complex finite f32. +- Capability reports use bounded opaque TLVs until a public driver contract exists. +- CRC-32/IEEE covers header and payload; the final four bytes carry the checksum. +- One envelope maps to one UDP datagram, capped at the IPv4 UDP payload maximum + of 65,507 bytes. Replay files prefix each envelope with a little-endian `u32`. +- Parsers reject unknown versions/types/formats, invalid dimensions/bandwidth, + multiplication overflow, inconsistent payload lengths, non-finite floats, + bad CRC, trailing datagram bytes, and frames above the cap. +- Flags distinguish calibrated, saturated, time-synchronized, dropped-predecessor, + and synthetic frames. Synthetic provenance cannot be cleared by downstream code. + +## Consequences + +### Positive + +- Deterministic simulator and future hardware use identical parsing and APIs. +- Explicit dimensions prevent ambiguous antenna or subcarrier interpretation. +- CRC, finite-value checks, and hard caps make network/replay ingestion robust. +- The format supports MT7981, MT7986, and MT7996 profiles without claiming their + undocumented firmware layouts. + +### Negative + +- A translation/copy step is required from a future kernel report. +- Maximum-size Wi-Fi 7 matrices may need segmentation in a later protocol version. + +### Neutral + +- Version 1 models one link per report; MLO correlation is a future extension. +- Capability TLVs are intentionally conservative until hardware metadata is known. + +## Links + +- [ADR-266: MediaTek Filogic CSI platform](ADR-266-mediatek-filogic-csi-platform.md) +- [ADR-018: ESP32 binary CSI framing](ADR-018-esp32-csi-frame-protocol.md) +- [ADR-264: RTL8720F radar wire protocol](ADR-264-rtl8720f-radar-wire-protocol.md) diff --git a/api-docs/adr/ADR-268-qualcomm-atheros-csi-platform.md b/api-docs/adr/ADR-268-qualcomm-atheros-csi-platform.md new file mode 100644 index 00000000..df515065 --- /dev/null +++ b/api-docs/adr/ADR-268-qualcomm-atheros-csi-platform.md @@ -0,0 +1,41 @@ +# ADR-268: Qualcomm Atheros CSI Platform Strategy + +- **Status**: accepted +- **Date**: 2026-07-18 +- **Tags**: qualcomm, atheros, csi, ath9k, ath11k, ath12k, simulator + +## Context + +RuView needs a Qualcomm path that is useful before vendor hardware access while +remaining honest about firmware boundaries. QCA9300 has demonstrated CSI tooling +through ath9k/PicoScenes-class systems. QCN9074 and QCN9274 have upstream Linux +connectivity drivers, but upstream ath11k/ath12k support does not by itself prove +that raw per-packet complex CSI is exported by public firmware. + +## Decision + +1. Use QCA9300 as the first physical baseline: 802.11n, up to 3x3 MIMO and + 20/40 MHz. Accept translated captures from established research tooling. +2. Model QCN9074 (Wi-Fi 6/6E, 4x4, up to 160 MHz) and QCN9274 (Wi-Fi 7, 4x4, + up to 160 MHz in protocol v1) as explicitly experimental simulator profiles. +3. Keep firmware/kernel formats behind a Rust adapter. RuView ingests only the + validated QCS1 application envelope defined by ADR-269. +4. Never label simulated frames as hardware. Physical support requires captured + fixtures, firmware provenance, antenna ordering, scaling and repeatability tests. +5. Prefer an upstream-reviewed Generic Netlink or relay-style export if modern + Qualcomm firmware exposes CFR/CSI; do not depend on undisclosed structs. + +## Consequences + +- Development, APIs and downstream sensing can be tested immediately. +- QCA9300 offers the shortest path to real Qualcomm data. +- Modern profiles may remain simulator-only until firmware cooperation exists. +- A translation copy is accepted in exchange for a stable, fuzzable boundary. + +## Links + +- [ADR-269: QCS1 wire protocol](ADR-269-qualcomm-csi-wire-protocol.md) +- [Linux ath11k supported devices](https://wireless.docs.kernel.org/en/latest/en/users/drivers/ath11k.html) +- [PicoScenes supported hardware](https://ps.zpj.io/manual/hardware.html) +- [ADR-270: vendor integration portfolio and acceptance gates](ADR-270-vendor-rf-sensing-integration-program.md) + diff --git a/api-docs/adr/ADR-269-qualcomm-csi-wire-protocol.md b/api-docs/adr/ADR-269-qualcomm-csi-wire-protocol.md new file mode 100644 index 00000000..a2ff83bd --- /dev/null +++ b/api-docs/adr/ADR-269-qualcomm-csi-wire-protocol.md @@ -0,0 +1,40 @@ +# ADR-269: Qualcomm CSI Wire Protocol + +- **Status**: accepted +- **Date**: 2026-07-18 +- **Tags**: qualcomm, csi, protocol, rust, udp, replay + +## Decision + +Define `QCS1` version 1 as a vendor-boundary envelope, not a Qualcomm firmware ABI. +It uses a 72-byte little-endian header plus payload and CRC-32/IEEE. The header +records report kind, total length, sequence, monotonic timestamp, device ID, +chipset profile, center frequency, bandwidth, flags, Tx/Rx counts, numeric format, +PPDU type, subcarrier count, noise floor, scale, subcarrier spacing, calibration +ID and payload length. + +CSI payloads contain one signed RSSI byte per receive chain followed by +`tx * rx * subcarriers` complex i16 or finite f32 values in Tx-major, Rx-major, +subcarrier-major order. Capability reports carry bounded opaque bytes. One QCS1 +frame maps to one UDP datagram; replay files prefix each frame with a little-endian +u32 length. + +Parsers fail closed on unknown enums, bad CRC, truncation, trailing datagram data, +non-finite values, inconsistent dimensions, chipset chain/bandwidth violations, +payload mismatches, arithmetic overflow and the IPv4 UDP payload ceiling. A +synthetic flag provides end-to-end simulator provenance. + +Version 1 profiles are QCA9300, QCN9074 and QCN9274. QCA9300 is capped at three +chains and 40 MHz; modern profiles are capped at four chains and 160 MHz. + +## Consequences + +- Simulator, replay and future hardware adapters share one validated Rust API. +- No private firmware layout is represented or redistributed. +- 320 MHz/EHT matrices require segmentation or a later protocol revision. + +## Links + +- [ADR-268: Qualcomm platform strategy](ADR-268-qualcomm-atheros-csi-platform.md) +- [ADR-267: MediaTek MTC1 protocol](ADR-267-mediatek-mimo-csi-wire-protocol.md) + diff --git a/api-docs/adr/ADR-270-vendor-rf-sensing-integration-program.md b/api-docs/adr/ADR-270-vendor-rf-sensing-integration-program.md new file mode 100644 index 00000000..311ca308 --- /dev/null +++ b/api-docs/adr/ADR-270-vendor-rf-sensing-integration-program.md @@ -0,0 +1,122 @@ +# ADR-270: Vendor RF Sensing Integration Program + +- **Status**: accepted +- **Date**: 2026-07-18 +- **Deciders**: RuView maintainers +- **Tags**: vendors, csi, telemetry, simulator, rust, hardware-validation + +## Context + +RuView is evaluating Qualcomm, RF Solutions, Origin AI, Plume, Linksys, +Electric Imp, Mist/Juniper, Luma, Google Nest, NETGEAR and Wifigarden. These +names do not represent equivalent integration surfaces: some expose raw CSI, +some expose derived sensing events or network telemetry, and some expose no +supported developer interface. A repeated implementation process must not turn +brand compatibility, Linux connectivity or synthetic fixtures into a false CSI +claim. + +## Decision + +Adopt a Rust-first provider portfolio with explicit capability negotiation: + +- `ComplexCsi`: calibrated per-packet complex channel matrices. +- `DerivedSensing`: vendor-produced motion, occupancy or location events. +- `RfTelemetry`: RSSI, radio, client and topology observations. +- `NetworkOnly`: useful as excitation/AP infrastructure but not a sensor. +- `Unsupported`: no stable, lawful or supportable integration surface. + +Every provider follows the same gated loop: + +1. Verify an authoritative API/SDK, exact model/chipset and licensing boundary. +2. Write provider and wire/contract ADRs before coupling core code to a vendor. +3. Implement bounded Rust types, explicit capabilities and synthetic provenance. +4. Test deterministic replay, corruption, loss, reconnect, backpressure, schema + evolution and secrets handling. +5. Promote to hardware support only after lawful physical capture on an exact + model/firmware, calibration and repeatability tests, and fixture publication + rights. Simulator success never satisfies this gate. +6. Publish code/release and an upstream or vendor collaboration announcement + that states the measured-versus-simulated boundary. + +### Portfolio decisions + +| Provider | Classification | Decision | +|---|---|---| +| Qualcomm QCA9300 | `ComplexCsi` candidate | Implement first physical baseline via established ath9k research tooling; QCS1 adapter ships simulator-first. | +| Qualcomm QCN9074/QCN9274 | experimental `ComplexCsi` | Simulator and protocol now; require confirmed ath11k/ath12k firmware export before hardware claim. | +| Origin AI | commercial `DerivedSensing`, possible CSI | Pursue NDA sandbox/API and raw-data rights; isolate proprietary engine behind provider trait/service boundary. | +| Plume/OpenSync | `RfTelemetry`; Plume Sense is gated `DerivedSensing` | Build optional OVSDB/control-plane adapter; negotiate Sense separately and do not infer raw CSI. | +| Mist/Juniper | `RfTelemetry` + location | Conditional read-only REST/webhook adapter for occupancy, RSSI and coordinates; no CSI claim. | +| NETGEAR | partner-gated `RfTelemetry` | Insight adapter only after API access; exact legacy OpenWrt models remain community experiments. | +| Luma | discontinued OpenWrt salvage target | Generic OpenWrt telemetry/pcap fixture only when already owned; no procurement or Luma CSI source. | +| Google Nest Wifi | `NetworkOnly` | Use as traffic/AP infrastructure; Device Access does not expose router CSI or radio telemetry. | +| Linksys | `Unsupported` for sensing | Linksys Aware reached end of support in 2024; record capability probe only, if needed. | +| Electric Imp | scalar IoT/RSSI telemetry | Optional agent/impCentral bridge for existing fleets; reject as CSI acquisition hardware. | +| RF Solutions | non-Wi-Fi RF/IoT telemetry | Exclude from sensing backend; optional RIoT environmental fusion is a separate future concern. | +| Wifigarden | commercial OEM, capability unknown | Hold implementation pending chipset, schema, offline, calibration and data-rights disclosure. | + +### Provider boundary + +Core code consumes a vendor-neutral `RfSource`-style contract whose capability +set prevents RSSI, location or derived occupancy from being represented as CSI. +Cloud adapters use bounded async queues, regional endpoints, secret-provider +credentials and explicit data provenance. Proprietary device SDKs live behind a +feature-gated FFI or sidecar boundary and are never redistributed without rights. + +## Consequences + +### Positive + +- The integration loop can be repeated without duplicating unsafe parsers. +- Product integrations remain useful even when only telemetry is available. +- Public releases make hardware confidence and simulator confidence distinct. + +### Negative + +- Several named vendors cannot produce a legitimate CSI implementation today. +- Commercial providers require contracts, subscriptions, test vectors or NDAs. +- Exact hardware revisions and firmware provenance increase validation effort. + +### Neutral + +- A no-go or telemetry-only ADR is a completed research outcome, not a failed port. +- Vendor status and APIs must be rechecked before each implementation begins. + +## Implementation Status + +The ADR-270 provider contract is implemented in Rust. Each portfolio entry has +a descriptor, bounded decoder or explicit fail-closed access state, deterministic +contract fixtures where lawful, registry coverage, and API exposure: + +- Origin AI: contract-configured derived-sensing decoder and request plan. +- Plume/OpenSync: read-only OVSDB request plan and RF telemetry decoder. +- Mist/Juniper: regional request configuration, paginated RF/location decoder. +- NETGEAR Insight: regional partner request configuration and telemetry decoder. +- Electric Imp and RF Solutions: bounded scalar telemetry bridges. +- Luma: explicitly experimental generic OpenWrt telemetry bridge. +- Google Nest: network-only contract events; never represented as CSI. +- Linksys: `Unsupported` decoder because Linksys Aware is end-of-support. +- Wifigarden: `ContractRequired` decoder pending a disclosed SDK/schema. + +`vendor-rf-sim` generates deterministic, provenance-labelled events for the +eight providers with a defined event contract and refuses to fabricate Linksys +or Wifigarden events. The sensing server exposes provider descriptors and latest +events under `/api/v1/rf/vendors` and accepts validated canonical simulator +events over its existing UDP port. Physical/vendor-cloud validation remains +separate from implementation completeness and is reflected by +`hardware_validated: false` until performed. + +## Evidence and Links + +- [ADR-268: Qualcomm strategy](ADR-268-qualcomm-atheros-csi-platform.md) +- [OpenSync developer sandbox](https://www.opensync.io/developer) +- [Origin AI Wi-Fi sensing architecture](https://www.originwirelessai.com/wifi-sensing/) +- [Juniper Mist webhook hierarchy](https://www.juniper.net/documentation/us/en/software/mist/automation-integration/topics/topic-map/webhook-hierarchy.html) +- [Linksys product end-of-life](https://www.linksys.com/pages/linksys-product-end-of-life) +- [Google Nest Device Access supported devices](https://developers.google.com/nest/device-access/supported-devices) +- [OpenWrt Luma WRTQ-329ACN](https://openwrt.org/toh/hwdata/luma/luma_wrtq-329acn) +- [NETGEAR Insight compatible devices](https://kb.netgear.com/000048452/What-devices-can-I-discover-monitor-and-manage-with-Insight) +- [Electric Imp imp005 hardware guide](https://developer.electricimp.com/hardware/imp/imp005_hardware_guide) +- [RF Solutions company portfolio](https://www.rfsolutions.co.uk/about-us-i1/) +- [Wifigarden service terms](https://policies.wifigarden.com/en-us/terms-of-service) + diff --git a/api-docs/adr/README.md b/api-docs/adr/README.md index 4826fef0..7c9e510e 100644 --- a/api-docs/adr/README.md +++ b/api-docs/adr/README.md @@ -1,6 +1,11 @@ # Architecture Decision Records -This folder contains 45 Architecture Decision Records (ADRs) that document every significant technical choice in the RuView / WiFi-DensePose project. +Latest proposed decisions: + +- [ADR-264: Versioned wire protocol for RTL8720F CFR and Range-FFT reports](ADR-264-rtl8720f-radar-wire-protocol.md) +- [ADR-263: Adopt RTL8720F 2.4 GHz FMCW radar as an optional RuView sensing platform](ADR-263-rtl8720f-2-4ghz-fmcw-radar-platform.md) + +This folder contains 182 Architecture Decision Records (ADRs) that document every significant technical choice in the RuView / WiFi-DensePose project. (The index tables below list a curated subset per domain; see the directory listing for the full set.) ## Why ADRs? @@ -120,6 +125,9 @@ Statuses: **Proposed** (under discussion), **Accepted** (approved and/or impleme | [ADR-097](ADR-097-adopt-rvcsi-as-ruview-csi-runtime.md) | Adopt rvCSI as RuView's primary CSI runtime (phased adoption) | Proposed | | [ADR-098](ADR-098-evaluate-midstream-fit.md) | Evaluate `ruvnet/midstream` for RuView's CSI / WebSocket / mesh pipeline | Rejected | | [ADR-099](ADR-099-midstream-introspection-tap.md) | Adopt midstream as RuView's real-time introspection + low-latency tap | Proposed | +| [ADR-263](ADR-263-ruview-npm-harness-deep-review.md) | `@ruvnet/ruview` npm harness — deep review + optimization strategy | Proposed | +| [ADR-264](ADR-264-rvagent-mcp-and-cli-npm-deep-review.md) | `@ruvnet/rvagent` MCP server + `@ruv/ruview-cli` — deep review + optimization strategy | Proposed | +| [ADR-265](ADR-265-ruview-npm-distribution-strategy.md) | RuView npm distribution strategy — CI gate, provenance, version single-sourcing, namespace | Proposed | --- diff --git a/api-docs/ddd/deployment-platform-domain-model.md b/api-docs/ddd/deployment-platform-domain-model.md index 4d883655..d693c7c7 100644 --- a/api-docs/ddd/deployment-platform-domain-model.md +++ b/api-docs/ddd/deployment-platform-domain-model.md @@ -425,7 +425,7 @@ pub enum WifiChipset { BroadcomBcm43455, /// Realtek RTL8822CS via modified rtw88 driver. RealtekRtl8822cs, - /// MediaTek MT7661 via mt76 driver modification. + /// Proposed MediaTek MT7661 research target; no public CSI export is verified. MediatekMt7661, } @@ -455,7 +455,7 @@ pub struct Esp32CompatFrame { ``` **Domain Services:** -- `CsiExtractionService` — Reads raw CSI from patched driver via Netlink socket (BCM43455), procfs (RTL8822CS), or UDP (MT7661) +- `CsiExtractionService` — Reads raw CSI from a validated chipset adapter. Nexmon/BCM43455 is the established Linux example; RTL8822CS and MT7661 remain unverified research targets and must not be advertised as working capture paths. - `SubcarrierResamplerService` — Resamples chipset-specific subcarrier counts to match ESP32 format (e.g., 256 → 128 via decimation or interpolation) - `ProtocolTranslatorService` — Converts `ChipsetCsiFrame` to `Esp32CompatFrame` with ADR-018 binary encoding - `CalibrationService` — Compensates for chipset-specific phase offsets, antenna spacing, and gain differences relative to ESP32 CSI @@ -625,7 +625,7 @@ pub struct EspNodeConnection { ### ESP32 Protocol ACL (CSI Bridge) -The WiFi CSI Bridge translates chipset-specific CSI formats (Nexmon, rtw88, mt76) into the ESP32 binary protocol (ADR-018). The sensing server never knows whether frames came from a real ESP32 or a TV box WiFi chipset. Virtual node IDs (200-254) prevent collision with physical ESP32 IDs but are otherwise treated identically by the ingestion context. +The WiFi CSI Bridge translates validated chipset-specific CSI formats into a versioned RuView envelope. Nexmon is the established Linux example; rtw88 and mt76 require a verified complex-CSI export before implementation. Virtual node IDs (200-254) prevent collision with physical ESP32 IDs but are otherwise treated identically by the ingestion context. ### Armbian Platform ACL diff --git a/api-docs/releases/v0.9.0-realtek-beta.1.md b/api-docs/releases/v0.9.0-realtek-beta.1.md new file mode 100644 index 00000000..9040a077 --- /dev/null +++ b/api-docs/releases/v0.9.0-realtek-beta.1.md @@ -0,0 +1,45 @@ +# RuView v0.9.0-realtek-beta.1 + +This prerelease introduces the Rust-first RTL8720F 2.4 GHz radar transport and +RuView ingestion path. It is intentionally simulator-validated until Realtek +hardware and the vendor SDK callback ABI arrive. + +## Included + +- ADR-263 records the upstream Ameba integration and licensing boundary. +- ADR-264 defines a versioned, bounded, CRC-protected radar envelope. +- `rtl8720f-sim` emits deterministic CFR, near-range, far-range, interference, + and capability reports to UDP or replay files. +- The sensing server validates RTL8720F datagrams, publishes bounded summaries + over `/ws/sensing`, and exposes the latest report at + `/api/v1/radar/latest`. +- Synthetic provenance is retained end to end as `realtek:simulated`; simulator + data is never presented as hardware data. + +## Compatibility + +The adapter tracks the radar control surface proposed by Ameba RTOS pull +request #1336 (`wifi_radar_config`, `AT+RAD`, and `AT+RADDBG`). The stable Ameba +RTOS v1.2.1 release does not yet expose the complete radar receive callback ABI, +so no vendor-private headers or binary libraries are copied into this release. + +## Validation status + +- Rust codec round trips, corruption rejection, size bounds, and deterministic + simulator tests pass. +- RuView server ingestion, REST reporting, and source provenance were exercised + end to end over loopback UDP. +- Windows release binaries are built from this branch and accompanied by + SHA-256 checksums. + +## Known limitations + +- No physical RTL8720F board has been flashed or measured. +- The vendor report callback and exact report layouts remain an SDK/hardware + validation gate; the adapter boundary may change when those arrive. +- This beta exposes transport and aggregate radar observability. Radar-to-pose, + vital-sign inference, RF calibration, and accuracy claims are not enabled. +- 2.4 GHz radar reports are not mislabeled as mmWave or Wi-Fi CSI events. + +Do not deploy this prerelease for safety-critical, medical, or occupancy billing +uses. It is an integration beta for SDK and hardware bring-up. diff --git a/api-docs/releases/v0.9.1-mediatek-beta.1.md b/api-docs/releases/v0.9.1-mediatek-beta.1.md new file mode 100644 index 00000000..778f69cc --- /dev/null +++ b/api-docs/releases/v0.9.1-mediatek-beta.1.md @@ -0,0 +1,32 @@ +# RuView v0.9.1-mediatek-beta.1 + +This simulator-first beta adds a Rust MediaTek Filogic MIMO CSI transport and +RuView ingestion path while preserving the boundary between demonstrated host +integration and unavailable physical CSI export. + +## Included + +- ADR-266 selects OpenWrt One (MT7981/MT7976) as the primary future hardware + target and BPI-R3 (MT7986/MT7975) as the secondary 4x4 target. +- ADR-267 defines the bounded, versioned, CRC-protected `MTC1` wire protocol. +- `mediatek-csi-sim` provides deterministic MT7981, MT7986, and MT7996 profiles, + complex MIMO CSI, per-chain RSSI, UDP streaming, and replay output. +- RuView validates MediaTek datagrams, publishes bounded WebSocket summaries, + and exposes `/api/v1/csi/mediatek/latest`. +- `mediatek:simulated` provenance is retained end to end. + +## Validation + +- Codec round trips, deterministic output, corruption/truncation rejection, + dimension limits, finite-value enforcement, and prefix parsing are tested. +- All hardware and sensing-server regression tests pass. +- All three profiles were streamed over loopback UDP and verified through the + RuView REST API. + +## Hardware boundary + +Upstream `mt76` and public MediaTek SDK material do not currently expose a +supported raw complex CSI API. This release does not redistribute private SDK +material, invent a firmware ABI, or claim physical MediaTek capture. Hardware +support requires a documented firmware/driver channel-estimate export followed +by calibration and repeatability validation. diff --git a/api-docs/releases/v0.9.2-qualcomm-beta.1.md b/api-docs/releases/v0.9.2-qualcomm-beta.1.md new file mode 100644 index 00000000..2d16179a --- /dev/null +++ b/api-docs/releases/v0.9.2-qualcomm-beta.1.md @@ -0,0 +1,22 @@ +# RuView v0.9.2-qualcomm-beta.1 + +This simulator-first beta adds a Rust Qualcomm Atheros CSI boundary without +claiming modern Qualcomm firmware exports that have not been physically verified. + +## Included + +- ADR-268 selects QCA9300 as the first physical baseline and treats QCN9074 and + QCN9274 as experimental modern profiles. +- ADR-269 defines the bounded, versioned, CRC-protected `QCS1` protocol. +- `qualcomm-csi-sim` emits deterministic MIMO CSI over UDP or replay files. +- The sensing server validates QCS1 datagrams, broadcasts bounded summaries and + exposes `/api/v1/csi/qualcomm/latest`. +- `qualcomm:simulated` provenance is retained end to end. + +## Validation boundary + +Codec, corruption, truncation, finite-value, dimensions, chipset bandwidth, +determinism and prefix parsing are automated. Loopback UDP/API validation covers +all profiles. Physical QCA9300 comparison and modern firmware export validation +remain hardware gates and will be published with firmware and calibration details. + diff --git a/api-docs/releases/v0.9.3-vendor-providers-beta.1.md b/api-docs/releases/v0.9.3-vendor-providers-beta.1.md new file mode 100644 index 00000000..c6c1001d --- /dev/null +++ b/api-docs/releases/v0.9.3-vendor-providers-beta.1.md @@ -0,0 +1,21 @@ +# RuView v0.9.3 Vendor Providers Beta 1 + +This beta implements ADR-270 as a capability-safe Rust provider program across +all ten researched vendors. + +## Included + +- Shared `VendorRfProvider` contract with bounded event validation. +- Origin AI, Plume/OpenSync, Mist/Juniper, NETGEAR Insight, Electric Imp, + RF Solutions, Luma/OpenWrt and Google Nest contract adapters. +- Explicit fail-closed Linksys (`Unsupported`) and Wifigarden + (`ContractRequired`) providers. +- Deterministic `vendor-rf-sim` JSONL/UDP fixtures for defined contracts. +- Provider registry, descriptors, latest-event REST endpoints and WebSocket + summaries through the sensing server. + +## Boundary + +This release implements and validates software contracts. It does not claim +vendor-cloud credentials, commercial SDK rights, physical hardware validation, +or complex CSI support for telemetry-only providers. diff --git a/api-docs/research/sota-nn-train-benchmark-brief.md b/api-docs/research/sota-nn-train-benchmark-brief.md new file mode 100644 index 00000000..6328f4a5 --- /dev/null +++ b/api-docs/research/sota-nn-train-benchmark-brief.md @@ -0,0 +1,147 @@ +# SOTA Evidence Brief — `wifi-densepose-nn` / `wifi-densepose-train` Benchmark ADR Seed + +| Field | Value | +|-------|-------| +| **Date** | 2026-06-14 | +| **Author** | deep-research (Opus) | +| **Purpose** | Seed a future benchmark/optimization ADR for the NN-inference (`wifi-densepose-nn`) and training (`wifi-densepose-train`) crates | +| **Scope** | The DELTA beyond what ADR-152 / ADR-150 / ADR-015 already establish — current published WiFi-CSI pose SOTA, winning architectures, edge-quantization SOTA, and a defensible benchmark-suite design | +| **Ethos** | Every claim graded PEER-REVIEWED / PREPRINT / VENDOR-CLAIM / BLOG, with MEASURED-on-public-benchmark distinguished from marketing. Numbers that could not be verified are flagged. No fabricated citations. | + +> **Citation discipline carried in from ADR-152 §2.2:** preprint accuracy numbers are CLAIMED until reproduced on our hardware. The project has already retracted its own "92.9% PCK@20" and "shipped-WiFlow-STD 97.25%" figures after measurement; this brief inherits that bar. + +--- + +## 1. Executive summary + +**Where the project stands vs the 2026 frontier.** The repo is, by the evidence already in-tree, *ahead of most academic groups on benchmark hygiene* and roughly *at parity on capability* — but the two are measured on incompatible yardsticks, which is the single biggest risk to any "beyond-SOTA" claim. + +- The project's headline reproductions (`benchmarks/wiflow-std/RESULTS.md`) are MEASURED and rigorous: WiFlow-STD retrained to **96.09–96.61% PCK@20** on the authors' own 360k-window 2D dataset (RTX 5080), shipped checkpoint REFUTED, dataset/code defects documented. This is a genuinely strong, reproducible result. +- **But that number is not on a standard public benchmark.** WiFlow-STD's dataset is self-collected (5 subjects, 15 keypoints, 2D, in-domain random split, hardware unspecified). The academic frontier on the *standard* public 3D benchmark (MM-Fi) reports **PCK@20 ≈ 61% / MPJPE ≈ 161 mm random-split** (GraphPose-Fi, Nov 2025) — a *harder* metric (3D, mm-scale, standard PCK normalization). The project's own AetherArena MM-Fi number (**81.63% torso-PCK@20 in-domain**, ADR-150) uses a *torso-normalized PCK* that is looser than GraphPose-Fi's standard PCK, so the three numbers (96% / 81.6% / 61%) **cannot be lined up** without a unified harness. Making them comparable IS the highest-value work item. +- The deployment frontier — **cross-subject / cross-environment generalization** — is where everyone collapses, the project included (ADR-150: 81.63% in-domain → ~11.6% leakage-free cross-subject). GraphPose-Fi independently confirms the cliff (61.1% random → 12.9% cross-environment PCK@20). This is the real research target, not in-domain PCK. + +**Top 3 highest-value optimization/benchmark targets:** + +1. **A unified, metric-locked accuracy harness in `wifi-densepose-train`** that scores any model under *one* explicit PCK definition (normalization, keypoint convention, split) so WiFlow-STD-repro, AetherArena/MM-Fi, and GraphPose-Fi numbers become directly comparable. Without this, no "beyond-SOTA" claim survives the "prove it" bar — the project has already been burned twice by metric ambiguity (the retracted 92.9% used absolute, not torso-normalized, PCK). +2. **A QAT path for the WiFlow-STD-class edge model.** The in-tree edge work (`RESULTS.md`) has *fully characterized PTQ* (static QDQ conv-only is the int8 sweet spot; dynamic int8 is a no-op on this all-conv architecture) and found the **half model (843k params) strictly dominates the published 2.23M** and **tiny (56k, 295 KB ONNX fp32) holds 94.1% PCK@20**. The one untested lever is **quantization-aware training**, which the general literature says recovers most of the PTQ accuracy gap. That is the next defensible edge win. +3. **Criterion-backed regression benches wired into CI** for the real Candle/ONNX forward path. The benches *exist* (`wifi-densepose-nn/benches/{inference,onnx,native_conv}_bench.rs`, `wifi-densepose-train/benches/training_bench.rs`) and `benchmarks/edge-latency/RESULTS.md` shows the methodology is sound (host≠ESP32 caveat made explicit). The gap is turning point-in-time captures into committed regression baselines. + +--- + +## 2. Findings per research question + +### RQ1 — Latest WiFi-CSI pose SOTA (2024–2026): published PCK@20 / MPJPE on the standard public benchmarks + +The crucial framing: **"WiFi pose SOTA" splits into two non-comparable tracks** — 3D pose on MM-Fi/Person-in-WiFi-3D (mm-scale MPJPE, standard PCK) vs 2D pose on self-collected sets (image-normalized PCK). The project's flagship reproduction lives in the second track; the academic frontier lives in the first. + +| Method | Venue / Year | Benchmark + split | PCK@20 | MPJPE | Grade | +|---|---|---|---|---|---| +| **GraphPose-Fi** (arXiv [2511.19105](https://arxiv.org/abs/2511.19105)) | PREPRINT, Nov 2025 | MM-Fi P1, **random split** | **61.1%** | **160.6 mm** (PA-MPJPE 105.0) | numbers MEASURED-in-study (preprint); beats MetaFi++, HPE-Li, DT-Pose | +| GraphPose-Fi | same | MM-Fi P1, **cross-subject** | 44.2% | 210.5 mm | same | +| GraphPose-Fi | same | MM-Fi P1, **cross-environment** | 12.9% | 302.7 mm | same — the generalization cliff | +| **DT-Pose** (arXiv [2501.09411](https://arxiv.org/abs/2501.09411)) | PREPRINT (ICLR'25 OpenReview [aPnLQ6WfQQ](https://openreview.net/forum?id=aPnLQ6WfQQ)), Jan 2025; code [cseeyangchen/DT-Pose](https://github.com/cseeyangchen/DT-Pose) | MM-Fi (domain-gap + topology focus) | not cleanly extractable from abstract | reports MPJPE; self-supervised masked pretrain + topology decode | numbers NOT verified at exact-table level here — flagged | +| **Person-in-WiFi-3D** (CVPR 2024, [openaccess](https://openaccess.thecvf.com/content/CVPR2024/html/Yan_Person-in-WiFi_3D_End-to-End_Multi-Person_3D_Pose_Estimation_with_Wi-Fi_CVPR_2024_paper.html)) | **PEER-REVIEWED**, CVPR 2024 | own 97k-frame multi-person set | — (multi-person, not single-PCK) | **91.7 mm (1p) / 108.1 (2p) / 125.3 (3p)** 3D joint error | MEASURED (peer-reviewed); own dataset, not MM-Fi | +| **WiFlow-STD** (arXiv [2602.08661](https://arxiv.org/abs/2602.08661), [DY2434 repo](https://github.com/DY2434/WiFlow-WiFi-Pose-Estimation-with-Spatio-Temporal-Decoupling)) | PREPRINT, Apr 2026 | self-collected, 5-subj, **2D, in-domain random** | 97.25% (claimed) | 0.007 m (image-norm) | claimed CLAIMED; **project reproduced 96.09–96.61% (MEASURED, RTX 5080)** after repairing dataset/code | +| **PerceptAlign** (arXiv [2601.12252](https://arxiv.org/abs/2601.12252)) | PREPRINT + MobiCom'26 acceptance | own 7-layout cross-domain 3D set | — | 222.4 mm (Scene4) / 317.1 (Scene5), claims −54% cross-env vs SOTA | CLAIMED (preprint); failure mode corroborated | +| **Project AetherArena** (ADR-150, [issue #876](https://github.com/ruvnet/RuView/issues/876)) | internal | MM-Fi, **random split**, **torso-PCK** | **81.63% torso-PCK@20** | — | MEASURED-internal; **torso-PCK ≠ GraphPose-Fi standard PCK** | +| **Project WiFlow-STD repro** (`benchmarks/wiflow-std/RESULTS.md`) | internal | their data, their split | **96.09–96.61%** | 0.0094–0.0098 m | MEASURED-internal (RTX 5080) | + +**How the project's ~96% compares to the frontier:** It is *not directly comparable*. The 96% is on an easier task (2D, in-domain, image-normalized PCK, single-environment, 5 subjects) than GraphPose-Fi's 61.1% (3D, standard PCK, mm-scale). The project's own MM-Fi-track number (81.63% torso-PCK@20) *appears* to beat GraphPose-Fi's 61.1%, **but only because torso-PCK is a looser normalization** — the project explicitly flags this (ADR-150 cites beating "MultiFormer's 72.25%" under the *same* torso metric, not GraphPose-Fi's). The honest statement: **the project is competitive on in-domain MM-Fi under its own torso metric, and collapses cross-subject exactly as the published frontier does.** No public number lets the project claim "beyond-SOTA" today. + +### RQ2 — What's winning architecturally now (2025–2026) + +The clear trend across the verified 2025–2026 papers: + +- **Graph / skeleton-aware decoders are the current academic SOTA on MM-Fi.** GraphPose-Fi (PREPRINT, Nov 2025) wins by injecting anatomical graph structure into the decoder — exactly the `GraphPose-Fi-style skeleton-aware graph head` ADR-150 §2.2 already names as the planned decoder. *The project's architecture direction matches the frontier.* +- **Self-supervised masked pretraining (MAE) is the cross-domain lever, not capacity.** UNSW MAE study (arXiv [2511.18792](https://arxiv.org/abs/2511.18792), PREPRINT, Nov 2025): cross-domain gains scale **log-linearly with pretraining data, unsaturated at 1.3M samples**; ViT-Base adds only 0.4–0.9% over ViT-Small. Recipe: **80% masking, (30,3) small patches**. DT-Pose (arXiv 2501.09411) independently uses masked pretraining + topology constraints for the domain gap. *Caveat (MEASURED in ADR-152 §2.3): UNSW's downstream tasks are classification, not pose — pose transfer remains a hypothesis. The project's own measurement (b) found WiFlow-STD pretrained features give optimization transfer but NOT feature transfer to ESP32 CSI.* +- **Spatio-temporal decoupling is the efficiency lever.** WiFlow-STD's whole contribution is decoupling spatial and temporal CSI processing to hit 2.23M params. The project verified the params/FLOPs (MEASURED) and then **beat it**: the half-model (843k) matches accuracy with 0.38× params (`RESULTS.md` efficiency sweep). +- **Geometry/layout conditioning is the cross-layout lever.** PerceptAlign (MobiCom'26): fusing transceiver-position embeddings + two-checkerboard calibration, claimed −60% cross-domain. ADR-152 §2.1 already adopted this (`NodeGeometry`, geometry embeddings). +- **NOT winning / absent:** diffusion models for CSI pose did not surface in the verified frontier. Full DensePose-UV regression from commodity WiFi remains undemonstrated (ADR-152 F5, MEASURED by full-text screening). No 2025–2026 paper was found that *beats the project's current direction* — the project is tracking, not trailing, the architecture frontier. + +**Verdict RQ2:** the winning stack (MAE pretrain → graph/skeleton decoder → geometry conditioning, ViT-Small-class capacity) is *already the planned ADR-150/152 stack*. The gain available is not a new architecture; it's (a) more heterogeneous pretraining data and (b) honest cross-domain measurement. + +### RQ3 — Edge/quantized inference SOTA for small CSI pose models + +The in-tree edge work (`benchmarks/wiflow-std/RESULTS.md` "Edge optimization" + "Static PTQ" + "Efficiency sweep") is already at or beyond what the public literature offers for this specific model class, and is MEASURED. Key findings to carry forward: + +- **Dynamic INT8 is a trap on all-conv CSI models.** WiFlow-STD has **zero `nn.Linear` layers** (21 Conv1d + 22 Conv2d + BatchNorm). `torch.quantize_dynamic` quantizes 0% of params (dynamic int8 has no conv kernels). MEASURED. +- **Static QDQ conv-only PTQ is the int8 sweet spot.** PCK@20 96.60–96.63% (vs fp32 96.68%, dynamic 96.52%), 2.53 MB. All-ops QDQ is strictly worse (−1.4 pt). MEASURED. +- **ONNX Runtime fp32 is the real CPU latency win**: 3.2 ms/window batch-1 vs torch 11.0 ms (~3.4×) at parity (2.4e-7). int8 is ~2× *slower* than ONNX fp32 at batch-1 (ConvInteger kernels). MEASURED. +- **Smaller-than-published dominates.** half (843k) ≥ full on accuracy; **tiny (56k, 295 KB ONNX fp32, 0.66 ms/win, 94.1% PCK@20)** is the smallest deployable artifact. At tiny scale int8 is a *bad* trade (−1.43 pt for −47 KB). MEASURED. +- **General QAT-vs-PTQ context (BLOG/VENDOR):** [NVIDIA TensorRT QAT blog](https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/), [Ultralytics QAT glossary](https://www.ultralytics.com/glossary/quantization-aware-training-qat), [ONNX Runtime quantization docs](https://onnxruntime.ai/docs/performance/model-optimizations/quantization.html): QAT "almost always" recovers accuracy PTQ loses on sensitive models; ONNX Runtime does NOT retrain (QAT must happen in PyTorch, then export QDQ). The [Onboard Optimization survey, arXiv 2505.08793](https://arxiv.org/pdf/2505.08793) (PREPRINT) covers on-device optimization broadly. These are *general* claims, not CSI-pose-specific — grade accordingly. +- **Hailo / Pi target (CLAUDE.local.md):** the 4× Pi+Hailo cluster (Hailo-8 @ 26 TOPS / Hailo-10 @ 40 TOPS) needs a **HEF** compile path, which is its own toolchain (not ONNX/Candle). No in-tree HEF benchmark exists yet — this is a genuine gap for the edge-inference claim. + +**Actionable for an inference-speed benchmark:** the honest comparand set is `{torch fp32, ONNX fp32, ONNX static-QDQ-conv-only int8, candle fp32}` × `{full, half, tiny}` on a fixed host, with the **host≠ESP32 / host≠Hailo caveat stated up front** (the `edge-latency/RESULTS.md` template already does this correctly). The one new datapoint worth producing: **QAT-int8 on the half model** to test whether QAT closes the PTQ −0.16 pt gap *and* keeps the size win. + +### RQ4 — Rigorous, reproducible benchmark methodology + +The repo already demonstrates the right methodology in three places — the ADR should codify it, not invent it: + +- **`benchmarks/wiflow-std/RESULTS.md`** — the gold standard already in-tree: pinned upstream commit, seed-42 file-level split documented, corruption masks committed as ground truth, every forced deviation recorded, mean-pose honesty baseline, MEASURED-vs-CLAIMED grading. +- **`benchmarks/edge-latency/RESULTS.md`** — criterion 0.5, explicit host machine, low/median/high brackets, contention caveat, host≠ESP32 separation, steady-state-vs-cold-start distinction. +- **Rust micro-bench:** criterion benches already exist in both crates (`wifi-densepose-nn/benches/`, `wifi-densepose-train/benches/`). + +What a credible "beyond-SOTA" claim requires (the bar that survives "prove it"): +1. **One locked accuracy definition** — PCK normalization (torso vs absolute vs bbox), keypoint convention (15 vs 17 COCO), and split (random / cross-subject / cross-environment) declared *before* the run. The retracted 92.9% died exactly because PCK normalization was unstated. +2. **A mean-pose / constant-output honesty baseline** on every split (already done in measurement (b) — a single-subject near-static set scored 95.9% torso-PCK@20 with a *constant* pose). Any claim must beat this. +3. **MEASURED-vs-CLAIMED grading** per number, with the exact command and raw-JSON path committed. +4. **Cross-domain, not just in-domain.** In-domain PCK is saturated and uninformative; the defensible claim is on cross-subject/cross-environment, where the frontier is 12–44% PCK@20. + +--- + +## 3. Proposed benchmark-suite design + +A two-part suite (`wifi-densepose-train` accuracy harness + `wifi-densepose-nn` latency harness), both committing raw JSON + a graded RESULTS.md. + +### 3.1 Accuracy harness (`wifi-densepose-train`) + +- **Metric module with one canonical PCK** (parameterized: `{torso, bbox, absolute}` normalization × threshold × keypoint-map), so a single function scores WiFlow-STD-repro, MM-Fi/AetherArena, and a GraphPose-Fi re-run identically. Lock the default to **torso-PCK@20 on 17-kp COCO** and *always* also print standard-PCK to expose the gap. +- **Fixed datasets/splits:** (i) WiFlow-STD cleaned 360k (their split, for repro parity), (ii) MM-Fi P1 random + cross-subject + cross-environment (to line up against GraphPose-Fi 61.1/44.2/12.9 and the project's 81.63), (iii) ESP32 paired eval set when ≥2k multi-subject windows exist. +- **Mandatory honesty baselines** emitted every run: mean-pose, constant-output, and (for cross-domain) source-only. +- **Output:** raw JSON + a RESULTS.md table with MEASURED/CLAIMED grades, mirroring `benchmarks/wiflow-std/RESULTS.md`. + +### 3.2 Latency/size harness (`wifi-densepose-nn`) + +- **Matrix:** `{torch fp32 (ref), ONNX fp32, ONNX static-QDQ-conv-only int8, candle fp32}` × `{full 2.23M, half 843k, tiny 56k}` × `{batch 1, 64}`, criterion-timed, host declared. +- **Report:** disk size, batch-1 + batch-64 ms/window (median + low/high), and PCK@20 on the locked 10k-window subset, so latency and accuracy never get cited apart. +- **Caveat block up front:** host ≠ ESP32-S3/WASM3, host ≠ Hailo HEF. No host number is presented as the edge number. +- **CI gate:** commit the current medians as regression baselines; fail PRs that regress latency >X% or accuracy >Y pt. + +### 3.3 What counts as a defensible "beyond-SOTA" result + +A claim is citable only if **all** hold: (1) scored under a pre-declared metric/split, (2) beats the relevant published frontier number *on the same metric definition* (e.g. >61.1% standard-PCK@20 on MM-Fi random, or >12.9% on cross-environment), (3) beats the mean-pose honesty baseline, (4) raw JSON + exact command committed, (5) graded MEASURED. The single most valuable "beyond-SOTA" target is **cross-environment MM-Fi**, where the published bar (12.9% PCK@20) is low enough that a real win is both achievable and unambiguous. + +--- + +## 4. Gap table + +| Capability | Project current (graded) | Published SOTA (graded) | Proposed target | Data / hardware needed | +|---|---|---|---|---| +| In-domain 2D PCK@20 (self-collected) | 96.09–96.61% (MEASURED, RTX 5080, WiFlow-STD repro) | 97.25% claimed (WiFlow-STD, CLAIMED) | match within noise + own architecture | cleaned 360k dataset (have); already met | +| In-domain MM-Fi PCK@20 (torso-norm) | 81.63% torso-PCK (MEASURED-internal) | GraphPose-Fi 61.1% *standard*-PCK (PREPRINT) — **not comparable** | re-score both under **one** PCK def | MM-Fi P1 (have); unified metric harness (gap) | +| **Cross-subject MM-Fi PCK@20** | ~11.6% torso (MEASURED, the cliff) | GraphPose-Fi 44.2% standard (PREPRINT) | close gap via MAE pretrain + graph decoder | 1.3M heterogeneous CSI corpus (ADR-150/152 §2.3), ViT-Small encoder | +| **Cross-environment MM-Fi PCK@20** | untested-internal | GraphPose-Fi 12.9% standard (PREPRINT) | **beat 12.9% → cleanest beyond-SOTA win** | MM-Fi cross-env split + geometry conditioning (ADR-152 §2.1) | +| ESP32 CSI→pose (17-kp) | no run beats mean-pose baseline (MEASURED, measurement b) | n/a (no public ESP32 pose benchmark) | beat mean-pose on temporal split | ≥2k multi-subject/multi-position paired windows (gap) | +| Edge int8 size/accuracy | static QDQ conv-only 96.61% @ 2.53 MB; tiny 94.1% @ 295 KB fp32 (MEASURED) | no model-matched public number | **QAT-int8 on half model** (untested lever) | PyTorch QAT + QDQ export; RTX 5080 (have) | +| Edge CPU latency | ONNX fp32 3.2 ms/win b1 host (MEASURED) | n/a (model-specific) | committed criterion regression baseline | host bench (have); ESP32/Hailo on-hardware (gap) | +| Hailo HEF edge inference | none in-tree (gap) | n/a | first MEASURED HEF latency | Hailo compile toolchain + Pi cluster (have hardware, CLAUDE.local.md) | +| Foundation encoder (MAE) | recipe adopted, untrained (ADR-152 §2.3) | UNSW: log-linear cross-domain scaling on *classification* (PREPRINT) | pose-transfer validation (hypothesis today) | 1.3M-sample corpus aggregation (priority per F3) | + +--- + +## 5. Sources (graded) + +| Source | Type | Grade | Used for | +|---|---|---|---| +| GraphPose-Fi, arXiv [2511.19105](https://arxiv.org/abs/2511.19105) | preprint | PREPRINT; table numbers MEASURED-in-study (fetched + quoted) | RQ1 MM-Fi frontier (61.1/44.2/12.9 PCK@20, 160.6/210.5/302.7 mm) | +| WiFlow-STD, arXiv [2602.08661](https://arxiv.org/abs/2602.08661) + [DY2434 repo](https://github.com/DY2434/WiFlow-WiFi-Pose-Estimation-with-Spatio-Temporal-Decoupling) | preprint+code | numbers CLAIMED; artifacts MEASURED; **project repro 96% MEASURED** | RQ1/RQ2/RQ3 | +| PerceptAlign, arXiv [2601.12252](https://arxiv.org/abs/2601.12252) | preprint + MobiCom'26 acceptance | CLAIMED numbers; failure mode corroborated | RQ1/RQ2 geometry conditioning | +| UNSW MAE, arXiv [2511.18792](https://arxiv.org/abs/2511.18792) | preprint | ablations MEASURED-in-study; pose transfer = hypothesis | RQ2 MAE recipe | +| DT-Pose, arXiv [2501.09411](https://arxiv.org/abs/2501.09411), OpenReview [aPnLQ6WfQQ](https://openreview.net/forum?id=aPnLQ6WfQQ), [code](https://github.com/cseeyangchen/DT-Pose) | preprint+code (ICLR'25) | exact MPJPE table NOT verified here — flagged | RQ2 masked-pretrain + topology | +| Person-in-WiFi-3D, [CVPR 2024](https://openaccess.thecvf.com/content/CVPR2024/html/Yan_Person-in-WiFi_3D_End-to-End_Multi-Person_3D_Pose_Estimation_with_Wi-Fi_CVPR_2024_paper.html) | peer-reviewed | MEASURED (91.7/108.1/125.3 mm); own dataset | RQ1 3D multi-person frontier | +| ONNX Runtime quantization [docs](https://onnxruntime.ai/docs/performance/model-optimizations/quantization.html) | vendor docs | VENDOR | RQ3 PTQ/QAT mechanics | +| NVIDIA TensorRT QAT [blog](https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/), [Ultralytics](https://www.ultralytics.com/glossary/quantization-aware-training-qat) | vendor/blog | BLOG/VENDOR; general, not CSI-specific | RQ3 QAT>PTQ context | +| Onboard Optimization survey, arXiv [2505.08793](https://arxiv.org/pdf/2505.08793) | preprint | PREPRINT | RQ3 on-device optimization landscape | +| In-tree `benchmarks/wiflow-std/RESULTS.md`, `benchmarks/edge-latency/RESULTS.md`, ADR-150, ADR-152, ADR-015 | internal MEASURED | MEASURED-internal | grounding, all RQs | + +**Unverified / flagged:** DT-Pose exact MM-Fi MPJPE table not extracted at primary-source precision (abstract-level only). GraphPose-Fi parameter count not reported in the paper. WiFlow-STD/PerceptAlign accuracy numbers are author-self-reported preprints. No CSI-pose-specific QAT benchmark exists in the public literature — the QAT recommendation rests on general (non-CSI) vendor/blog evidence. diff --git a/api-docs/user-guide.md b/api-docs/user-guide.md index fc62b728..d9542a6e 100644 --- a/api-docs/user-guide.md +++ b/api-docs/user-guide.md @@ -522,6 +522,25 @@ Base URL: `http://localhost:3000` (Docker) or `http://localhost:8080` (binary de | `GET` | `/api/v1/mesh` | ADR-110 fleet-wide mesh sync map ([iter 29](adr/ADR-110-esp32-c6-firmware-extension.md)) | `{"nodes":{"9":{...},"12":{...}},"total":2}` | | `GET` | `/api/v1/nodes/:id/sync` | Single-node mesh sync snapshot (or 404) | `{"offset_us":1163565,"is_leader":false,...}` | | `GET` | `/api/v1/mesh/metrics` | ADR-110 mesh state in Prometheus exposition format ([iter 36](adr/ADR-110-esp32-c6-firmware-extension.md)) | `wifi_densepose_mesh_offset_us{node="9"} 1163565\n…` | +| `GET` | `/api/field` | ADR-262 P3 — latest **signed RuField `FieldEvent`s** from the live sensing cycle, plus the signer pubkey + a `dev_signing_key` flag. Only egress-safe (P1/P2) events are surfaced; identity/biometric (P4/P5) and raw (P0) are held edge-local | `{"spec":"rufield","signer_pubkey_hex":"…","dev_signing_key":true,"events":[…]}` | + +### RuField surface (ADR-262 P3) + +RuView's live WiFi-CSI sensing now also speaks the standalone **RuField MFS** wire format. Each governed sensing cycle is converted (via the `wifi-densepose-rufield` anti-corruption bridge) into a **signed** `FieldEvent` (`Modality::WifiCsi`, ed25519 `ProvenanceRef`) and surfaced on two additive endpoints: + +- `GET /api/field` — the most recent signed events (JSON). +- `GET /ws/field` — a WebSocket that streams each cycle's signed event (mirrors `/ws/sensing`). + +```bash +curl -s http://localhost:3000/api/field | python -m json.tool # latest signed FieldEvents +python -c "import asyncio,websockets; asyncio.run((lambda: websockets.connect('ws://localhost:8765/ws/field'))())" # stream +``` + +Privacy is fail-closed: only egress-safe **P1/P2** events leave the box — raw (P0) and identity/biometric/aggregate (P3–P5) cycles are held **edge-local** and never appear on these endpoints; a no-presence cycle emits **no event**. + +**Signing key:** the surface signs with a **dedicated dev/sensing key**, seeded from `WDP_RUFIELD_SIGNING_SEED` (a 64-char hex string or a ≥32-byte value); when unset it falls back to a deterministic dev default and logs a `WARN` (the `dev_signing_key` flag in `/api/field` reflects this). This is a standalone key pending the ADR-262 §8 Q1 key-ownership decision — set `WDP_RUFIELD_SIGNING_SEED` for any real deployment. + +> **Honesty (ADR-262 §0/§6):** this is real plumbing on a live endpoint, **not an accuracy claim.** It is the single-link CSI sensing with its existing caveats (no validated room-coordinate accuracy — positions are the "strongest field peak", not calibrated triangulation). ### Example: Get fleet mesh state (ADR-110) @@ -1842,6 +1861,23 @@ node scripts/eval-wiflow.js \ --data data/paired/*.jsonl ``` +> **Model format boundary:** `train-wiflow-supervised.js` produces the +> JavaScript WiFlow model `wiflow-v1.json`. There is currently no supported +> command that converts that JSON model into the sensing server's binary RVF +> container, and renaming the file to `.rvf` does not convert it. Use the JSON +> model with the JavaScript evaluation/inference tools. To train a model that +> the Rust sensing server can load, use its native training path, which writes +> RVF directly: +> +> ```bash +> cargo run -p wifi-densepose-sensing-server --release -- \ +> --train --dataset data/mmfi --dataset-type mmfi \ +> --epochs 100 --save-rvf models/room-model.rvf +> ``` +> +> The camera+CSI paired JSONL workflow and the native RVF trainer are separate +> pipelines today. A JSON-to-RVF exporter is future work. + **Evaluation protocol matters.** Use `eval-wiflow.js` (torso-normalized PCK@20, the metric comparable to published WiFi-pose results) on a temporal hold-out, and sanity-check that predictions actually vary across frames diff --git a/api-docs/vendor-rf-providers.md b/api-docs/vendor-rf-providers.md new file mode 100644 index 00000000..e7d07a7d --- /dev/null +++ b/api-docs/vendor-rf-providers.md @@ -0,0 +1,60 @@ +# ADR-270 Vendor RF Providers + +RuView exposes a capability-safe Rust provider layer for vendor sensing and RF +telemetry. It never converts RSSI, occupancy, location or network inventory into +complex CSI. + +## API + +- `GET /api/v1/rf/vendors` — all provider descriptors and access states. +- `GET /api/v1/rf/vendors/latest` — latest validated event per vendor. +- `GET /api/v1/rf/vendors/:vendor/latest` — latest event for one stable vendor ID. +- `POST /api/v1/rf/vendors/:vendor/events` — ingest the vendor's documented + sidecar/webhook payload through its strict provider decoder. This `/api/v1/*` + route uses the server's bearer-token policy when configured. + +Stable IDs are `origin_ai`, `plume`, `mist`, `netgear`, `electric_imp`, +`rf_solutions`, `linksys`, `luma`, `google_nest`, and `wifigarden`. + +## Deterministic simulator + +```bash +cd v2 +cargo run -p wifi-densepose-hardware --bin vendor-rf-sim -- \ + --vendor plume --frames 100 --output plume.jsonl + +# Stream canonical synthetic events to the sensing server UDP port. +cargo run -p wifi-densepose-hardware --bin vendor-rf-sim -- \ + --vendor mist --frames 100 --udp 127.0.0.1:5005 --realtime +``` + +Supported simulator names are `origin-ai`, `plume`, `mist`, `netgear`, +`electric-imp`, `rf-solutions`, `luma`, and `google-nest`. Linksys is refused +because its sensing service is discontinued. Wifigarden is refused until a +contracted event schema exists. + +Every synthetic event includes `synthetic: true`, a deterministic sequence and +timestamp, and a source ending in `-sim-01`. + +Canonical UDP JSON is accepted only when `synthetic: true`. Live vendor payloads +must use the HTTP ingestion route so provider-specific schemas, metric allowlists, +access states and bounds cannot be bypassed. + +## Live/provider payloads + +Provider decoders are strict, bounded and reject unknown schema fields. Origin +paths and credentials are supplied by the commercial contract. Plume uses a +read-only allow-listed OVSDB request plan. Mist and NETGEAR configurations use +regional HTTPS endpoints with redacted tokens. Electric Imp, RF Solutions and +Luma accept only allow-listed scalar metrics. Google Nest remains network-only. + +Credentials are never embedded in fixtures or descriptors. Linksys returns +`Unsupported`; Wifigarden returns `ContractRequired`. These are usable, +test-covered provider outcomes—not simulated integrations. + +## Hardware honesty + +All descriptors remain `hardware_validated: false` until exact hardware/cloud +versions, lawful access, repeatable captures, calibration where applicable, and +fixture publication rights have been verified. Passing the simulator and API +tests validates RuView software only.