Status: PLAN / DESIGN + Phase-1 scaffold shipped behind a flag (default OFF). Needs the owner's explicit go-ahead on the approach before any wiring into live agents — this stores biometric face data of real people.
Date: 2026-07-10
Author track: kai-engineer
Flag: KAI_FACE_RECOGNITION (env, default OFF). Nothing in this subsystem runs unless it is =1.
Companion flag it rides on: KAI_VIDEO_MODE (the camera path added to oracle.html / the Gemini-Live bridge on 2026-07-10, also default OFF).
---
The agents (Leo, KAI, etc.) should recognize people. When Leo sees the owner (via the call camera
feed / any camera), he learns that face, and next time he sees them — anywhere — he knows it's them.
Same for anyone he sees or talks with: he builds a profile of that person (who they are, what they
talked about, memories made together) to know them better over time. This profile/memory is **shared
across all agents** — any agent can access it and explore the memories they made with that person. And
when multiple agents see/recognize the same person, they confirm identity AI-to-AI, which raises a
confidence score against that face+person so recognition gets better over time.
---
This is feasible and, importantly, it slots into infrastructure that already exists — we are not
building person-memory from scratch, we are adding a face modality to a person-memory system the
fleet already runs.
Three facts make this low-risk to design (higher-risk to *enable* — see Privacy, §7):
1. There is already a biometric enrollment pattern to mirror. shared/voice-biometrics.mjs +
state/biometric_profiles.json + per-person .npy "DNA" files in state/dna_signatures/ +
a Python helper shared/vocal_dna.py (librosa → ~60-dim L2-normalized voiceprint, cosine match,
1-to-many identify). Face recognition is the same shape with a different embedder. We copy the
proven pattern rather than invent one.
2. There is already a shared, cross-agent person store. Every agent (Leo, KAI, Groq, Gemini,
Claudey, X, Oracle — separate Node processes) already reads/writes the same on-disk stores:
transcripts.db (user_profile_memories, entity_facts, user_facts, user_fingerprints,
user_profiles) via shared/transcript-memory.mjs, and the structured fact layer
shared/user-warehouse.mjs (keyed by Discord user ID, alias-resolved). "Shared between all for
memory" is already how the fleet works — we attach faces to the *same person key*.
3. There is already a camera frame source. As of 2026-07-10, oracle.html VIDEO MODE v1 captures
the call camera and sends {type:'frame', data:<base64 jpeg>} every ~2s over the existing
/ws/voice socket → server → OracleLiveVoiceSession.sendFrame() → gemini-live-bridge.sendFrame()
(realtimeInput.video), all gated by KAI_VIDEO_MODE (default OFF). That server-side frame choke
point is exactly where a face-ingest hook belongs — one place, one guarded call.
The unifying design decision: do not create a parallel identity namespace. Key face profiles by
the same canonical person id the fleet already uses (Discord user id, alias-resolved via
user-warehouse.resolveAlias). The moment a face is tied to person P, that face is automatically
linked to P's voice DNA, P's user_facts, and P's transcript memories — i.e. it is *already*
"shared across all agents and all the memories they made with that person," for free.
---
┌──────────────────────── CAPTURE (browser, KAI_VIDEO_MODE) ────────────────────────┐
│ oracle.html getUserMedia(video) → <canvas> → JPEG │
│ every ~2s: ws.send({type:'frame', data:<base64 jpeg>}) over /ws/voice │
└───────────────────────────────────────┬───────────────────────────────────────────┘
▼
┌──────────────────────── SERVER frame choke point ────────────────────────────────┐
│ command-center-server.mjs ws 'frame' handler → OracleLiveVoiceSession.sendFrame │
│ │
│ ── PHASE-2 HOOK (guarded, no-op when flag OFF) ── │
│ if (KAI_FACE_RECOGNITION==='1') personRecognition.ingestFrame(base64Jpeg, { │
│ agent: session.botName, channelId, sessionUserId }) // fire-and-forget │
└───────────────────────────────────────┬───────────────────────────────────────────┘
▼
┌──────────────────── FACE PIPELINE (shared/person-recognition.mjs) ────────────────┐
│ 1. detect + embed: face_dna.py --embed <jpeg> → 512-d L2-normed vector (+ bbox, │
│ det-score). 0 faces → drop. >1 → take most-central/largest. │
│ 2. match: cosine vs every known person's stored embeddings │
│ best ≥ MATCH_HI → recognized(personId) │
│ MATCH_LO..HI → tentative (needs corroboration) │
│ < MATCH_LO → unknown → provisional person (unlabeled) │
│ 3. record sighting: append {ts, agent, channelId, cos, det} to person.sightings; │
│ store the embedding (cap N per person, keep most diverse) │
│ 4. confidence: recompute person.confidence from #embeddings, #sightings, │
│ #distinct agents, and cross-agent confirmations (§4) │
└───────────────────────────────────────┬───────────────────────────────────────────┘
▼
┌──────────────────── SHARED PERSON STORE (all agents read/write) ──────────────────┐
│ state/face_profiles.json ← canonical per-person face record (§3) │
│ state/face_signatures/<id>/ ← that person's embedding vectors (.npy / .json) │
│ keyed by the SAME person id as user-warehouse / transcripts / voice DNA │
└───────────────────────────────────────┬───────────────────────────────────────────┘
▼
┌──────────────────── RETRIEVAL (mid-conversation, §5) ─────────────────────────────┐
│ leo.mjs / kai.mjs / oracle-live-voice: │
│ const who = personRecognition.whoIsOnCamera(channelId) │
│ → {personId, name, confidence, lastSeen, sharedNotes, recentMemories} │
│ → injected into the system/context so the agent greets by name + references │
│ past memories ("good to see you again, Ryan — last time we talked about …") │
└────────────────────────────────────────────────────────────────────────────────────┘
**Face capture → detection + embedding → person store → match by similarity → confidence that rises with
sightings and cross-agent confirmations.** Each stage above maps 1:1 to that requirement.
---
One canonical record per person, keyed by the fleet's existing person id. New file
state/face_profiles.json (mirrors state/biometric_profiles.json), with heavy embedding vectors
offloaded to state/face_signatures/<personId>/ (mirrors state/dna_signatures/).
// state/face_profiles.json
{
"schemaVersion": 1,
"people": {
"<personId>": { // SAME id as user-warehouse / transcripts (Discord snowflake,
// or "prov:<uuid>" for an as-yet-unlabeled stranger)
"personId": "<personId>",
"displayName": "Ryan", // null until known / labeled by an agent or the owner
"handles": ["nastermodx"], // known aliases (kept in sync w/ user-warehouse.resolveAlias)
"provisional": false, // true = a stranger we can re-recognize but haven't named
"firstSeen": "2026-07-10T19:50:15.514Z",
"lastSeen": "2026-07-10T20:14:02.010Z",
"embeddings": [ // vectors live in face_signatures/<id>/; this is the manifest
{ "file": "0001.npy", "dim": 512, "model": "insightface:buffalo_l",
"capturedAt": "…", "agent": "Leo", "detScore": 0.94 }
],
"sightings": [ // capped ring buffer (e.g. last 200), pruned by retention
{ "ts": "…", "agent": "Leo", "channelId": "…", "cos": 0.71, "detScore": 0.94 }
],
"confidence": 0.0, // 0..1, see §4 (identity confidence for this face↔person link)
"crossAgentConfirmations": [ // AI-to-AI agreements that raised confidence (§4)
{ "ts": "…", "byAgent": "KAI", "matchedPersonId": "<personId>", "cos": 0.69, "delta": +0.05 }
],
"notes": {
"shared": [ // notes any agent can read/append — "shared between all"
{ "ts": "…", "agent": "Leo", "text": "Owner. Prefers to be called Ryan." }
],
"perAgent": { // an agent's own private impressions (still readable by all)
"Leo": [ { "ts": "…", "text": "Talked about the KAIVERSE terminator work." } ]
}
}
}
}
}
Storage decision — JSON registry + .npy sidecars, NOT a new SQLite table. Justification:
per-person .npy (biometric_profiles.json + dna_signatures/). Mirroring it means the face path
reads exactly like the voice path — same mental model, same enrollment ergonomics, same backup story.
the fleet's established cross-process shared medium (the whole state/ "shoebox"). Writes use
temp-file + atomic rename (the pattern already used in shared/metrics-store.mjs,
kaiverse-life.mjs, etc.) to survive a crash mid-write. Concurrency is low (a sighting every ~2s per
active call), so a short advisory lock + last-writer-merge is sufficient; no DB contention.
.npy sidecarsexactly like voice DNA, keeping the JSON small and diffable.
(what you talked about) already lives in transcripts.db / user_facts keyed by that id. We do not
duplicate it — retrieval (§5) joins on the id. This is what makes the profile "shared across all
agents" without any new sharing mechanism.
*(If volume ever grows past what JSON comfortably holds — thousands of people, millions of sightings —
the migration path is a face_profiles/face_embeddings table inside transcripts.db alongside
user_profile_memories. Deferred; not needed for Phase 1–3.)*
---
Agents don't message each other synchronously about a face; they **corroborate through the shared
store**, which is exactly how the fleet already coordinates. The confidence on a face↔person link is a
function of independent agreement:
confidence = clamp01(
0.25 * f(numEmbeddings) // more enrolled views of this face
+ 0.25 * f(numSightings) // seen more often
+ 0.30 * f(numDistinctAgents) // ← the cross-agent term: how many DIFFERENT agents agreed
+ 0.20 * meanMatchCosineAboveHi ) // how strong the matches were
P (cos ≥ MATCH_HI), and agent A had already labeled/recognized P, B writes a crossAgentConfirmation for P. A
confirmation from a distinct agent applies a positive delta (with diminishing returns — the 5th
agent to agree adds less than the 2nd). This is the "multiple agents see the same person → confidence
goes up" mechanic.
P strongly *and* Q non-trivially (an ambiguous face), or two agents attach conflicting displayNames to the same embedding cluster, confidence is
lowered and the record is flagged needsReview rather than silently guessing. Better to be unsure
than confidently wrong about who someone is.
prov:<uuid> with a low confidence. Once anagent (or the owner) attaches a real name/handle, the provisional record is merged into the canonical
person id (embeddings + sightings carried over), and future cross-agent hits accrue to the named
person.
MATCH_HI, MATCH_LO, per-agent decay, max embeddings/person)— tuned during Phase 2 with the owner, the same eyeball-iterate loop the visual work uses.
---
The retrieval surface is a single read the voice/text paths can call cheaply:
// shared/person-recognition.mjs
whoIsOnCamera(channelId) -> null | {
personId, name, provisional, confidence, lastSeen,
sharedNotes: [...], // notes.shared, newest first
recentMemories:[...] // JOINED from the shared stores by personId:
}
recentMemories is assembled by joining on the person id into the memory the fleet already keeps:
user-warehouse.listFacts(personId) (home, preferences, relationships) + a recall query against
transcript-memory (user_profile_memories / transcript_fts) for recent topics with that person.
Nothing new is stored to answer "what did we talk about" — it's already there under the same key.
Hook points (all Phase 2+, guarded, additive — never on the audio pacer path):
oracle-live-voice.mjs / leo.mjs, when a call starts or a face is first recognized on a channel, fetch whoIsOnCamera(channelId) and prepend a short, budgeted context line
to the model turn: *"You can see Ryan on camera (confidence 0.82). Recent: KAIVERSE terminator work,
prefers 'Ryan'."* The agent then naturally greets by name / references the memory. **This is a
system-context injection only — it does not touch leoPacedFeed, the throttle, or the audio timing.**
whoIsOnCamera/whoIs(personId) read can enrich a reply when arecognized person is present. Optional, later.
lattice-bridge (best-effort, non-blocking) so KAI's associative recall also "knows the face,"
consistent with how user facts are already mirrored.
---
Context: the box is a limited Windows laptop (RTX 4050, but treat CPU as the floor). The existing
biometric path already shells out to Python (vocal_dna.py, via execSync('python …')) using
numpy/librosa/scipy. So a Python face helper is the *native* fit — same enrollment ergonomics,
same .npy artifacts, same call pattern.
Options considered:
insightface (ArcFace, buffalo_l) on onnxruntime ← recommended | ArcFace, 512-d | Python helper face_dna.py, mirrors vocal_dna.py | CPU-feasible; SOTA-grade discrimination; pip install insightface onnxruntime; auto-downloads model; detection + embedding in one lib; produces .npy exactly like voice DNA | first-run model download (~300 MB); onnxruntime CPU install to verify on the box |face_recognition (dlib) | 128-d | Python helper | very common, simple API | dlib build on Windows is painful (cmake/VS build tools); heavier to install |mediapipe face embedder | ~192-d | Python helper | pip-installable, CPU-light, small model | slightly lower discrimination than ArcFace; API churn across versions |@vladmandic/face-api (tfjs) in Node | 128-d | in-process in the fleet | no Python; keeps it all Node | tfjs-node native bindings finicky on Windows; adds a big dep to every bot |oracle.html | 128–512-d | the browser, on the already-captured frame | strong privacy win — raw face never has to leave the browser; only the embedding vector crosses the wire; zero server model install | logic split across browser+server; only covers the dashboard camera path, not other cameras |Recommendation: **Option A (InsightFace ArcFace buffalo_l, 512-d, onnxruntime-CPU) as the
server-side embedder**, because it is the most consistent with the proven vocal_dna.py pattern
(Python helper → .npy → cosine match → 1-to-many identify), CPU-feasible, and the most discriminative
of the easy-install options. Strongly consider Option E as a Phase-3 privacy upgrade: move
embedding into the browser so raw biometric imagery never needs to reach the server for the dashboard
camera — the server would then store only vectors. A + E compose (same 512-d space if the browser uses
the same model family; otherwise keep them as separate model tags on the embeddings).
Phase-1 reality: the scaffold ships face_dna.py as a guarded stub — it documents the InsightFace
path and only attempts a real embed if the flag is on *and* the library imports successfully; otherwise
it returns a clear "embedder-not-installed" status and the pipeline no-ops. **No model is downloaded and
no dependency is installed until the owner approves the approach.**
---
This subsystem stores biometric identifiers (face embeddings) of real people. Treat it accordingly.
KAI_FACE_RECOGNITION (env, default OFF) and rides on KAI_VIDEO_MODE (also default OFF). With the flag unset, the module is inert: no capture, no embed,
no store, no retrieval. Phase 1 changes zero live-agent behavior.
himself/known people), not silent background harvesting of every face on camera. Strangers may be
stored only as provisional, unlabeled re-identification records, and only while the flag is on;
naming a person is an explicit step. A consent/notice string should be surfaced in the dashboard when
the camera + recognition are both on. *(Owner to confirm the consent UX before Phase 2.)*
state/ and are read only by the localfleet. No face data, embedding, image, or profile is ever sent outside the fleet, posted to
Discord, put in a URL, or included in any external API payload beyond the already-consented
Gemini-Live video frames that KAI_VIDEO_MODE itself sends.
base64 image data, never an embedding dump, never a face crop path in a URL. (Same discipline as the
"no secrets in logs" rule.)
.env values.forgetPerson(personId) API removes the JSON record and the .npy sidecars (mirrors voice-biometrics deletion expectations). Retention prunes old sightings.
needsReview over aconfident wrong identification* (§4). An agent should never assert an identity below a comfortable
confidence threshold — "I think this might be Ryan" vs. silently addressing the wrong person.
UX, and (c) enabling the flag, before any wiring into live agents. Flagged again in the report.
---
Phase 1 — Foundation (this deliverable; shipped, flag OFF, non-breaking).
shared/person-recognition.mjs: the shared person-profile store (schema + atomic read/write API) and a matchEmbedding stub. All exports are **no-ops returning
null/empty when KAI_FACE_RECOGNITION!=='1'**.
shared/face_dna.py: guarded embedder stub (documents the InsightFace path; real embed only iflib present).
guarded ingest hook and the retrieval hook are documented here for Phase 2 but not applied.
Phase 2 — Live recognition + cross-agent confirmation (needs owner go-ahead).
face_dna.py stub with the real InsightFace embed + the matchEmbedding cosine/identify implementation.
command-center-server.mjs / OracleLiveVoiceSession.sendFrame) — fire-and-forget, no-op when flag OFF.
voice-biometrics enrollment).Phase 3 — Memory retrieval + agent greeting + privacy upgrade.
whoIsOnCamera retrieval into the voice/text context (system-context injection only; never thepacer). Agents greet by name and reference past memories joined from the existing shared stores.
lattice-bridge.server for the dashboard camera path.
Phase 4 — Hardening.
needsReview queue + owner tooling to merge/split/forget identities. liveness heuristic already in vocal_dna.py.
---
Phase 1 (created):
shared/person-recognition.mjs — shared person-profile store + gated read/write API + match stub.shared/face_dna.py — guarded face-embedder stub (InsightFace path documented; inert without lib).state/face_profiles.json — created lazily on first write (not committed with data).state/face_signatures/ — created lazily (per-person .npy sidecars).Phase 2+ (documented, NOT yet edited):
command-center-server.mjs — one guarded ingest line at the type:'frame' handler.shared/oracle-live-voice.mjs / bots/leo.mjs / bots/kai.mjs — guarded retrieval context injection.---
1. Approve InsightFace (ArcFace buffalo_l, onnxruntime-CPU) as the embedder? (vs. dlib / mediapipe
/ browser-side.) This is the only real dependency install.
2. Consent UX for enrollment + a camera-on/recognition-on notice — what do you want it to say/show?
3. Stranger policy: store unknown faces as provisional re-id records while the flag is on, or only
ever store people you explicitly enroll?
4. Green light to wire the guarded hooks into live agents once Phase 1 is reviewed? (Until then this
subsystem is dormant.)