Research
Evidence-backed item

Product systems research · evaluation-backed

System

The Scriben system: from physical capture to governed action

A product systems report on how Scriben turns conversations into inspectable context, durable memory, and user-approved action, with a five-session evaluation of the intelligence layer.

Critical-term recall
84%
Speaker diarization
95%

System thesis

Scriben is a real-world context system, not a transcription interface with an agent attached. The product begins with a physical source event, carries that evidence through interpretation and memory, and turns it into work only at a visible, user-controlled action boundary.

This report connects three bodies of product research that are easy to misunderstand in isolation: the observable reliability requirements for ambient hardware and firmware, a multi-metric evaluation of conversation intelligence, and the context-to-action architecture behind Scriben Agent.

The public evidence layer contains five paired recordings totaling 155 minutes across product engineering, venture investing, informal overlapping speech, legal mediation, and a board meeting. Scriben reached 84% critical-term recall, 95% speaker diarization, 73% evidence-based speaker attribution, and 74% meeting-summary accuracy on the retained corpus.

The product implication is larger than any one score. Reliable real-world intelligence requires a chain of custody for context: a legible capture event, independently evaluated interpretation, source-linked memory, and review before the system changes the world on the user's behalf.

Procedural research poster for The Scriben system: from physical capture to governed actionSCRIBEN / PRODUCT SYSTEMSCONTEXT SYSTEMPHYSICAL EVENT → INSPECTABLE INTELLIGENCE → REVIEWABLE ACTION01CAPTUREDEVICE + FIRMWARE02UNDERSTANDEVALUATED INTELLIGENCE03REMEMBERSOURCE-LINKED CONTEXT04ACTAGENT + HUMAN REVIEWPROVENANCE / SEPARATE METRICS / DURABLE MEMORY / APPROVAL BOUNDARYSCRIBEN / CONTEXT / SYSTEM / 001

One context system, four technical boundaries

Scriben is engineered as a closed loop from a physical source event to a reviewable action. Each boundary has a different failure mode, so capture, interpretation, memory, and execution are designed and evaluated as separate layers.

  1. 01Device + firmware boundary

    Capture

    The system begins beyond the screen. Recording state, continuity, integrity, recoverability, consent, and user control define the public capability boundary; proprietary components, protocols, and thresholds remain undisclosed.

  2. 02Conversation intelligence

    Understand

    Scriben preserves critical terms, separates speakers, resolves identities from conversational evidence, and produces summaries whose coverage and unsupported claims can be scored independently.

  3. 03Context + memory boundary

    Remember

    Source-linked notes, speakers, decisions, and commitments become durable context with provenance. The system preserves the path back to evidence instead of treating generated interpretation as unquestioned truth.

  4. 04Agent + action boundary

    Act

    Scriben Agent can recover captured context, prepare documents and spreadsheets, and stage work across email, calendars, Slack, Linear, and CRM workflows. Externally consequential actions remain visible for user review.

The technical advantage is the integration: physical provenance, evaluated interpretation, durable memory, and human-controlled action operate as one inspectable system, not as a transcript handed to a generic agent.

The product systems problem

Most knowledge-work agents begin after the useful context has already been compressed, fragmented, or lost. They see a screen, a transcript, or a prompt, not the source event, the speaker structure, or the evidence behind a commitment.

That gap creates compounding failure. An interrupted capture can remove evidence; a merged speaker can change ownership; an unsupported summary statement can enter memory; and an agent can then execute a polished action from a false premise.

The system question is therefore end-to-end: how can a product preserve trustworthy context from the physical world, make each transformation inspectable, and still convert that context into useful work without turning uncertainty into autonomous side effects?

Architecture and evaluation method

At the capture boundary, Scriben defines reliability through observable outcomes: recording state is understandable, the source event remains intact through ordinary interruption, handoff can recover safely, and consent, retention, and user control remain legible. This report intentionally does not disclose component selection, sensor geometry, firmware architecture, protocols, thresholds, or power characteristics.

At the intelligence boundary, references were frozen before product outputs were scored. Critical-term recall used 204 checkable names, numbers, dates, and domain terms. Diarization, evidence-based attribution, and summary faithfulness were measured separately so one strong layer could not mask another failure.

Summary scoring used 153 frozen subjects at roughly one subject per 90 seconds. Coverage received credit only when the report retained a quotation from the product output; each unsupported statement incurred a three-point penalty. Model-based grading ran three times and reported the median.

At the memory boundary, interpreted context retains a path back to source evidence. At the action boundary, Scriben Agent converts notes and accepted commitments into discrete, reviewable artifacts and staged actions. Preparation remains separate from execution so outward side effects can be inspected before approval.

System evidence and current capabilities

The evaluation layer reached 84% weighted critical-term recall across 204 frozen names, numbers, dates, and domain terms. This is exact-match recall on the benchmark term set, not word error rate.

Speaker diarization reached 95% on sampled utterances, while evidence-based speaker attribution reached 73% without enrolled voices or a supplied participant list. The 22-point separation identifies an important system boundary: keeping voices distinct was materially easier than resolving those voices to people from conversational evidence alone.

Meeting-summary accuracy reached 74% across the four meeting sessions included in the aggregate. The continuous-crosstalk recording remains a separate acoustic stress case rather than part of that aggregate.

The current Scriben Agent implementation can recover note context and action items, prepare documents and spreadsheets, and stage work across email, calendars, Slack, Linear, and CRM workflows. Multi-action requests remain separated into reviewable work items rather than becoming one opaque approval.

Together, these results support a layered product architecture. Capture integrity, speaker structure, identity resolution, summary faithfulness, memory provenance, and action approval are treated as distinct control points before real-world context is allowed to drive downstream work.

Study design and evaluation evidence

Research questions

  1. RQ01

    Can Scriben preserve the names, numbers, dates, and domain terms that make a conversation operationally useful?

    Tests recall of checkable lexical evidence rather than treating transcript fluency as sufficient.
  2. RQ02

    Can Scriben keep speakers distinct and recover identities from conversational evidence without prior voice enrolment?

    Separates the acoustic task of diarization from the semantic task of attribution.
  3. RQ03

    Can Scriben produce summaries that cover the recorded subjects without introducing unsupported statements?

    Scores coverage and faithfulness together instead of relying on presentation quality or preference alone.

Evaluation protocol

  1. 01
    Paired capture

    Five conversations were recorded under matched conditions across in-room, remote, hybrid, and crosstalk settings.

  2. 02
    Frozen references

    Terms, subjects, sampled speakers, and accepted commitments were fixed before product outputs were scored.

  3. 03
    Metric separation

    Critical terms, diarization, attribution, and summary faithfulness were evaluated independently so one strength could not mask another failure.

  4. 04
    Boundary reporting

    Excluded sessions, unpublished measures, privacy transformations, and reproducibility limits remain part of the reported result.

Conversation recorded
155 min

Five sessions, 12–57 minutes each

Speakers per session
2–7

In-room, remote, hybrid, and crosstalk

Checkable evidence
204 + 153

Critical terms + summary subjects

Scriben aggregate scores reported in this evaluation
MeasureScriben
Critical-term recallWeighted exact match on names, numbers, dates, and domain terms84%
Speaker diarizationUtterances kept on a consistent, non-collapsed speaker label95%
Speaker attributionAnonymous labels resolved from evidence inside the conversation; no enrolment73%
Summary accuracyCoverage minus a three-point penalty per unsupported statement; four meeting sessions74%
Five-session evaluation set
SessionEvaluation setting
Product and engineering57 min · 7 speakers · group call
VC investor meeting25 min · 4 speakers · hybrid
Informal conversation32 min · 2–3 speakers · continuous crosstalk
Lawyer discussion29 min · 7 speakers · fully remote
Board meeting12 min · 4–5 speakers · one room

Interpretation boundary

  • The crosstalk session is an acoustic stress test and is not included in the reported 74% meeting-summary aggregate.
  • Two unnamed baseline approaches were evaluated under the same paired-capture protocol. Their identities and scores are intentionally omitted from this public report.
  • Action extraction, named-owner, latency, memory, recall, and helpfulness measures are withheld until their public scoring protocols are sufficiently specified.
  • Scoring code, frozen references, grader prompts, model versions, and raw outputs are not included in this public report.

The end-to-end Scriben system

Limitations

This five-session benchmark has no external adjudicator, confidence intervals, or statistical significance test. It reports performance on this corpus and these stored runs, not a population estimate.

Critical-term recall is exact-match recall on a curated set of 204 checkable terms, not word error rate. The reference terms were established from the retained transcripts and surrounding context rather than by an independently commissioned verbatim transcript, creating a possible consensus bias.

Device generation, firmware, app version, transcription model version, network conditions, grader model, prompts, and inter-rater agreement are not disclosed publicly. Product updates or different acoustics could change the results.

Ground truth, scoring code, and stored runs are retained by Scriben but are not published with this report. The aggregate scores cannot be independently reproduced from the public page alone.

No component selection, schematic, sensor geometry, firmware architecture, protocol, threshold, power characteristic, or patent-relevant implementation detail is disclosed. No public hardware reliability or field-performance benchmark is claimed here.

Scriben Agent remains in staged rollout. This report does not provide a public benchmark for end-to-end task completion, latency, connector coverage, partial-failure recovery, or adoption.

Safety, privacy and deployment considerations

Consent, retention, visible capture state, and user control belong inside the system boundary rather than being deferred to policy text after collection.

The investor session replaces people, organizations, figures, and subject matter with consistent fictional equivalents across every product output and reference. The remaining sessions are described as publicly available recordings.

Baseline identities and scores are withheld deliberately. The evaluation should be read as a characterization of Scriben on one retained corpus, not as a public competitive ranking.

An 84% critical-term recall score is not sufficient evidence for consequential automation without review. The deployment hierarchy remains: preserved source evidence first, structured interpretation second, durable context with provenance third, and externally consequential action only after the user can inspect what the system relied on.

Citation and links

Scriben Research. (2026). The Scriben system: From physical capture to governed action. Scriben product systems report.

No public paper, code, demo, or dataset is attached to this item.