Skip to content

PRISM

Six capabilities. One production reality.

PRISM helps you test before launch, understand real behavior after launch, and turn important failures into changes you can validate.

PRISM · run_9f3aOutcome mismatch
User intent
Exchange order #4821
Agent state
intent = refund
Tool call
process_refund → 200 OK
Outcome
Refund issued, expected exchange

A request can return 200 OK and still fail the user.

Connect PRISM

Connect the app you already built

Use a verified SDK, framework, OpenTelemetry path, connector, import, or API.

ConnectionReceiving
SDK
PRISM SDK
Standard
OpenTelemetry
Frameworks
LangChain · LangGraph
Import
Supported historical activity

Every supported connector is available on every plan, including Free.

The artefact

What a month of this actually produces

Not a dashboard you are meant to check. A score with six weighted dimensions under it, every one of them traceable to the rules that moved it, and a delta against last month that somebody can be held to.

  • 78/100

    Reliability score

    +6 vs May

  • 8.4%

    Failure rate

    was 12.7%

  • 74

    Satisfaction proxy

    was 65

  • 1.2%

    Guardrail trigger rate

    was 1.8%

  • 0

    High-severity incidents

    3 medium, resolved

Scorecard · six weighted dimensionsOverall 72 → 78

Every dimension moved up. Improvement velocity moved most (+12) because five fixes shipped against two the month before — the score is measuring whether the team is getting faster, not just whether the agent is behaving.

  • Task success & resolutionweight 25%7682+6

    KB fixes lifted completion on top intents

  • Correctness & groundednessweight 20%7279+7

    Fewer stale-content answers after KB refresh

  • User friction & satisfactionweight 15%6574+9

    Loops down; still the weakest dimension

  • Stability & error-free executionweight 10%8588+3

    Timeout retry cut tool errors sharply

  • Guardrails & policy adherenceweight 20%8791+4

    Zero high-severity events

  • Improvement velocityweight 10%4961+12

    5 validated fixes vs 2 in May

Green at 85. The weights are fixed per client, stated in every report, and change only at quarter boundaries — a score whose weights move is a score you can flatter.

Figures from a sample FinVault reliability report. Design-partner data, synthetic and illustrative of real deliverables.

Reliability score · 90 daysweekly

72 to 78 in a month, and 85 is Green. The two backlog items already in flight are projected to close that gap on their own.

KB remediation wave shipped · Jun 978

The marker is the remediation wave that shipped on June 9. An inflection with no cause named on it is a chart asking to be believed.

PRISM weekly score snapshots

Volume and completion, together+23% over the quarter

Traffic up 23% over the quarter and completion up 3.5 points at the same time. Either number alone proves nothing; the pair is the only shape that means growth is producing better outcomes rather than more failures.

  • 39.1K83.6%
    April
  • 44.5K85.9%
    May
  • 48.2K87.1%
    June

Traffic up 23% and completion up 3.5 points at the same time. Either number alone proves nothing; the pair is the only shape that means more usage is producing better outcomes rather than more failures.

PRISM session telemetry

The loop

48,210 sessions in, five validated changes out

The narrowing is the product. Anything that turns production traffic into alerts has made your week worse; this ends in a small number of changes that shipped, with a human signing off at every gate and a measurement afterwards proving each one worked.

  1. 4,050

    Failures clustered

  2. 3,822

    Root causes classified

  3. 14

    Backlog items opened

  4. 9

    Fixes proposed

  5. 7

    Human approved

  6. 5

    Shipped & validated

Human escalations avoided

Escalation rate fell 9.2% → 7.6% of 48,210 sessions = 771 avoided escalations

~154 support hours · ~$5,400

Engineering triage eliminated

Failure clustering + root-cause classification replaces manual log review

~48 engineering hours · ~$4,300

Completions recovered

Refund KB fix recovered ~380 sessions; refusal tuning adds ~120/month

Revenue-linked

~$9,700/month measurable value against $5,000/month program cost ~1.9× program cost. Deliberately conservative: placeholder loaded rates, measured effects only, and revenue from recovered completions excluded entirely until the per-completion value is agreed with your finance team in week one.

Underneath all of it

One session, held together

Every number on this page comes from the same structure: a production session kept as one record with its intent, its state, its tool calls and its outcome still attached to each other. Split those into four systems and none of the analysis above is possible.

Observealways on
coverage100% of production sessions
capturedTraces · evaluations · guardrail events
volume48,210 this month
Understanddaily
classifiedIntent, failure pattern, root cause
clustered4,050 failures → 5 patterns
flagged5.6% below confidence → human
Improveweekly
backlog14 items, owners and ETAs attached
proposed9 fixes · KB, prompt, PR
gateHuman approval requiredALWAYS
Provemonthly
shipped5 changes, re-measuredVALIDATED
score72 → 78 · +6
artefactExecutive report + Trust Pack

The unit of analysis. Not a log line, not a span, not a prompt — a session, with what the user wanted still attached to what they got.

Always on, not on request

Coverage is 100% of production sessions rather than a sample, which is the only reason a pattern can be counted instead of noticed. “214 sessions with this signature” is a sentence sampling cannot produce.

A human gate before anything ships

Knowledge and guardrail changes happen in-platform, prompt changes ship as versioned diffs with rollback, code changes ship as pull requests carrying their evidence. All three need explicit sign-off. PRISM never merges on its own.

Proof is a deliverable, not a tab

The month ends in a document: the score, the deltas, what shipped, what it moved, and the three decisions leadership has to make. Dashboards are the supporting evidence behind it.

Plan notes

Evaluator creation and management begin on Builder. End User Management and User Risk Profiles begin on Growth.

Start with one AI app

Connect a real workflow and find the first thing worth improving.