Domain scenarios
The examples each demonstrate one part of mekik. The domain
scenarios do the opposite: each one is a small, realistic desk that uses
several features together, the way an application would. Each runs
offline. A scripted model stands in for the LLM (the same seam as the
--probe modes), the probe plays the client, and every step asserts the frame
stream it got back. So each scenario is also an integration test, and CI runs
all five.
They live under ts/examples/, one directory each, with a README:
cd ts
node examples/banking/banking.ts
node examples/insurance/insurance.ts
node examples/healthcare-triage/triage.ts
node examples/travel-booking/travel.ts
node examples/support-desk/support.ts
pnpm run examples:domain # all five; also part of `pnpm check`
The output follows the existing probes: the frames of each turn (→ a tool
call, ← its result, ⏸ a pause, ▦ a component, ✦ a skill), a ✓ line per
assertion, and a final ✅. A failed assertion exits 1 with the message.
At a glance
| Scenario | Graph | What it pins on the wire |
|---|---|---|
banking | router → accounts / history / payments (agent) → approve → execute money tools owned by their skills, two pauses in one node for a second approver, exactly-once execution, a genui-table history, a hidden tool that fails without a trace | |
insurance | intake (form pause) → assess | a genui-form interrupt, decision tools owned by a skill and offered only after load_skill, a sign-off pause that keeps them, a typed rejection reason |
healthcare-triage | intake (agent) → escalate (chips) → book (client tool, agent) | redacted identifiers on every frame, skill-held scoring and booking, an allowlisted client-declared skill, a chips-only pause, the page's calendar as a client tool |
travel-booking | router → search → compare → book (agent); cancel (agent) | a reconnect that replays exactly the missed frames with the unlocked tools intact, welcome.pending, a skill-held cancellation that runs once |
support-desk | router → billing / tech → handoff / chat | tag-scoped skills per route, per-node tool scoping, an MCP knowledge base, an A2A hand-off whose pause is relayed to the desk's human |
Skills in every scenario
Every scenario has its own §12 skill catalog (the skills app option, announced
in the skills frame on connect). In each catalog
a skill owns its tools: the entry
lists them next to its instructions (SkillEntry.tools), so a skill reads as
"instructions + tools", and the model cannot call a tool before it has read the
rules that govern it. The tools never leave the server; every probe asserts
the skills frame names none of them.
| Scenario | Skill (tag) | Tools |
|---|---|---|
| banking | wire-transfer-rules | check_transfer_limit, transfer_funds |
| banking | dispute-handling | open_dispute |
| insurance | water-damage-assessment (home) | check_coverage, approve_claim |
| insurance | rejection-letter (home) | check_coverage, reject_claim |
| healthcare-triage | triage-protocol (triage) | score_triage |
| healthcare-triage | appointment-booking (booking) | book_appointment |
| travel-booking | fare-rules (booking) | book_flight |
| travel-booking | cancellation-policy (cancellation) | cancel_booking (built per request, held with skillTools) |
| support-desk | refund-policy (billing) | issue_credit |
| support-desk | incident-runbook (tech) | escalate_to_specialist |
The probes record the tool list offered in every model round. Every scenario asserts that a held tool is absent before its skill loads and offered from the round after. Several also check that a premature call is refused with an observation and does not run. Beyond that, each scenario adds its own check:
- tag scoping across routes (support desk);
- the unlocked set surviving a pause or a reconnect (insurance, travel);
- redaction holding on the skill frames (healthcare).
Every tool is built once, at module level. One that needs the run (to mount a
card, or to key state by the conversation) reads the calling run's ctx with
toolContext(config) from its LangChain config. Only one tool in all five
scenarios is built per request: travel's cancel_booking refunds from the
conversation's booking in graph state, so the cancel node holds it under
cancellation-policy with runAgent({ skillTools }), and says why in a
comment.
Most agents use runAgent({ skills }). The support desk's tech node shows the
hand-wired form, because its MCP tools come pre-wrapped by withMcpTools:
withSkills(ctx, filter, { onLoaded }) names each entry's tools in the
load_skill observation and reports each successful load, and the loop offers
that skill's tools (mekik.skillTools) from the next round.
Banking
The payments node is a runAgent loop over a catalog whose skills own the
money tools:
const SKILLS: SkillEntry<StructuredToolInterface>[] = [
{
name: "wire-transfer-rules",
description: "Rules for sending money: fraud screening, limits, second approvers. Load before any transfer.",
instructions: "1. Screen every transfer with fraud_screen. …",
tools: [checkTransferLimit, transferFunds],
},
{
name: "dispute-handling",
description: "How to open a card or account dispute for a transaction the customer does not recognise.",
instructions: "Confirm the date, description and amount from the history, then open_dispute. …",
tools: [openDispute],
},
];
The node passes only its always-on tools (lookup_payee, fraud_screen,
flag_for_review) and skills: true. Balance and history tools stay
always-on in their own nodes. The probe asserts that
transfer_funds is not offered before its skill loads, and that a premature
call is refused without staging anything. It also checks that the skill
frame's seq comes before the transfer trace, and that loading
dispute-handling unlocks only open_dispute.
transfer_funds only stages the transfer; the approvals come next. Over the
$5,000 limit, the approval node calls mekik.approve twice with different
keys: the customer first, then a second approver. On the resume that answers
the second pause, the node re-runs from the top and the customer's answer
comes back from the journal, so the customer is not asked again. The probe
asserts that, plus one execute_transfer per transfer and a ledger debited
once.
The history is a genui-table mounted from inside the tool. The fraud check
runs with a show: false policy, so no frame for it reaches the wire. When
the service is down the tool throws, and runAgent hands the model Error from fraud_screen: fraud screening service timed out
as an observation. Following the skill, it calls flag_for_review instead of
transfer_funds, and no money moves.
Insurance
The claim form is the pause. mekik.approve with
ui: mekik.genui.form.ref({ fields }) and no actions mounts the form, and the
submitted values come back as the resume answer. The app's skills option
holds three skills: the skills frame announces them on connect (no
instructions), and the adjuster node offers only the home-tagged ones.
The adjuster is a runAgent loop. Only lookup_policy is always offered.
water-damage-assessment owns check_coverage and approve_claim, and
rejection-letter owns check_coverage and reject_claim. The probe asserts
four things:
- Before
load_skill, onlylookup_policyandload_skillare offered. - A premature
check_coveragecall is refused with an observation naming both skills that own it. The tool doesn't run and nothing is traced. - From the round after the
skillframe, the skill's tools are offered, and the other skill's decision tool never is. - Across a pause. Each decision needs a senior adjuster's sign-off (an
approvepolicy). On the resume, the earlier rounds replay from the journal, the next round is still offered the unlocked tools, and the decision runs exactly once.
A rejection carries a typed reason. reject_claim's zod schema pins code to
EXCLUDED_PERIL | POLICY_LAPSED | BELOW_DEDUCTIBLE, and
{ code: "EXCLUDED_PERIL", clause: "4.2(b)" } lands on the typed
claim-decision component.
Healthcare triage
The medical record number arrives as a verified claim
(StaticTokenAuthenticator, read with mekik.authClaims), never in the chat.
The intake tools use the same redaction technique as
sql-agent.ts --probe: a redact policy, here on runAgent. The model reads
the real MRN, date of birth and name, while the traces show «redacted». The
probe checks that no identifier appears on any frame: skill frames (which
carry only id, name, status and source) and a second tab's full
transcript replay included.
Both agents use skills, each scoped to its own tag. triage-protocol owns
score_triage, the red-flag rules, and appointment-booking owns the
server-side book_appointment. The calendar itself stays a
client tool.
This is also the scenario with
client-declared skills. The portal declares two in
hello.skills, and the server opts in with a clientSkills allowlist that
accepts plain-language and drops override-triage, a skill that would tell
the model to ignore red flags. The probe asserts both outcomes:
plain-languageis in the prompt and itsskillframe sayssource: "client";override-triagenever reaches the model, and loading it is an unknown-skill observation with no frame.
Client skills are per connection: another patient's portal, which declared
none, does not get plain-language.
An urgent score parks the run with mekik.choose (three chips, no form). The
booking slot comes from the page's own calendar: mekik.callClientTool(ctx, "open_calendar", …) parks on an interrupt carrying data.tool, and the page
answers { ok: true, result }. The booking agent then loads its skill, and
book_appointment runs once against the real MRN. An emergency answer mounts
a genui-alert and never opens the calendar.
Travel booking
Search, compare and book are separate nodes. The comparison is a
genui-table with a literal chunk id, so the replay after the pick
re-renders the same element. The book and cancel nodes are runAgent
loops, each tool gated by an approval policy. fare-rules owns book_flight.
cancel_booking is the one tool built per request: it refunds from the
conversation's booking in graph state, so the cancel node holds it under
cancellation-policy with skillTools.
The reconnect test: the client records a watermark, then its socket dies while
the fare-rules skill frame and the booking approval stream. A new connection
with { conversationId, watermark } gets welcome.pending re-announcing the
same approval id, and a replay that starts at watermark + 1, has no gaps, and
is exactly the frames the dead socket missed, skill frame included. The
approval is answered from the new socket. The next model round is still
offered book_flight (the unlocked set is derived from the journal), and the
booking runs once.
The skill-held cancellation runs before a second pause ("rebook?") in the
same node, so the resume that answers it replays the node, agent loop and all.
The model is not asked again, and the journal keeps cancel_booking at one
call. Two tabs confirming at once get one run and one refusal (busy), and
asking to cancel again finds the cancelled booking in graph state.
Support desk
A router sends each turn to a node with its own tools, and the probe asserts
what each node bound. The skills are tag-scoped per route:
refund-policy (billing) owns issue_credit, and incident-runbook (tech)
owns escalate_to_specialist. The probe asserts tag scoping both ways:
- each node's prompt lists only its own skill;
- loading the other route's skill is an unknown-skill observation;
- the only
skillframe on each turn is the node's own.
The other two agents in this scenario are mekik apps too:
- the knowledge base is served with
MekikMcpServerand consumed by the tech node throughwithMcpTools, askb__search, which surfaces as an ordinarytool_calltrace (MCP); - the network specialist is served with
MekikA2aServer. The desk hands the case over withmessage/sendinsidemekik.tool, which journals it (A2A).
When the specialist's task comes back input-required, the desk parks on its
own interrupt carrying the specialist's question and chips. It then forwards
the human's answer as a data part on the same task. The probe checks one
hand-off message and one answer, the same task and context ids, a single
reboot on the specialist's side, and a resume that replays only the hand-off
node, so the knowledge base is not called again.
Writing your own
The scenarios share
ts/examples/lib/probe-kit.ts:
ScriptedModelholds per-node queues of scripted decisions (say,call) and records what each node's model was shown: its toolbox, its system prompt, every observation, and the tool list offered in each round.asChatModel(node)wraps the same script as a chat model forrunAgent.runToolsis a hand-wired model↔tool loop, with decisions journaled per node. Itsskillsoption appliesrunAgent's skill gating by hand: the skills' own tools are held untilwithSkills'onLoadedreports a successful load;skillToolsadds pre-wrapped extras.Collector,checkanddescribecover frame capture and the output style.
To write a new scenario, copy a directory, swap the domain, and assert the frames your users depend on.