Skip to main content

Domain scenarios

The examples each demonstrate one part of mekik. The domain scenarios do the opposite: each one is a small, realistic desk that uses several features together, the way an application would. Each runs offline. A scripted model stands in for the LLM (the same seam as the --probe modes), the probe plays the client, and every step asserts the frame stream it got back. So each scenario is also an integration test, and CI runs all five.

They live under ts/examples/, one directory each, with a README:

cd ts
node examples/banking/banking.ts
node examples/insurance/insurance.ts
node examples/healthcare-triage/triage.ts
node examples/travel-booking/travel.ts
node examples/support-desk/support.ts

pnpm run examples:domain # all five; also part of `pnpm check`

The output follows the existing probes: the frames of each turn (→ a tool call, ← its result, ⏸ a pause, ▦ a component, ✦ a skill), a ✓ line per assertion, and a final ✅. A failed assertion exits 1 with the message.

At a glance​

ScenarioGraphWhat it pins on the wire
bankingrouter → accounts / history / payments (agent) → approve → execute money tools owned by their skills, two pauses in one node for a second approver, exactly-once execution, a genui-table history, a hidden tool that fails without a trace
insuranceintake (form pause) → assessa genui-form interrupt, decision tools owned by a skill and offered only after load_skill, a sign-off pause that keeps them, a typed rejection reason
healthcare-triageintake (agent) → escalate (chips) → book (client tool, agent)redacted identifiers on every frame, skill-held scoring and booking, an allowlisted client-declared skill, a chips-only pause, the page's calendar as a client tool
travel-bookingrouter → search → compare → book (agent); cancel (agent)a reconnect that replays exactly the missed frames with the unlocked tools intact, welcome.pending, a skill-held cancellation that runs once
support-deskrouter → billing / tech → handoff / chattag-scoped skills per route, per-node tool scoping, an MCP knowledge base, an A2A hand-off whose pause is relayed to the desk's human

Skills in every scenario​

Every scenario has its own §12 skill catalog (the skills app option, announced in the skills frame on connect). In each catalog a skill owns its tools: the entry lists them next to its instructions (SkillEntry.tools), so a skill reads as "instructions + tools", and the model cannot call a tool before it has read the rules that govern it. The tools never leave the server; every probe asserts the skills frame names none of them.

ScenarioSkill (tag)Tools
bankingwire-transfer-rulescheck_transfer_limit, transfer_funds
bankingdispute-handlingopen_dispute
insurancewater-damage-assessment (home)check_coverage, approve_claim
insurancerejection-letter (home)check_coverage, reject_claim
healthcare-triagetriage-protocol (triage)score_triage
healthcare-triageappointment-booking (booking)book_appointment
travel-bookingfare-rules (booking)book_flight
travel-bookingcancellation-policy (cancellation)cancel_booking (built per request, held with skillTools)
support-deskrefund-policy (billing)issue_credit
support-deskincident-runbook (tech)escalate_to_specialist

The probes record the tool list offered in every model round. Every scenario asserts that a held tool is absent before its skill loads and offered from the round after. Several also check that a premature call is refused with an observation and does not run. Beyond that, each scenario adds its own check:

  • tag scoping across routes (support desk);
  • the unlocked set surviving a pause or a reconnect (insurance, travel);
  • redaction holding on the skill frames (healthcare).

Every tool is built once, at module level. One that needs the run (to mount a card, or to key state by the conversation) reads the calling run's ctx with toolContext(config) from its LangChain config. Only one tool in all five scenarios is built per request: travel's cancel_booking refunds from the conversation's booking in graph state, so the cancel node holds it under cancellation-policy with runAgent({ skillTools }), and says why in a comment.

Most agents use runAgent({ skills }). The support desk's tech node shows the hand-wired form, because its MCP tools come pre-wrapped by withMcpTools: withSkills(ctx, filter, { onLoaded }) names each entry's tools in the load_skill observation and reports each successful load, and the loop offers that skill's tools (mekik.skillTools) from the next round.

Banking​

The payments node is a runAgent loop over a catalog whose skills own the money tools:

const SKILLS: SkillEntry<StructuredToolInterface>[] = [
{
name: "wire-transfer-rules",
description: "Rules for sending money: fraud screening, limits, second approvers. Load before any transfer.",
instructions: "1. Screen every transfer with fraud_screen. …",
tools: [checkTransferLimit, transferFunds],
},
{
name: "dispute-handling",
description: "How to open a card or account dispute for a transaction the customer does not recognise.",
instructions: "Confirm the date, description and amount from the history, then open_dispute. …",
tools: [openDispute],
},
];

The node passes only its always-on tools (lookup_payee, fraud_screen, flag_for_review) and skills: true. Balance and history tools stay always-on in their own nodes. The probe asserts that transfer_funds is not offered before its skill loads, and that a premature call is refused without staging anything. It also checks that the skill frame's seq comes before the transfer trace, and that loading dispute-handling unlocks only open_dispute.

transfer_funds only stages the transfer; the approvals come next. Over the $5,000 limit, the approval node calls mekik.approve twice with different keys: the customer first, then a second approver. On the resume that answers the second pause, the node re-runs from the top and the customer's answer comes back from the journal, so the customer is not asked again. The probe asserts that, plus one execute_transfer per transfer and a ledger debited once.

The history is a genui-table mounted from inside the tool. The fraud check runs with a show: false policy, so no frame for it reaches the wire. When the service is down the tool throws, and runAgent hands the model Error from fraud_screen: fraud screening service timed out as an observation. Following the skill, it calls flag_for_review instead of transfer_funds, and no money moves.

Insurance​

The claim form is the pause. mekik.approve with ui: mekik.genui.form.ref({ fields }) and no actions mounts the form, and the submitted values come back as the resume answer. The app's skills option holds three skills: the skills frame announces them on connect (no instructions), and the adjuster node offers only the home-tagged ones.

The adjuster is a runAgent loop. Only lookup_policy is always offered. water-damage-assessment owns check_coverage and approve_claim, and rejection-letter owns check_coverage and reject_claim. The probe asserts four things:

  • Before load_skill, only lookup_policy and load_skill are offered.
  • A premature check_coverage call is refused with an observation naming both skills that own it. The tool doesn't run and nothing is traced.
  • From the round after the skill frame, the skill's tools are offered, and the other skill's decision tool never is.
  • Across a pause. Each decision needs a senior adjuster's sign-off (an approve policy). On the resume, the earlier rounds replay from the journal, the next round is still offered the unlocked tools, and the decision runs exactly once.

A rejection carries a typed reason. reject_claim's zod schema pins code to EXCLUDED_PERIL | POLICY_LAPSED | BELOW_DEDUCTIBLE, and { code: "EXCLUDED_PERIL", clause: "4.2(b)" } lands on the typed claim-decision component.

Healthcare triage​

The medical record number arrives as a verified claim (StaticTokenAuthenticator, read with mekik.authClaims), never in the chat. The intake tools use the same redaction technique as sql-agent.ts --probe: a redact policy, here on runAgent. The model reads the real MRN, date of birth and name, while the traces show «redacted». The probe checks that no identifier appears on any frame: skill frames (which carry only id, name, status and source) and a second tab's full transcript replay included.

Both agents use skills, each scoped to its own tag. triage-protocol owns score_triage, the red-flag rules, and appointment-booking owns the server-side book_appointment. The calendar itself stays a client tool.

This is also the scenario with client-declared skills. The portal declares two in hello.skills, and the server opts in with a clientSkills allowlist that accepts plain-language and drops override-triage, a skill that would tell the model to ignore red flags. The probe asserts both outcomes:

  • plain-language is in the prompt and its skill frame says source: "client";
  • override-triage never reaches the model, and loading it is an unknown-skill observation with no frame.

Client skills are per connection: another patient's portal, which declared none, does not get plain-language.

An urgent score parks the run with mekik.choose (three chips, no form). The booking slot comes from the page's own calendar: mekik.callClientTool(ctx, "open_calendar", …) parks on an interrupt carrying data.tool, and the page answers { ok: true, result }. The booking agent then loads its skill, and book_appointment runs once against the real MRN. An emergency answer mounts a genui-alert and never opens the calendar.

Travel booking​

Search, compare and book are separate nodes. The comparison is a genui-table with a literal chunk id, so the replay after the pick re-renders the same element. The book and cancel nodes are runAgent loops, each tool gated by an approval policy. fare-rules owns book_flight. cancel_booking is the one tool built per request: it refunds from the conversation's booking in graph state, so the cancel node holds it under cancellation-policy with skillTools.

The reconnect test: the client records a watermark, then its socket dies while the fare-rules skill frame and the booking approval stream. A new connection with { conversationId, watermark } gets welcome.pending re-announcing the same approval id, and a replay that starts at watermark + 1, has no gaps, and is exactly the frames the dead socket missed, skill frame included. The approval is answered from the new socket. The next model round is still offered book_flight (the unlocked set is derived from the journal), and the booking runs once.

The skill-held cancellation runs before a second pause ("rebook?") in the same node, so the resume that answers it replays the node, agent loop and all. The model is not asked again, and the journal keeps cancel_booking at one call. Two tabs confirming at once get one run and one refusal (busy), and asking to cancel again finds the cancelled booking in graph state.

Support desk​

A router sends each turn to a node with its own tools, and the probe asserts what each node bound. The skills are tag-scoped per route: refund-policy (billing) owns issue_credit, and incident-runbook (tech) owns escalate_to_specialist. The probe asserts tag scoping both ways:

  • each node's prompt lists only its own skill;
  • loading the other route's skill is an unknown-skill observation;
  • the only skill frame on each turn is the node's own.

The other two agents in this scenario are mekik apps too:

  • the knowledge base is served with MekikMcpServer and consumed by the tech node through withMcpTools, as kb__search, which surfaces as an ordinary tool_call trace (MCP);
  • the network specialist is served with MekikA2aServer. The desk hands the case over with message/send inside mekik.tool, which journals it (A2A).

When the specialist's task comes back input-required, the desk parks on its own interrupt carrying the specialist's question and chips. It then forwards the human's answer as a data part on the same task. The probe checks one hand-off message and one answer, the same task and context ids, a single reboot on the specialist's side, and a resume that replays only the hand-off node, so the knowledge base is not called again.

Writing your own​

The scenarios share ts/examples/lib/probe-kit.ts:

  • ScriptedModel holds per-node queues of scripted decisions (say, call) and records what each node's model was shown: its toolbox, its system prompt, every observation, and the tool list offered in each round. asChatModel(node) wraps the same script as a chat model for runAgent.
  • runTools is a hand-wired model↔tool loop, with decisions journaled per node. Its skills option applies runAgent's skill gating by hand: the skills' own tools are held until withSkills' onLoaded reports a successful load; skillTools adds pre-wrapped extras.
  • Collector, check and describe cover frame capture and the output style.

To write a new scenario, copy a directory, swap the domain, and assert the frames your users depend on.