Make your engineering software agent-ready in eight weeks.

We wrap your existing scripting surface, manuals and support history into a tested agent layer — deployed inside your product, benchmarked on your own cases, metered so you can sell it, owned by you.

Book a pilot call Fixed scope, fixed price. On-prem or air-gapped.
run 0417 · spec-checkverdict.json

00:00 task: design misses spec limit on the primary check

00:02 → load_project iot_node.prj

00:03 → search_docs "response peak near upper bound"

00:04 → run_solver full sweep

00:41 worst margin: −12% vs limit

00:42 → propose_fix two component values

00:43 → run_solver full sweep

01:19 worst margin: +4% vs limit · verdict: pass

Three layers you keep after we leave

No new UI to sell, no lock-in to our models. The harness runs against whatever LLM you or your customers choose.

Tool adapters

Your TCL, Python, COM or batch interface becomes a set of typed, sandboxed tools an agent can call — exposed as an MCP server so any agent can drive your product.

Knowledge layer

Manuals, app notes, forum posts and ticket history become retrieval and reusable skills. Setup and convergence questions get answered before they reach support.

Verifier harness

A golden-case runner scored against your experts' judgments, with regression tracking. Every agent action ends in a verdict you can show a customer.

Ship it as a SKU, not a feature

Most vendors want to sell an AI tier and have no way to price or meter it. Every run the harness executes is metered, attributed and logged — so you can charge for it from the first pilot, before you know what a run is worth to your users.

  • Per-run metering — tokens, solver calls and wall time on every run
  • Cost attribution — by customer, seat and workflow, not one opaque monthly bill
  • A billable verdict log — an auditable record of what ran and what it returned
run 0417 · spec-checkmetered
model tokens18.4k in · 2.1k out
solver calls2
wall time79 s
run cost$0.42
billed toAI tier · seat 14/50

How an eight-week pilot runs

One engineer embedded with your team, one deliverable per fortnight, and a benchmark you run yourself at the end.

  1. Weeks 1–2

    Map the surface

    We inventory your scripting interface, file formats and docs, and pick the 20 workflows your users ask about most.

  2. Weeks 3–4

    Generate adapters

    Our agent writes and tests wrappers against a live install of your product. Your engineers review, not write.

  3. Weeks 5–6

    Build the golden set

    Support tickets become labelled cases. Retrieval and skills are tuned until recall clears your bar.

  4. Weeks 7–8

    Benchmark and hand over

    You run the harness on held-out cases. If the numbers hold, we convert to a licence; if not, you keep the code anyway.

Reference deployment

99%top-5 retrieval recall on the vendor's own support corpus
99.7%correct tool selection across the diagnostic workflow
28 toolswrapped from the vendor's existing scripting interface
8 wksfrom first call to a customer-run benchmark
Built for a simulation software vendor whose users needed targeted fixes, not another chat window. The agent diagnoses from their solver output and proposes changes their own tool then verifies. The physics changes from vendor to vendor; the three layers do not. Case study on request.

Who this is for

A good fit

  • Simulation, CAD, EDA or test software with a scripting interface and real users — structural, thermal, fluids, circuit, any physics
  • $5M–$500M revenue, no in-house AI team, and a renewal cycle where customers are asking about AI
  • A support queue full of setup, meshing and convergence questions
  • Experts whose checklists live in their heads

Not a fit

  • You want us to replace your engine — we make it easier to drive, we don’t rebuild it
  • You need a generic chatbot rather than a verified workflow

No scripting surface yet? We can add a minimal agent-facing one first — enough for the harness to drive your product, not a full public API. That is a shorter, separately scoped engagement that runs before the pilot, not a bolt-on to it.

Bring one workflow. We'll show you the harness running on it.

A 30-minute call: you describe your product's scripting surface and the question your users ask most, we scope the pilot.

Book a pilot call