Deterministic step
Computes.
Can produce numbers. Golden-tested: same input, same output.
A reference for the people evaluating the architecture: the catalogue you compose an agent from, the three node types you wire together, the type system that keeps a composition valid, and the engine, scheduler and compute layer that run it.
Computes.
Can produce numbers. Golden-tested: same input, same output.
Decides its own steps and uses tools.
Cannot produce numbers. Budgeted. Traced.
Runs a published, signed pipeline.
Runs deterministically regardless of what calls it, and carries its own guarantee.
Agent → workflow. A signed, tested workflow sits in the agent’s toolbox. The agent decides when to run it; the workflow still runs deterministically. The agent’s freedom is bounded by approved capabilities, not raw SQL — your security team approves the workflow list once, and agents compose freely after that.
Workflow → agent. A step in a deterministic pipeline can say “solve this with an agent”: the agent uses its tools, loops as needed, and returns a typed result.
You choose from this catalogue — you do not write code. Transform steps are catalogue items too, and every one of them is deterministic: the numbers in a report still come from code, never from the model.
Every step in the catalogue declares what it accepts and what it produces — a comparison step might take table → table, a summarize step table → text — and the composer refuses an incompatible connection outright. It is the one mechanism that lets you wire steps together freely without ever producing a broken agent.
A template ships from the vendor as a starting graph. The moment you edit it, the agent counts as forked, and it carries that mark everywhere it appears in the interface. There’s no silent drift — but drift is the default path once you start editing.
Edited — no longer follows the template.
It takes the run request, walks the template’s step graph, and writes every step’s input, output and status as it runs.
One run engine, one box. Concurrency is limited, not distributed across a cluster.
The engine doesn’t know what a step does — it resolves that from the Step Kit, the layer that owns every driver, transform and delivery channel.
A scheduled run and a manually triggered one write the exact same record. There is one way in, and one code path from there.
If the box was off when a run was due, each instance declares what happens next: skip it — the default — or run once when the box comes back.
If the previous run of an instance is still going, the new one closes immediately as skipped_overlap. It never starts a second run on top of the first.
The expected run calendar is actively watched. A report that should have been produced and wasn’t is itself an alarm event — silence is treated as a failure, not as nothing happening.
Production use surfaces five things the scheduler doesn’t yet do: dependencies between agents (run B only after A succeeds), backfilling a past period without mixing it into live reports, a maintenance window that defers a run rather than skipping it, an alarm that escalates through an ordered list of people when nobody acknowledges it, and a legal hold that exempts specific runs from retention deletion. None of these are built yet.
A site is rarely one box. Other Spark clusters, GPU servers, n machines serving m models behind vLLM, TGI, Ollama or NIM — Kogei can send work to any of them. Discovery is layered, most trustworthy first: adding a machine by address always works, and it is what most regulated customers choose; discovery through Kubernetes or Consul comes next; a subnet scan is last, and only runs against a range an admin supplies, with a narrow, rate-limited port list that is logged to the audit trail — an unannounced port scan on a bank’s network is exactly what its security team built its monitoring to catch. Add more than one machine and three things change: an agent step asks for a class of model, not a specific one; every machine carries a trust-zone tag an agent cannot cross; and the run record now states which machine did the work.
Most people in finance, operations or sales don’t draw flowcharts. The canvas still exists, but it isn’t the only way in: you can describe what you need in your own words, and a local model drafts the graph. The risk here is different — a wrong graph is caught by a person before it’s saved, not after it has shipped a number in a report, and the model can’t produce anything the canvas wouldn’t also accept: it emits typed nodes and edges from the same catalogue, so an invalid graph can’t be produced. It asks about the things that actually change the graph — which source, what threshold, who sees it, whether to add a research step — and assumes the rest, showing what it assumed rather than hiding it. It will not invent a source that isn’t connected; if you mention a system you haven’t connected, it says so and asks whether to add a placeholder or connect it first.