How the platform is put together, what each part is responsible for, and what it can prove about its own output.
Arrow keys or click to advance
The rest of this deck answers them in that order.
Documents arrive from wherever they already are. Everything past that line runs in one private account, and every enterprise customer has their own.
One service is the only writer to the record. Every other component asks it, so tenant scope is decided in one place, and that is the only place it has to be right.
The process graph, the leases and the retries are durable relational state, so a process that is halfway through has a position you can read and report on.
Work executes in sandboxes that see one document, hold no lasting credential, and are handed an identifier they must resolve through the control plane.
There is one route out to model providers, and two doors you use: the API and the data lake. Both are places you can put a control.
A document is one portable, open-format file. The original bytes, the values read out of them and the evidence for each of those values all sit inside it.
Retention, export, legal hold and deletion all act on one thing.
A check that passes while a reviewer is looking at it passes identically when the server reprocesses it, because both run the same compiled code, built once and shipped to both places.
A reviewer works against the real document in the tab, not a rendering of it, and a document of tens of megabytes opens and edits there without a round trip per keystroke.
Your taxonomy is the business meaning of the work, authored once. The prompt schema and the storage schema are both generated from it, so there is nothing to keep in sync by hand.
You author the taxonomy once, in the words the business already uses for these fields.
One generation pass produces all four at the same time, off the same tree, so they cannot disagree with each other. Changing what a field means is one edit to that tree, reviewed as a change to a governed resource, and the same pass rewrites all four.
Most of what makes a document readable sits outside it: what your team knows about the counterparty, about the layout, about the exception that comes up twice a year. In the platform that knowledge is a resource in its own right, versioned and owned by your people.
A knowledge item is one piece of your operation's expertise: a taxonomy entry, a prompt, a worked example, a lookup row, a rule. Items are grouped into versioned sets, and each set carries an expression over the features content picks up as it moves, so the right knowledge reaches the right content with nobody routing it by hand.
The matching knowledge is merged into the document as a new version before extraction runs. That makes a given extraction reproducible from one artefact, and it means a value can name the instruction that shaped it as precisely as it names the words it came from.
A correction made in the course of review becomes a candidate item. Automation can propose one; a person with the authority promotes it into the live set. Nothing automated writes to a set that is in force, which is what keeps the loop governed as it compounds.
The schema the model receives will not accept a bare value. Every field has to arrive with the lines it was read from, so an uncited answer is invalid before anything looks at it.
Deterministic code then looks for the claimed value at exactly those lines, converting types as it goes.
If it finds it, the value is tagged onto the exact words on the page and becomes typed data, and the tokens it came from stay attached to it.
If it cannot, the value is withheld; there is no confidence threshold that lets it through. It becomes a work item naming what the model claimed, which lines it cited, and what those lines actually said.
Where a business decides a field needs a looser rule, that is a setting on the field, applied by the person who owns the definition and recorded against the field.
A process is a graph of typed steps. Each one declares what has to be true before it can run, so the definition fixes the order and any worker that picks the step up honours it.
Human work is the same kind of node. It waits as a row, which costs nothing, survives every deployment, and stays queryable while it waits.
The reviewer's named action decides which edge is taken. Approve, rework and reject are separate paths through the same graph, so the graph itself shows why a document went the way it did.
These are six of the failures we design for, and what the platform does on each of them.
The work was never held inside the worker. It holds a lease on a row and extends it while it is alive. When the lease expires the work is reclaimed and retried, under the retry policy stamped on it when it was planned.
Queues carry notifications; the state is in the database. A background pass looks for work that stopped moving and replays the completion, but only once it can prove the underlying work actually settled.
Every dispatch slot goes to the least loaded tenant by weight, so a large batch cannot take the head of the line. Work that cannot get capacity goes back on the queue and waits, so a saturated system runs slower and keeps its error rate flat.
The step is a row with an owner and a status. It costs nothing to leave sitting, it comes through deployments intact, and you can report on it the whole time.
Retry, timeout and backoff are resolved when the work is planned and written onto it, so a configuration change cannot rewrite what is already in flight. A stage can be declared non-fatal, in which case exhausting its retries advances the process to the next step.
Calls retry with backoff behind the one gateway, and every attempt is priced and recorded against the document. There is no automatic failover between providers. If one goes down, someone repoints the gateway by hand.
Extraction logic, third-party modules and agents all execute in the compute plane, and that plane is treated as hostile.
Extraction and reasoning calls leave through one gateway, which holds the provider credentials, so cost, routing and rate limits are enforced in one place.
All of it is in the file, so it inherits the file's access control and the file's retention. Two years later the answer does not depend on a log that has since rolled or a service that has since been replaced.
Every enterprise customer gets a private tenant account in the region they choose. We operate it, and nothing in it is shared with any other customer.
Separate storage, separate database, separate compute, separate model route. There is no pooled tier underneath and no multi-tenant table your documents sit in next to somebody else's.
Every resource the platform holds is reachable through the API, under the same access control the screen uses. Extracted data lands continuously in the lake in open format, partitioned, carrying the definition it was read against.
Optional private DR: a second environment in another region, kept in sync, private to you on the same terms as the first.
Schema changes ship as versioned SQL applied under a lock at start-up. Nothing in the platform issues DDL of its own accord, which is what makes promotion between environments a reviewable event.
We operate the infrastructure. You own the knowledge and the record it produces, and you reach both through two documented surfaces. If your controls need to sit somewhere else, say so now.
Four answers to the exit question. The short version is that everything you would want on the way out is already in an open format, and most of it is already moving.
The documents are ordinary files.
Open format, one per document, carrying its own content, data, schema and history. Our command-line tool reads one and prints all of that with no platform running.
The configuration is text.
Definitions, processes and prompts round-trip to manifests in your own repository, and promote between environments as reviewed changes. Git is the versioning system, deliberately.
The data is already arriving.
Every version lands in the lake as an open, partitioned file carrying the definition it was read against, alongside the audit trail. A document from three years ago still explains itself.
There is one account to hand over.
Because your work lives in a single private account, a handover is a copy of things already in open formats, on infrastructure with one owner.
None of that is a migration path we would build for you on the way out. It is how the system stores things while you are using it.
Four properties decide whether a coding agent can be useful against a platform on its first day. They are ordinary properties of a well-built API, and we had them before the agents arrived.
A resource is referred to by a readable URI, never by an environment-specific id, so the same reference means the same thing in your sandbox and in production. Resolution happens server-side, under the same access control as every other call.
The API description is generated from the source. Every change regenerates it and fails the build if it differs from what is committed, and a route with no entry in it does not merge. An agent reading the spec is reading the running system.
One declaration per resource type produces the REST surface, both halves of access control, the audit record and the spec entry together. Seventy-seven resource types go through it, so the surface an agent learns once holds everywhere.
Definitions round-trip to YAML in your own repository, with environment-specific identifiers rewritten to slug references. A dry run puts a change through the server's real validators without writing it, so an agent can check its own work before it asks.
A spec says what a valid request looks like. It cannot say which of two valid-looking spellings is the live one, in what order things have to be applied, or how to write a definition a model can act on. So that ships as an installable pack: one skill per resource type, with the shape, worked examples and a mistakes table.
An agent here runs with no shell, no file write and no outbound network, and that is enforced at the tool boundary by a check that fails closed, not by a system prompt. Its capability is an allowlist per conversation, default deny. Its calls go through the same access control a person's do.
For configuration it drafts into your editor and a person presses Save. The agent-side save tools were removed so that stays structural. A delegated token whose authority can only ever be a subset of the invoking human's, and a staged change a human approves against a content digest, are built and in review.