The short read
A long report is only partly captured, and the part travels on as if it were the whole.
LLMs respond differently each time for the same prompt; they are not predictable and repeatable. This is one reason agents are not doing mission critical work without human supervision. We solved this problem.
An atomic workflow is an AI agent built as a chain of small (indivisible) steps, each of which finishes validly, or stops and says so. Every LaunchAI agent works this way. A step is an action, a decision, or a human review, with named inputs, an expected result, and its own record.
Origins
Atomos is Greek for uncuttable. In 1808 John Dalton published the first part of A New System of Chemical Philosophy, which gave each element its own atom with its own weight. Chemistry got a unit it could count.
Computing borrowed the word for the same reason. In 1981 Jim Gray described the transaction: a unit of work that commits entirely or not at all, and stays committed. In 1983 Theo Härder and Andreas Reuter named the full set of guarantees ACID. The A is atomicity. It is the reason a bank transfer cannot debit one account and forget to credit the other.
LaunchAI applies the same discipline to agents. The unit is the step. A step does one thing, checks that it happened, and writes its own record. A half-finished step never hands its output to the next one.
The problem
Hand a model a large job, such as "reorder whatever is low," and it starts well. Then it meets what nobody mentioned: a second Save button, a quantity in the wrong unit, a warning dialog. With no instruction left, it improvises. Sometimes harmlessly. Sometimes by converting the wrong recommendation into a purchase order, with full confidence.
The public benchmark agrees. On OSWorld, which scores agents on open-ended tasks on real computers, the best agents in September 2026 still missed roughly one task in seven, by their own self-reported numbers. Remarkable progress since 2024, and still a miss rate no controller would accept on a posting.
Back-office work multiplies the misses. Chain 30 steps that each succeed 99 percent of the time and the run finishes clean 74 percent of the time. The misses that cost money are the silent ones: a run that reports success after posting half a batch.
A long report is only partly captured, and the part travels on as if it were the whole.
Correct values, Save clicked, and a validation message opens behind another window. Nothing was saved. The log says done.
Twenty of forty receipts post, the session drops, and a rerun from the top posts the first twenty again.
A decision made on an incomplete list: nothing looks low because half the items were never read.
The answer is structural. Make every step small enough to check, then check it.
Building blocks
Every workflow is a chain of three step types. Each has a narrow purpose, named inputs, an expected result, and a record of its own.
| Type | What it does | Done when | Example |
|---|---|---|---|
| Action | Operates software: click, type, read, open, save, send | The expected screen or value appears | Open the ERP and wait for the home page |
| Decision | Examines the current state and picks a branch, by exact rule or model judgment | A branch is chosen and logged with its reason | Is available quantity below safety stock? |
| Human review | Stops the run and waits for a person | A person approves, rejects, or corrects | Approve this week's replenishment recommendations |
Three sorting questions place any piece of work in the right type
Which decisions get an exact rule and which get judgment is the job of the Dual-Path Rule Engine. How the review step pauses, routes, and resumes is on Human review.
Anatomy of a step
A step is a contract. It says what it needs, what it will do, and what the screen must show when it is done. The runtime holds it to that contract on every run.
The expected result is mandatory. A step with no proof of done is not a step. If the inquiry returns no rows, step 3000 fails. It does not report "zero items below safety stock" and let the run conclude that nothing needs ordering.
The outputs are held back. A value exists for later steps only after the step that produced it has verified. That single rule is what keeps a short read from turning into a confident branch.
| Part | What it declares | Example: step 3000, Monitor inventory levels |
|---|---|---|
| Name and purpose | One sentence a manager can read | Read available quantity and safety stock for every stocked item |
| Inputs | Named, typed values the step needs before it starts | The warehouse to read, set when the run starts |
| Actions | The ordered moves on screen, each checked before the next | Navigate to the balance inquiry, extract the grid, open the report, save a snapshot, send a notice |
| Expected result | What must be true when the step is done | A results grid with at least one row, and every row carrying all five fields |
| Outputs | Named values later steps may use, held back until the step verifies | {{inventory_balances}} |
| Boundary | What the step may touch | Only the applications and folders the workflow declares |
| Record | What it leaves behind | Before and after screenshots for each action, the extracted values, timing, outcome |
Execution
The same sequence runs when the Studio validates a step during building and when the agent runs alone at night.
A step approved in the Studio has already passed through the engine that will run it in production.
Every named input exists and has the right type. A missing warehouse code fails the step before the agent touches a screen.
One action at a time through the five moves of the Vision Layer, ending with a fresh capture. No target, or two candidates? It does not guess.
Extracted values are staged, not handed forward.
The expected result is checked against the screen, the file, or the value.
Staged outputs become available, the record closes, and the connector's condition picks the next step.
Staged outputs are discarded, the failure is recorded with what the agent saw, and policy decides: retry, send to a person, or stop.
Guarantee
One boundary, stated plainly: LaunchAI does not roll back your ERP. Atomicity governs the agent's own work, meaning its outputs, its record, and its place in the chain. If one step posted a receipt and the next step failed, the receipt stays posted and the record says so.
A step completes and verifies, or fails and flags. There is no third outcome called "mostly worked."
If a read captured 212 of 300 rows before a dialog stole focus, the 212 are discarded. The rule holds in the Studio and at night.
Fix the cause and it resumes at the failed step. Steps that already committed do not run again.
The record shows which steps committed and which never started.
Effect states
Reading is safe to repeat. Writing is not. Click Post twice and an ERP may post twice. So every step that changes something in another system ends in one of three plain states.
The agent saw the proof: a new receipt number, a changed status, the posted line in the grid. The step commits.
commitClear evidence of no change: a validation message, a record-locked warning, the form still open with no document number. The step may try again if policy allows.
retry allowedThe screen froze after Post. The session dropped. A confirmation flashed and vanished. The agent does not click again. It stops, holds its place, and asks.
hold & askThe third state is where duplicate postings are born. A retry that assumes the first attempt failed turns one receipt into two. LaunchAI treats an unknown result as unknown until something settles it: a person looking at the before and after screenshots and the ERP itself, or a business key the target application checks on its own.
An operation is idempotent when doing it twice has the same effect as doing it once. Stripe's API accepts an idempotency key and, when the same key arrives again, returns the saved result instead of charging twice. A screen has no such field, so the key lives in the business data: a vendor invoice number the ERP rejects as a duplicate, a packing slip number stamped on a receipt. With such a key, "did it land?" becomes a lookup. Without one, a person answers it, because a guess is how the same invoice gets paid twice.
The same thinking guards the front door: a run request carries its own key, so a double-clicked Run button cannot start two runs.
Resume semantics
A resumed run is the same run. Its identity, its committed steps, and the values those steps produced carry forward.
Values extracted before the failure are not read again. If step 3000 captured the balances at 6:02 and step 12000 failed at 6:40, the resumed run still works from the 6:02 snapshot, the same numbers the approved recommendations were based on. The record says which snapshot fed which decision.
Resume is not a new run with a shortcut. A new run starts at step 1000 and reads everything fresh. A resumed run continues the story the record has already told.
| Situation | What runs again | What never runs again |
|---|---|---|
| A read step failed partway through a report | The whole read step, from its first row | Every step before it |
| A write step did not land (validation error, locked record) | That step, once the cause is fixed | The committed writes before it |
| A write step's result is unknown | Nothing, until a person or a business key settles it | The write itself, once confirmed as landed |
| A reviewer rejected an item | Nothing; a rejection is a branch, not a failure | Not applicable; the run follows the rejection path |
| The machine lost power mid-step | The interrupted step, its last action treated as unknown | Anything already committed |
Loops
Back-office work arrives as lists: 140 count lines, 60 receipts, 300 recommendations. A loop in LaunchAI names the collection it walks and a maximum number of passes. There is no unbounded loop anywhere in the runtime.
Set the maximum from real volumes, with room for a heavy month.
Worked example
An inventory workflow as drafted in the builder for an Infor ERP. Eighteen nodes read left to right on one screen.
Start
Open the ERP
Monitor inventory levels
Configure replenishment and run MRP
Review replenishment recommendations
Record the rejection
End: recommendations rejected
Convert approved recommendations
Receive and post materials
Manage work in progress
Control finished goods
Conduct cycle counts
Approve inventory adjustments
Escalate the rejected adjustment
End: adjustment escalated
Post approved adjustments
Publish the summary
End
The count-and-adjust portion is inventory adjustment work on Infor CloudSuite Industrial (SyteLine), which can run to hours a day for busy plants. Screen names vary by version.
Step 2000 holds three actions: launch, wait for the home page within a timeout, set the process name. Step 3000 holds six, and its extraction names five fields (item number, warehouse, location, available quantity, safety stock) written to {{inventory_balances}}.
Nothing converts to a purchase order until a person approves at 5000, and no count adjustment posts until a person approves at 13000. Each rejection ends on its own recorded path.
If step 12000 fails on a mislabeled bin, the fix is step 12000. Materials posted at 9000 stay posted and are not posted twice.
Recovery
Each of these stops one box and leaves the rest of the chain alone.
In each case the record answers the auditor's question before it is asked: which step, what the agent saw, who intervened, what they found, and what ran afterward.
A warehouse filter was left on the balance inquiry the evening before. The grid comes back empty, the step fails on its expected result, and MRP at 4000 never runs on an empty picture. A person clears the filter; the run retries from 3000.
A buyer has the PO open in another session and the ERP refuses the receipt: clear evidence it did not land. Two policy retries; the lock persists; the step flags. The buyer closes the order, the run retries 9000, and the materials post once.
No receipt number, no error. The result is unknown, so the agent does not click Post again. A person gets both screenshots, searches by packing slip, finds the receipt posted at 6:31, and settles it as landed. One receipt exists.
The workflow builder
The builder is where a process becomes a map the agent follows and a person can read. It is also where an agent spends most of its working life: one thing changed, the rest untouched.
Steps run left to right, numbered 1000, 2000, 3000. Boxes mark actions, decisions, human reviews, and endpoints. Connectors carry their conditions in plain text, such as human_review_status = Approved. A published change governs the next run, so the map is never a picture of what the agent used to do.
{{name}}.Versions
Two people editing at once are detected, not silently overwritten. Any earlier version can be restored. Each run records the version it executed, and a run cannot execute against workflow material that changed after it was queued. Building and approving for production are separate permissions, and publishing can demand step-up authentication.
When a single step changes, the process owner re-shows it in the Studio and it is validated again on the live application before the new version goes up for approval. The change types are on the Run console page.
| State | What it means |
|---|---|
| Draft | Being built or changed. Copilot edits land here. |
| Pending approval | Submitted and waiting for a person with the right to approve for production. |
| Approved | Signed off and read-only. Changing it starts a new draft. |
| Published | The version that scheduled and on-demand runs execute. |
| Archived | Retired from use, kept, and restorable. |
Why the ceremony matters
It is written into a federal enforcement order. On August 1, 2012, Knight Capital's order router sent more than 4 million orders into the market in 45 minutes while trying to fill 212 customer orders. The SEC traced the damage to a deployment: a technician did not copy new code to one of eight servers. No second technician reviewed it. No written procedure required one. The old, dormant code on the eighth server woke up.
An agent that posts to your ERP deserves the discipline that router lacked: a second person before anything goes live, and no way for an approved version to drift.
Copilot
Type a change. The Copilot proposes the branch, shows the difference on the canvas, validates it against your permissions and the current version, and applies it only when you approve.
It also answers questions about the map: which steps write to the ERP, what step 4000 does, where the reviews sit.
If a colleague published a newer version while you were typing, the proposal is checked against theirs, not yours. And an approved proposal lands in a draft, which follows the same approval path as any other change. The Copilot shortens the typing, not the controls; it cannot edit an approved version in place.
| You type | What the Copilot proposes |
|---|---|
| "Add a check for items on quality hold before converting recommendations" | A decision step between 5000 and 8000, with a branch that skips held items and records them |
| "Send adjustments above the dollar limit in the approvals table to the controller" | A decision before 13000 that reads the limit from a table and routes by amount |
| "Save the count sheet as a PDF in the inventory folder" | A file action appended to step 12000, with the folder named |
| "Rename step 14000 to 'Send the rejected adjustment to the controller'" | A renamed node, with nothing else on the map changed |
Reuse
A validated step, such as signing in to a supplier portal, serves every agent that needs it. The portal's quirks are learned once.
Reuse runs deeper than steps. When LaunchAI trains on an application, the navigation map it builds is shared by every agent in the company. Rules and tables are shared too: the approvals table that routes inventory adjustments can route purchase requisitions. One edit, one owner, every agent that reads it.
Failure-mode catalog
None ends in silent success. The failures already walked through above are not repeated here.
| Failure | What the step does | What the owner sees | Recovery |
|---|---|---|---|
| Readiness check fails before the start | The run does not begin | Which check failed: version, machine, or inputs | Fix the gap; start the run |
| Target not found (a button moved) | Does not guess; fails | The screen with the missing target | Re-show the step in the Studio, or retrain the map |
| Many targets missing after a vendor release | Fails at the first affected step | A screen that no longer matches the map | Retrain the application map; workflows stay as they are |
| Two plausible targets (Approve and Approve All) | Refuses the ambiguity; fails | Both candidates marked | Tighten the action's wording; re-validate |
| Validation error on save | Records "did not land" with the message | The message text on screen | Correct the data; retry the step |
| Session expired | Fails at the login screen | The login page | Sign in again, or use the credential vault entry |
| One-time code requested | Waits for the named person | A request for the code | The person types it; the run continues |
| Extracted value fails its type | Rejects the value | The field and what was read | Goes to human review |
| Model read below its confidence threshold | Routes to a person | The source region and the reading | Reviewer confirms or corrects |
| A table has no row for this case | The decision cannot evaluate; fails | The lookup that found nothing | Add the row, or give the rule a no-row branch |
| Review waits too long | Escalates; the run holds its age | The waiting item and assignment | Next person on the path decides |
| Portal down for maintenance | Fails to reach its page | The maintenance notice | Pause for the window; resume after |
| Action aimed at an undeclared application | Refused before it reaches the screen | The refusal in the record | Declare the application in a new version, if it belongs |
The system
Each component does one job for the chain. Its own page explains how.
Governance and security
Atomic steps make controls enforceable because every control has a place to attach.
Drafts, approvals, and publication are recorded with who and when.
A workflow declares its applications, folders, and databases; the runner refuses an action aimed anywhere else, even if a document on screen asks for it.
Generated code is audited for forbidden imports and runs in a restricted interpreter. That is hardening, not a formal sandbox, and we say so.
Credentials, one-time codes, and model chain of thought are never logged.
An auditor reads the same numbered boxes the agent executes, not a description written after the fact.
What to measure
Atomicity produces clean numbers because every outcome is pinned to a step. These measures tell you whether a chain is healthy.
| Measure | How to compute it | What it tells you |
|---|---|---|
| Clean-finish rate | Runs finished clean ÷ runs started | Overall health; the run console shows it per workflow |
| Failures by step number | Count of flags per step, per month | Where the chain is weak; one step with most flags is one fix |
| Retries per step | Retry attempts ÷ step executions | Whether a timing or locking problem is building |
| Unknown results | Steps ending "unknown" per month | Whether writing steps need a business key |
| Time to resume | Resume time minus flag time | Whether flags reach someone who can act |
| Steps changed per version | Nodes edited ÷ versions published | Whether changes stay surgical |
| Model calls per run | Model-path steps executed per run | Cost drift; exact-rule steps make no model call |
Limits
Proof after every step costs something. Here is what, plainly.
Every action is checked before the next, and a vision-driven step is slower than an API call. Where a clean API exists, the agent can use it inside the same atomic contract. For a high-volume, API-only subset, an integration platform may run faster and cost less; LaunchAI also reaches the screens those platforms cannot.
The Studio carries that load by asking only questions that change what the agent does, but someone still has to say what "done" looks like.
A posting that committed stays committed. Reversing it is a business decision, and a reversal can be built as its own workflow with its own review step.
That minute is the price of never posting twice.
Per transaction, on licenses you already own. LaunchAI runs the same kind of steps with a proof after each one, and reaches the screens a recorded selector cannot.
FAQ
No. A step is the smallest unit you can verify, fix, and read on its own. Opening an ERP takes three actions: launch, wait, confirm. They live in one node because they succeed or fail together. Each action inside is still checked against the screen before the next one runs.
Small enough that you can say what "done" looks like on screen. If the expected result needs a paragraph, split the step. A good test: could a new hire check this step's outcome in five seconds from its screenshot?
A flag on one numbered box. Open it and you see the screen the agent saw, the value it expected, and what it found instead. Fix the cause, then retry from that step. Everything before it stays done.
No. A decision reads the current state and picks a branch. Writing is an action's job. Keeping them apart means the record always shows which step chose and which step changed something, which is the first thing an auditor asks.
Yes. Its inputs are the agent's record and the source document. Its expected result is a decision from an assigned person. Its record holds who decided, what they saw, and what they changed. It cannot half-complete: a decision arrives and the run takes that branch, or nothing moves.
The run stops at that step, or routes the item to a review if the workflow says so, and an alert goes out by the Portal, email, or chat. Nothing downstream runs on bad data, and nobody has to reconstruct the night from a log file.
No prompts. No babysitting. No dumb questions.