Accountability
Every action, decision, and approval has a record with a time and, where a person was involved, a name.
Everything an Agent does is verifiable. Every granular action or decision is visible and retractable for a human to audit.
Before-and-after screenshots with the target marked, the coordinates, the text typed, the values extracted, the timestamp. Every decision records its inputs, the rule that fired, and the branch taken. Every review records who decided, what changed, and how long it waited. Any run exports as one tamper-evident package, stored on your own machine.
The problem
An agent that posts invoices, edits vendor records, or files a pay application does work that used to carry a person's initials. The ERP still logs the change: a field on a voucher, a new value, a user ID, 2:14 a.m.
That log cannot say what the agent was looking at, which rule sent the invoice down one branch instead of another, or whether a person checked it first. Those are the questions an auditor, a controller, or an investigator asks.
Every action, decision, and approval has a record with a time and, where a person was involved, a name.
When a number is wrong, the record shows which link failed: the document, the reading, the rule, or the entry.
An auditor, lender, or surety can test the agent's work from the record without taking anyone's word for it.
What is kept
Nothing is sampled, and nothing depends on a busy afternoon. Five layers of evidence, from a single click up to a full recording of the run.
| Level | What is kept |
|---|---|
| Every action | Screenshots before and after, the target marked in red, the click coordinates, the text typed, the values extracted, the time |
| Every decision | The inputs, the rule that fired or a summary of the model's reasoning, the branch taken |
| Every review | Who looked, what they saw, what they changed and decided, when, and how long it waited |
| Every run | The full path, every exception, every retry, and the outcome |
| Optional, per workflow | Full video of the run |
A plain screenshot shows what was on the screen. A marked one shows where the agent decided to act, which proves it pressed the Save on the invoice form and not the one on the vendor panel beside it.
A log of what a model said is not evidence of what happened on screen. This record is the screen itself, before and after.
The screenshots cost the agent no extra step. The Vision Layer captures the screen before it acts and again to verify the expected change; the audit trail keeps what verification already produced.
The evidence schema
Each record lists the fields it holds and the question each field settles. Pick a record type.
| Field | What it holds | Question it settles |
|---|---|---|
| Before screenshot | The screen as the agent saw it, immediately before acting | Was the right record open? |
| Target mark | A red box on the element acted on | Which of two look-alike controls was pressed? |
| Coordinates | Where the click or keystroke landed | Can the input be reproduced exactly? |
| Typed text | The characters entered; credential vault values appear masked | What went into the field? |
| Extracted values | Each value read, tied to the screen region it came from | Does the reading match the source? |
| After screenshot | The screen once the expected change appeared, or failed to | Did the action take effect? |
| Time | When the action happened | Where does it sit against the ERP's own log? |
| Field | What it holds | Question it settles |
|---|---|---|
| Inputs | The values the decision used, each linked to where it was read | What did the decision know? |
| Path | Deterministic rule or model judgment | Was this arithmetic or interpretation? |
| Basis | The rule that fired, or the model's typed answer with a short summary of its reasoning | Why this branch? |
| Routing | Whether a model answer fell below the step's confidence threshold and went to a person | Did a doubtful reading reach a human? |
| Branch | Which way the workflow went | What happened next? |
Rules are versioned with who changed them and when, so the rule's own history shows which wording was in force at the time of any run. Exact-rule decisions make no model call; the rule is the whole reason.
| Field | What it holds | Question it settles |
|---|---|---|
| Reviewer | The person who acted | Who is accountable for the call? |
| What they saw | The agent's record and the source document, as displayed | Was the reviewer shown the evidence? |
| What they changed | Each corrected value, old and new | Did a person override the agent? |
| Decision | Approve, reject with a reason, correct, or supply a missing fact | What did the control decide? |
| When | The time of the decision | Did approval come before the posting? |
| Wait | How long the item sat in the queue | Is the review step a bottleneck? |
| Field | What it holds | Question it settles |
|---|---|---|
| Workflow version | The approved version that ran | Which logic produced this work? |
| Path | Every step taken, in order | What did the agent actually do? |
| Exceptions | Each flagged item, with its step and screenshot | What did it refuse to finish alone? |
| Retries | Each repeated attempt | Was the clean result clean the first time? |
| Outcome | Finished clean, finished with exceptions, or stopped, with what committed | What state were the systems left in? |
| Video (optional) | A full recording, where the workflow's setting turns it on | What did the whole session look like? |
Three things never enter any record: credential values, one-time codes, and the model's chain of thought.
Validation evidence
An agent's evidence does not begin on its first unsupervised run. Every step was performed and approved in the Studio, and each approval saved a screenshot and a clip. That validation evidence is the agent's first audit trail. By the time it runs alone, the record already shows each step working, and who approved it.
The history continues through every change. Every revision of the workflow is kept, and an approved version is read-only. The evidence behind version 7 does not vanish when version 8 ships, and nobody can quietly edit the version an auditor is testing.
Answers, not guesses
Each answer comes with a screenshot, a value, or a name attached. No one has to take the agent's word for anything.
Run console
The trail is built to be read by a controller, not parsed by an engineer. A typical inspection goes like this.
In the run console's history, by workflow and date.
The path the agent took is laid out step by step, in the same boxes as the workflow's flowchart.
Before and after screenshots side by side, with the red mark on the target, the text typed, and the values read.
The inputs, the rule or the model's summary, and the branch appear together.
The reviewer's name, the values they changed, their decision, and the wait time.
Follow its lineage back to the source.
The run, or a range of runs, as one package.
Need help? The Support panel lets you reference the exact run and step in a thread, so a LaunchAI support engineer knows exactly which one you mean.
Value lineage
Value lineage is the record of where one number came from and everything that happened to it before it landed in a system of record. Up to six stops.
A supplier invoice PDF shows freight of $412.50 beside a merchandise subtotal of $6,120.00. The figures are an illustration. Follow the freight charge forward.
Page 1 of the PDF, the totals block, with the region around "Freight 412.50" outlined.
Extracted as currency, 412.50, held as an exact decimal. No rounding happens on the way in.
The controller's exact rule: freight above 5% of the merchandise subtotal goes to review. 5% of $6,120.00 is $306.00; the charge is 6.74%. The decision record shows both inputs, the rule, and the review branch.
The AP lead compares the charge with the carrier quote on the PO and approves. The record logs the name, the time, and an 11-minute wait.
The agent types 412.50 into the freight or miscellaneous charge line of the ERP's invoice entry screen. The after screenshot shows the value in place.
After the save, the ERP displays its voucher number. The agent reads it into the run record: the join key to the ERP's own change log.
That is how a question usually arrives. An auditor points at a voucher in the ledger. Find the run by workflow and date, open the step that saved that voucher, and follow the freight value back through the entry, the review, the rule, and the reading to the outlined region of the PDF. Every stop carries a screenshot or a name.
Diagnosis
A vendor invoice posts with a date of March 4. The vendor meant April 3.
Without lineage, all three look identical: a wrong date in the ledger and an afternoon of guessing. A fourth failure hides between source and rule: the reading. The pixels say 8; the extracted value says 3. The lineage shows it the same way, side by side.
| Failure | What the lineage shows | The fix | Who makes it |
|---|---|---|---|
| Source | The document itself carries the wrong value | Ask the vendor for a corrected document | AP clerk |
| Reading | Pixels and extracted value disagree | Route that field to review, or tighten the step with support's help | Process owner |
| Rule | Correct reading, wrong transformation or branch | Edit one rule or one table row; versioned and reversible | Rule owner |
| Entry | Correct value, wrong field | Re-show one step in the Studio | Process owner |
Then fix the ledger and prove the fix. The posted invoice is corrected in the ERP by your normal correction procedure. The next run's decision record shows the new table row firing, which is the evidence that the fix took.
PCAOB · AICPA
Auditors of public companies work under PCAOB standards. Most mid-market companies are private and are audited under the AICPA's clarified standards (the AU-C sections), which carry parallel requirements for evidence and sampling. Either way, the requests arrive on a provided-by-client list, and they fall into a handful of kinds.
| Auditor request | What they are testing | What you hand over |
|---|---|---|
| Walk one transaction through | How transactions are initiated, authorized, processed, and recorded | One run export, with the lineage of the chosen transaction from source document to ledger |
| The population for the period | That every item had a chance to be selected | A range export covering every run in the period, clean, with exceptions, or stopped |
| Support for selected items | Each sampled item's accuracy and authorization | The action, decision, and review records for each selected item |
| Evidence a control operated | That the review happened before posting, by an authorized person | Review records with names, times, decisions, and wait times |
| Changes to automated logic | That changes were authorized and tested | Rule and workflow version histories: who, when, what, and the approver |
| Who can change or approve | Segregation of duties | The permission assignments for build, approve, edit rules and tables, review, and run |
| Reports the agent produced | Accuracy and completeness of company-produced information | The run record and the lineage of each figure in the report |
AS 2201 · walkthroughsA walkthrough follows a transaction “from origination through the company's processes… until it is reflected in the company's financial records.”
AS 2315 · AU-C 530 · sampling“All items in the population should have an opportunity to be selected,” and the sample must be one the auditor can expect to be representative.
AS 1105 · company-produced informationThe auditor tests its accuracy and completeness, or tests the controls over it.
Worked example
The controller of a 120-person mechanical contractor prepares for the first audit in which AP ran AI agents for vendor invoice coding to job cost. All volumes and thresholds are illustrations.
The request list names six items. Here is how each is answered.
The auditor picks one August invoice. The controller exports that run and opens the lineage for the total: the PDF region, the reading, the job and cost code lookup, the review (it was over threshold), the entry, and the ERP's voucher number. The auditor follows it in the ERP with the clerk's own screens.
A range export from July 1 to September 30, 2026. The population list comes from the runs' extracted invoice numbers, and it includes the 19 exceptions.
The auditor selects 25 invoices, including 3 that went to review and 1 exception. For each: invoice read, coding decision, review if any, entry, voucher number.
Invoices over $10,000 go to the project accountant. The 3 reviewed items show the reviewer's name, the approval time, and that approval preceded entry. On one, the reviewer changed the cost code; old and new are recorded.
On July 22 the review threshold dropped from $15,000 to $10,000, approved by the controller. On August 14 the coding table gained rows for a new project. Both sit in the version histories; the auditor tests items on either side of July 22.
The controller lists who holds build, approve, edit rules and tables, and review permissions. The clerk who built the workflow cannot approve it for production.
Each names its step and carries the screenshot the agent saw; the ERP's own log shows the clerk's manual posting. The two records join on the invoice number.
Export package
Any run, or any range of runs, exports as one package: screenshots, timestamps, extracted values, decisions, and approvals.
The records are append-only and hash-chained per run. Each record carries a cryptographic fingerprint of the one before it. Change or delete any record and every fingerprint after it stops matching. Tampering is not prevented by a promise; it is exposed by arithmetic.
Edit record 4 and its fingerprint changes; record 5 no longer points at it. The break is visible to anyone checking the chain.
| Contents | Detail |
|---|---|
| Run records | Path, exceptions, retries, outcome, and the workflow version |
| Action records | Before and after screenshots with the red target mark, coordinates, typed text (masked where it came from the vault), extracted values, times |
| Decision records | Inputs, the rule that fired or the model's answer and reasoning summary, the branch |
| Review records | Reviewer, what they saw and changed, decision, time, wait |
| Lineage | The stops for each value, from source region to destination field |
| Video | Included where the workflow records it |
Screenshots are digitized evidence. The PCAOB's audit evidence standard, AS 1105, ranks original documents above converted ones, whose reliability depends “on the controls over the conversion and maintenance of those documents.”
The controls here are specific: capture at the moment of action, a chain that breaks when edited, and storage on hardware you control.
Agree three things with the auditor before the first export
Bring a sample export to that conversation during your pilot; the format and verification steps are worth confirming with us then.
Retention design
The evidence is stored on the machine that ran the agent, inside your network. You set the level of detail and the retention period per workflow. For example:
Match the settings to your own record-keeping obligations. The product does not decide them for you.
A retention plan takes five decisions per workflow.
A posted invoice, a payroll register, a vendor master change, a status report.
Payroll is a common driver: the U.S. Department of Labor requires payroll records for at least three years, and the time cards and wage rate tables behind them for two.
Every screenshot, run summaries, or full video.
And write down who approved it.
Under FRCP 37(e), a court can impose measures when ESI that should have been preserved for anticipated litigation is lost. When a dispute touches a workflow, extend retention or export the runs first.
| Workflow | Record it supports | Detail | Retention (illustrative) |
|---|---|---|---|
| Vendor bank-detail changes | Vendor master changes, fraud controls | Every screenshot | 1 year, then longer if your policy requires |
| Payroll data entry | Payroll register | Every screenshot plus full video | At least 3 years, per the payroll record rule |
| Vendor invoice posting | AP vouchers | Every screenshot | Your financial records schedule |
| Daily status report | Internal report | Run summaries | 30 days |
Because the store is the machine's own disk, the machine becomes part of your records infrastructure. Put agent machines under the backup policy that covers other financial records. Before a machine is retired or wiped, export or retain the evidence it holds.
Access
Screenshots show whatever the agent saw: bank details, employee records, tenant ledgers. Access is controlled the same way as the rest of the product.
Separate from building, approving, reviewing, and running.
Every permission, every time. A hidden button is not a control.
Evidence belongs to the tenant (a company, division, or client) and stays inside it.
Screens that showed a vault-filled field carry a mask where the value was.
Sensitive actions can require it at that moment.
Keeps sensitive screens from outliving their purpose.
Governance
The trail matters most as proof that designed controls operated. Five LaunchAI controls leave their own evidence.
| Control | Evidence it leaves |
|---|---|
| Human review step before posting | Review records with names, times, and decisions |
| Exact rules for money decisions | Decision records naming the rule; the rule's version history |
| Approval before production | Workflow versions with the approver; only approved versions run |
| Workflow fence to declared applications | The declared scope on each approved version; the Runner refuses any other application |
| Enrolled machines only | The enrolled-device list in the Portal; a revoked machine's agents stop |
Metrics
The run console shows success rate, exceptions, and pending reviews per workflow. The rest come from the records themselves.
| Measure | How it is computed | What a bad trend means |
|---|---|---|
| Exceptions per 100 items, by step | Exceptions divided by items, grouped by the step that flagged them | One step, one screen, or one vendor needs attention |
| Review wait time | The median wait on review records | Reviewers are overloaded or assignment is wrong |
| Reviewer correction rate | Reviews with a changed value divided by all reviews | A reading or rule needs work |
| Retries per run | Retry count from run records | A screen or session is unstable |
| Root-cause mix | Share of fixes by source, reading, rule, and entry | Where maintenance time actually goes |
| Time to answer an audit request | Hours from request to package | Whether the evidence is findable |
Edge cases
Retries, stops, doubtful readings, masked values. The record is designed for the awkward runs, not just the clean ones.
Both attempts are in the run record. The failed attempt's half-produced output was cleared first, so the evidence shows one clean result and one failure, never a blend.
The record shows which steps committed and which never started. Stopping is not rolling back: a committed posting stays posted, and the record says so.
The decision record shows the answer fell below the step's threshold and went to review; the review record shows what the person decided.
A vault credential appears masked in the typed text and in every screenshot. The record shows that something was typed, never what.
Screenshots still carry the full action-by-action evidence; video adds the moments between actions.
Extend the workflow's retention or export the runs before the date passes.
Honesty
The platform
FAQ
What auditors, controllers, and IT ask most about the record.
For each model-path decision, it keeps the inputs, a summary of the model's reasoning, and the branch taken. It does not store the model's full internal chain of thought. Exact-rule decisions need no summary: the rule itself is the reason.
That is the auditor's call. What they receive is a screen-level record of every action, the rule or review behind every branch, and a package that shows any alteration. Ask them to review a sample export during your pilot, before the first period that relies on it.
No. The Studio's documentation describes the workflow as designed and validated: the flowchart, one box per step, the validation screenshot and clip. The audit trail records what each run actually did. An auditor reads the first to understand the control and the second to test it.
No one edits a record in place. Records are append-only and hash-chained per run, so a changed or removed record breaks the chain after it. The retention setting on each workflow is the sanctioned way evidence ages out, and it is set by people with the right permission.
The agent's record covers the agent's actions and shows the run's full path. Keystrokes a person makes by hand are logged by the application they used, such as the ERP's change log under that person's user ID. Join the two records by time and document number.
Yes. Detail and retention are set per workflow, so a workflow that only reads and reports can keep less, for less time, than one that moves money. The agent still verifies every action as it works; the setting decides how much of that evidence is kept afterward.
Trust is not expected. It's recorded.