Legacy desktop ERPs
Including thick clients delivered over remote desktop, where the local machine receives only pixels.
Software built for humans uses a screen. Our Agents do the same.
VLI is how your Agent can see and operate everything on your computer. It captures the screen, finds the target, acts with real mouse and keyboard events, and captures again to confirm the change.
Coverage
An API is an entrance the vendor controls. The vendor has to build it, maintain it, and agree to open it for you. Many never built one. Some sell it only above a certain license tier. Some expose half the screens and none of the reports.
The screen is the one interface every vendor ships, because people need it.
Including thick clients delivered over remote desktop, where the local machine receives only pixels.
The basic IBM 3270 display showed 24 rows of 80 characters, and emulators still support that layout.
Supplier, customer, government, bank.
Read on screen as an ordinary step, with no separate OCR pipeline to build.
Including the workbook with twenty years of macros in it.
"Is there an API?"
"Can a person do this with a screen, a mouse, and a keyboard?"
The test for whether an agent can do a job changes.
How it works
If the change did not happen, the step does not pass. A click that landed on nothing is caught at move five, the moment it happens.
Windows, tabs, forms, tables, dialogs.
Which pixels are a button, a disabled field, a table row, and exactly where each sits.
By exact rule when the step is exact, with a model when it needs judgment.
Through the operating system, with real mouse and keyboard events.
Capture again and confirm the expected change happened.
Every row lands in the run's record. What the record holds, and how to trace a value back through it, is covered on the Audit Trail page.
Grounding
Grounding turns “the Approve button” into a precise spot on the screen. The VLI works through four layers, cheapest and most exact first.
The model comes last. Most targets are settled by cheap, direct evidence before any model is asked anything.
The layers are ordered by cost and exposure. The first three run on the machine: nothing leaves it. The fourth sends the screen region to the model account your company configured, and it is the slowest to answer. Putting the model last keeps most steps fast, keeps most screens off the network, and makes the common case repeatable, because an element name read from an accessibility tree does not vary from run to run.
| Layer | What it reads | Where it earns its place |
|---|---|---|
| 1. Accessibility tree | The element structure an application publishes for screen readers (Microsoft UI Automation on Windows, the Accessibility API on macOS, AT-SPI on Linux) | Native applications and browsers that expose one |
| 2. OCR | Text read directly from the pixels | Labels, values, and anything printed on screen |
| 3. Local element detector | A small model on the machine that finds buttons, fields, and rows | Custom-drawn controls, terminal screens, remote desktop windows |
| 4. Vision model | The screenshot and the target description, sent to the configured model | Only when the first three cannot resolve the target |
Safety
Ambiguity is never settled by guessing.
If no candidate qualifies, or two do, the target is unresolved. The workflow then re-plans, routes the step to a person, or stops, as its policy says. The agent does not pick the first match and hope.
The cost of the alternative is concrete. “Delete” and “Delete filter” can sit on the same toolbar. “Post” and “Post and print” can share a dropdown. A tool that takes the first match will eventually take the wrong one, on the one record where it matters.
Coordinates
Screens lie about coordinates.
Windows defaults to 96 dots per inch, and at a 120 dpi setting everything grows by 25 percent. A button designed at (100, 48) then sits at physical pixels (125, 60), while its logical coordinates stay (100, 48). UI Automation reports physical coordinates, so a tool that mixes the two systems clicks beside the target. Multiple monitors with different scaling multiply the problem.
The VLI resolves each target on the screen it just captured and confirms the result on the screen it captures next. Whatever the scaling, a click that lands beside its target produces no expected change, and move five fails it on the spot.
Real workdays
Scrolling is the famous test case, and it has its own worked example below. The same discipline covers the rest of the mess an ordinary workday throws at a screen.
| Situation | What the VLI does | Where it is set |
|---|---|---|
| An expected popup, such as a confirmation dialog | A step handles it like any other screen | Taught in the Studio |
| A popup or session warning nobody expected | The capture sees it; verify fails because the expected change never happened; the step routes to re-plan, review, or stop instead of clicking through | The workflow's exception policy |
| A report or grid still loading | wait_for_element holds the step until the named element appears, with a timeout | The action's wait condition |
| A disabled button | Interpretation reads the disabled state; a press would change nothing, so the step fails rather than moving on | Verify, on every action |
| Two elements that match one description | Treated as unresolved, as described above | The workflow's exception policy |
| A crooked scan or a photo of paper | Read on screen as a model-path step with typed outputs | Dual-Path Rule Engine |
| A terminal screen | No element tree, so OCR and the local detector read it; navigation uses keys, because the screen was built for a keyboard | The actions taught in the Studio |
| A session that expired overnight | The sign-in screen appears instead, so verify fails; a workflow taught to sign in draws on a vault entry, and any one-time code is requested from a person | Security |
Application training
An application can be trained once, before any workflow depends on it. In an automated session, a small model explores the software: menus, dialogs, pages, and the moves between them. The result is a navigation tree, a map of every route the model found.
Training is separate from teaching. The map knows how to reach a screen. The workflow, taught in the Studio, knows what to do once there.
Cloud ERPs ship releases on the vendor's calendar, and portals change without notice. Small changes, such as a button that moved, are absorbed, because a target is found by what it is, not where it used to be. A large change fails at the step that meets it, with the screenshot of what the agent found, instead of wandering.
The fix is scoped: retrain the part of the map that changed, and have the process owner re-show any step whose screen moved.
Menus differ by role in most ERPs; a clerk's account and an administrator's account see different trees. Train under an account with the same rights the agents will use, so the map holds the routes the agent can actually take.
The action language
More than fifty actions cover pointer, keyboard, reading, scrolling, waiting, verifying, file dialogs, uploads, and annotation. The Studio writes them. Anyone can read them, which is the point: a controller reviewing a workflow should not need a developer to translate it.
Values in double braces are variables. A scrape writes what it read into a named variable, and any later step can use it. A demonstration value never hides inside an action; it is always a named variable, so the same action serves tomorrow's record.
Worked example
A buyer runs the open purchase order report every Monday. It returns 2,400 lines. About 30 fit on the screen at once.
Reading it is harder than it looks. Many modern grids render only the rows in view and swap them out as you scroll; AG Grid's documentation, for one, describes rendering just 10 extra rows past each edge by default. There is no whole table sitting in memory to read. The only way through is to scroll and read, roughly 80 screens of it.
A naive bot fails in three quiet ways:
The stakes are concrete: a skipped line is a late purchase order nobody chases, and a duplicated line inflates open commitments.
Open the report screen in the ERP and set the buyer and date filters from variables. ERP screen and menu names vary by version.
Run the report, then wait_for_element "Report ready," so the first capture is never a half-painted grid.
Read the report's printed line total into {{report_line_total}}.
Scroll and scrape the grid through the viewport ledger, writing PO number, line, release, supplier, due date, and open quantity for every row into {{open_po_lines}}.
verify_value: the ledger's row count must equal {{report_line_total}}. Equal passes; anything else fails the step.
Hand {{open_po_lines}} to the expediting steps that follow.
Viewport ledger
The VLI keeps a viewport ledger:
A row that appears at the bottom of one screen and the top of the next is recognized as one row. Each new screen is checked against what the ledger already holds before the agent moves again. When a scroll brings no new row into view, the grid has ended, and the count goes to step 5.
| Screen | Rows in view | Already in the ledger | New rows captured | Ledger total |
|---|---|---|---|---|
| Screen 1 | 1 to 30 | None | 30 | 30 |
| Screen 2 | 29 to 58 | 29 and 30, matched by PO number, line, and release | 28 | 58 |
| Screen 3 | 57 to 86 | 57 and 58 | 28 | 86 |
| Last | 2,371 to 2,400 | Rows carried over from the screen before | The remainder | 2,400 |
The figures are illustrative; the scroll distance depends on the grid and the row height.
Containment
A mismatch fails step 5, and the failure is contained. The partial list is discarded, so the expediting steps never see 2,391 lines dressed as 2,400. The retry reads the report again from the top. If the count still disagrees, the workflow's policy sends it to a person, with screenshots of the grid and of the printed total, because a report that disagrees with itself twice is a question for someone who knows the ERP.
The row key is chosen when the step is taught. PO number, line, and release identify an open line uniquely; a key that could repeat (supplier and due date, say) would merge two real lines into one, and the count check would catch it.
Shortcut
If the ERP can export the same report to a file, the agent can run the export and read the file instead; files are ordinary work for an agent that uses the desktop.
The ledger exists for grids that offer no export, exports that leave out a column the buyer needs, and portals that show data they will not let you download.
Your part
Most of the VLI's work is visible during the run and after the fact, not configured by hand.
You do not position clicks or write selectors. The choices you make are the ones a manager would make: which workflows run on which machines, which applications to train, and what a step should do when it meets something it does not recognize.
Architecture
The VLI is the agent's hands and eyes. Everything else decides what the hands do and keeps the record.
Governance and security
An agent that can operate any application needs fences that do not depend on the application.
Metrics
These numbers come from run history and the per-action record. Read them per application and per screen, because a problem on one ERP form hides inside a fleet average.
| Measure | What it tells you |
|---|---|
| Unresolved targets per screen | Where look-alike controls or cramped layouts need a clearer step description |
| Verify failures per application | Which applications changed, lag, or throw unexpected dialogs |
| Model calls per run | Cost and latency; it falls as maps are trained and routes are reused |
| Seconds per step, minutes per run | The speed side of the tradeoff, measured against the person's time for the same work |
| Rows read against report totals | Whether long reads reconcile, run after run |
| Retraining events per vendor release | The maintenance cost of each application |
Tradeoffs
LaunchAI gives up latency to win coverage.
Vision costs more per step than an API call. The agent looks, understands, acts, and verifies, every time. Three facts keep that cost in proportion:
Through the screen, LaunchAI can do the work every narrower method does, plus the work none of them reach. Where a narrower method covers a subset faster or is already paid for, that is a legitimate reason to use it there.
| Method | Per-step speed | What it reaches | When it can make sense |
|---|---|---|---|
| Vendor API or integration platform | Fastest per call | The objects and actions the vendor chose to expose | High-volume transfers between systems with good APIs; a LaunchAI step can call the same API and use the screen for the rest of the job |
| RPA with selectors | Fast on stable screens | Applications with stable, exposed selectors | Licenses and developers already in place; LaunchAI covers the same screens plus remote desktops, terminals, and custom controls, verifying every action |
| Cloud browser agents | Fast in parallel | Web pages | Public-web collection at scale; LaunchAI covers the same pages from your own session, plus every desktop application |
| LaunchAI VLI | Slower per step | Any application a person can use | Work that crosses systems, lacks APIs, or needs per-action evidence |
Questions
LaunchAI promises no accuracy percentage, and a vendor that does is quoting a benchmark, not your screens. Perception fails at the edges: low contrast, tiny fonts, two nearly identical labels. The design answer is to refuse ambiguity, verify every action, and stop at the step that failed. Test it on your own screens during a pilot.
Yes, as another window on the machine. No element tree crosses the connection, so the accessibility layer sits out and the other three layers do the work. Display scaling and color settings vary between remote sessions, so test on your own during a demo.
It reads them on screen, like any other window. Crooked scans and handwriting go down the model path, which returns typed values checked against a schema. A read below the confidence threshold goes to a person with the document beside the value, and is not posted until that person confirms it.
The process owner or IT starts it, and LaunchAI's support engineers help on request; retraining after a redesign is part of what support covers beyond break-fix. The process owner then re-shows any step whose screen moved. Rules and tables need no change, because a redesign moves screens, not policy.
It depends on the workflow. Exact-rule decisions make none. Targets resolved by the accessibility tree, OCR, or the local detector make none. Model calls come from reading messy documents, judging unfamiliar screens, and targets the local layers cannot settle. Training an application removes most of the calls that navigation would otherwise need.
The design is the same on all three: capture, interpret, decide, act, verify. What differs is the accessibility layer each operating system publishes, and how much of it a given application fills in. Confirm your operating system version before a pilot.
No prompts. No babysitting. No dumb questions.