How this started
The agents got good. The trust didn’t.
Browser agents can already click through your apps. Give one a task and it will find the buttons. That is real progress, and it has a flaw you notice on the second day: it acts live. You learn what it changed after it changed it, from a screenshot if you are lucky.
The other path is to wire the agent into each app by hand: tools, guardrails, a review step. That works, and it costs an integration per app, per team, per quarter.
EEVEE is our answer: a workbench with one key. The agent builds the tool and works it through typed actions. Every write shows you its consequence before it happens. Only your passkey lets anything publish or land.
- 01
What did it build?
Every version is kept. You see a live preview and the source on request, never a mystery bundle.
- 02
Does it work?
Scenarios run the app in a real browser: fill, click, restart, assert. You read verdicts, not code.
- 03
What will this change?
Each write is rehearsed on a copy of current data. The card shows the exact fields, before and after.
The Decisions ledger in this demo already works this way. It shows the diff before you ask.
Independent hackathon prototype for The WebMCP Challenge. Meridian Ops, its customers, and its numbers are fictional and seeded.
