Fixed-scope service

Find where your AI agent fails before your users do.

A fixed-scope reliability audit for browser agents and autonomous workflows. I test your critical flows, expose silent failures, and deliver evidence-backed findings, verification designs, and a remediation plan you can act on.

$750 fixed price · One system · Up to three critical workflows

The verification engine is public and MIT licensed: skynet-desktop, @zek21/skynet-cli. See a real failure caught, step by step.

The failure your tests do not catch

The problem is not that your automation breaks loudly. It is that it can report success while silently failing, and you have no reliable way to prove its output was correct.

What you receive

  1. Your critical workflows reproduced and instrumented, up to three.
  2. A map of every silent-failure and unverifiable-success path found in them.
  3. Evidence from actual test runs, with artifacts, not a consulting memo.
  4. A severity-ranked failure map you can hand straight to an engineer.
  5. The verification gates and receipts those workflows need, designed and prototyped where the scope allows.
  6. A concrete remediation plan.
  7. One handoff session covering what failed, what passed, and what should change.

Target: five business days from the point where access and scope are ready.

What an audit actually looks like

A real sequence from this site, run on 9 August 2026, while shipping the page you are reading:

  1. The run. A template change was deployed to the live server.
  2. The apparent result. The deploy reported success and proved it: the installed file hashed byte-for-byte identical to the source.
  3. The verification check. A separate probe fetched the public page and read what a visitor would actually be served.
  4. The discrepancy. The served page was still the previous version. The file was deployed; the page was not live. The tool's own success signal could not tell the difference.
  5. The artifact. The check failed closed and named the exact missing element, instead of reporting "deployed".
  6. The remediation. The gate held the "done" claim until an independent fetch of the rendered HTML matched what was promised.

That gap between "the tool said it worked" and "the user got it" is the entire subject of this audit.

Proof, and what it does not prove

The verification engine behind this work is public and you can read it before you buy anything:

Being explicit about what this evidence is: every example on this page comes from my own systems and open-source work. It demonstrates engineering capability, not customer results. I have no client case studies yet, and I am not going to dress up logos or numbers to suggest otherwise. Until I do, the public engineering evidence and audit artifacts above are the proof available to evaluate this work.

How it runs

  1. Intake. You describe the system and the workflows that matter most. No credentials at this stage.
  2. Scope and access. We agree in writing what is covered, and you get an invoice before work starts. Staging or test environments are preferred; read-only wherever it is possible.
  3. Test runs. I reproduce the workflows and instrument them until the failure modes are visible rather than theoretical.
  4. Findings. Severity-ranked map, artifacts, and the verification design.
  5. Handoff. One working session, then the written plan is yours to execute or to hand back to me as a separate build.

Scope and price

$750 fixed

Included: one system, up to three critical workflows, the test runs and their artifacts, the severity-ranked failure map, the verification design, the remediation plan, and the handoff session.

Not included: implementation beyond a small verification prototype, ongoing monitoring, and anything that is really a security assessment or penetration test. If the work grows past the agreed scope, it is re-quoted before it starts rather than absorbed quietly.

Fair questions

Do you have client case studies?

No. The public evidence comes from my own systems and open-source work, and I say so rather than hiding it.

What if nothing fails?

You get the evidence of what was tested, what passed, where coverage is still thin, and what verification would keep it that way. A clean result you can prove is worth having.

Do you fix everything you find?

No. The audit proves and prioritises the problems and prototypes the verification. Larger implementation is scoped and quoted separately, so a fixed-price audit never turns into an open-ended build.

Is this a security audit?

No. This is reliability and verification work. It is not penetration testing and it is not a security assurance sign-off.

What access do you need?

As little as will do the job: a staging or test environment where possible, read-only access where that is enough, and a written note of what is retained and for how long. Never paste secrets into the contact form.

Who does the work

Exzil Calanza. You are hiring a person who operates a system, not an anonymous AI company. Based in the Philippines (GMT+8), with overlap into US Pacific business hours. More about the operator and the track record.

Start with the workflow that worries you most

Tell me what the agent is supposed to do and how you would know it truly did it. That is enough to scope the audit.

Chat with us
Hi, I'm Exzil's assistant. Want a post recommendation?