Skip to content
Back to blog

Your AI agent says “done”: how to check its work

Your AI agent says the job is done. Learn how to check the actual result, find supporting evidence and catch missing work before you sign it off.

Stellary Product Desk4 min read

Last reviewed on August 30, 2026

Your AI agent says “done”: how to check its work

“Done.” The new field appears on the form. But leave it empty and submission fails. The client asked for an optional field: the task is not finished.

To verify an AI agent’s work, start with the observable result, not its closing message.

  1. What was requested?

    Return to the original criteria.

  2. What actually works?

    Try the relevant user journey.

  3. What is still unchecked?

    Name the limits, then decide.

A claim is not yet evidence

Match the evidence to the promise. A screenshot may be enough to check alignment. It cannot prove that a form works.

The agent says

“The form is ready”

The new field appears in a screenshot.

Useful evidence

Both submissions succeed

With and without a value in the optional field. The data reaches the right place.

The agent says

“The page is updated”

A file was saved on its machine.

Useful evidence

The right version is accessible

Open the relevant URL. Distinguish the preview from the public site.

Anthropic draws this distinction between an execution trace and the final outcome. A tool call proves an attempt; its effect still needs checking.

Run the journey in four steps

Return to our fictional client email: replace a PDF and add an optional Company field, keeping the price and logo unchanged.

  1. Open the right version. Did the agent change a preview or the public site? Reload the specified URL.
  2. Try both cases. Submit with a company name, then without one. Use test data without creating real registrations or notifying clients.
  3. Check the receiving end. Was the data stored in the right place? Also download the PDF through the page to check its version.
  4. Check what should be unchanged. Are the price and logo intact? A successful change can hide an unrelated edit nobody asked for.

Missing access? Record the limit. Without access to the system receiving registrations, you can check the confirmation screen but not the final stored record.

A handover you can scan

Ask for three short sections. This is a fictional example, not a test result produced for this article.

Illustrative handoverPreview ready for review
Verified
The correct PDF downloads. Both submissions, with and without a company name, were recorded. Price and logo are unchanged.
Not checked
The confirmation email: access to the test inbox was unavailable.
Pending
Your approval to publish. The public site has not changed.

For a real assignment, include the preview URL and test references. “Verified” without accessible evidence leaves you to reconstruct the investigation.

Ask for a verifiable handover

Compare your result with the original request, including anything that should remain unchanged.

Verified: list the checks actually performed and links to the evidence.

Not checked: state what still needs testing and which access is missing.

Pending: list the decisions or permissions needed.

Do not invent evidence. An attempt is not a success; a preview is not a published change. Take no external action without the required authorisation.

Accept, correct or wait

The criteria are met

Accept the work

Evidence is accessible and any remaining limits do not prevent acceptance of the agreed scope.

A result is missing

Request a correction

Name the defect: “Submission fails without a company name.” That is more useful than “check again.”

If an essential check is impossible, leave the task pending. A second AI can help review it, but its agreement does not replace a successful test.

In Stellary, keep acceptance criteria in card checklists and evidence with the task. Define permissions using the AI agents guide.

For recurring assignments, keep a few cases to rerun after changing models: an ordinary request, a missing attachment and contradictory instructions. One successful demonstration is not enough to judge reliability.

Practical questions

Do I need to check everything myself?

Match the review to the risk: read through a rewrite, test a functional journey, and explicitly approve messages or publications before execution.

Is a screenshot enough?

Sometimes, for a visual state. For delivery, stored records or forms, check the behaviour and the result at the receiving end.

What if the agent cannot finish the checks?

It should state what is verified and what remains unknown. A missing essential check prevents treating the task as fully approved.

You might also like

Go further with Stellary

Get started

Ready to pilot your projects with AI?

Stellary brings together your board, docs, and AI agents in one command center.