Start where “done” is checkable and undo is possible.
Part 5 was the cold water. This part is the filter. Agents are not a personality upgrade for every workflow. They pay off where a goal can be tested, tools can be narrow, and a human can still say no.
The technical beat is fit criteria, not a vendor list.
A short test before you buy the demo
Ask four questions:
- Checkable goal — Can a stranger tell whether the job finished? (tests pass, ticket filed, row reconciled)
- Limited tools — Can you name the five actions it may take — and exclude the rest?
- Human gates — Where does money, reputation, or production change hands?
- Rollback — If it is wrong, can you undo or at least notice fast?
If you cannot say what “done” looks like, you do not have an agent job yet. You have a demo.
Where the criteria usually hold
Software development. Tests, refactors, PR drafts. The scoreboard is CI. Review before merge is a natural gate. Rollback is git revert.
IT operations. Diagnostics and runbooks. Propose the restart; do not flip the cluster without a pageable human. Logs are already the culture.
Customer support. Gather account context, draft replies, escalate edge cases. Send is the gate. The CRM is a bounded tool if you lock writes.
Research. Collect sources, cluster notes, stop before “final truth” claims. Citation and human reading are the product. Agents shine at fetch-and-organize, not at being the oracle.
Finance operations. Reconcile and flag. Never auto-wire money without policy that would embarrass you in an audit. Checkable diffs beat fluent summaries.
Data analysis. Explore, chart, summarize. Keep PII and credentials out of casual tools. A notebook plus review is closer to a copilot with extra steps — still a good fit if queries are allowlisted.
Business workflows. Multi-step busywork with clear inputs and outputs: “move this invoice from received to coded,” not “handle finance.” Audit logs are the feature.
Where they usually do not
Open-ended strategy. Irreversible legal send. Anything whose success is “the customer felt heard” with no transcript standard. Anything whose tools include “the entire internet plus every internal admin API.”
Example
Good: “Open a draft PR that adds a null check matching these three failing tests.”
Bad: “Make the app better.”
Good: “List invoices unmatched for 14 days; do not pay.”
Bad: “Handle accounts payable.”
Same company. Different fitness.
Conclusion
Fit is a design choice. The next three parts assume you passed this filter: prompt engineering and policy, evaluation and testing, shipping and deployment. Skip them and you are back in demo land.
Takeaway: Agents make sense where goals are checkable, tools are limited, humans gate irreversible steps, and rollback exists.
Part 5: The Uncomfortable Reality
Part 7: Prompt Engineering and Policy