





























Key Insights
- Choose the first AI workflow from repeated operational work, not from the most impressive technology demonstration.
- Exclude workflows that lack ownership, stable rules, trustworthy inputs, observable completion, or controllable impact.
- Score outcome value, volume, stability, observability, access, reversibility, and exception burden with documented evidence.
- Break broad processes into the smallest useful unit that still reaches a business outcome.
- Use your manual baseline and production pilot instead of treating a borrowed productivity benchmark as a forecast.
Your team already has a list of things it wants AI to fix. Sales wants faster follow-up. Operations wants fewer manual updates. Customer service wants the backlog under control. The question is not whether there is work to automate. The question is where to start.
Many teams choose the first workflow by visibility. They pick the process leadership hears about most, the task with the most impressive demo, or the broadest promise on a planning slide. That choice often produces a large project with unclear rules and no agreed definition of done.
Start with the queue, not the technology.
The better first workflow is usually quieter: repeated work with a clear trigger, stable decisions, a visible final state, and an owner who feels the cost of delay. It gives the company a contained place to learn how an agent should act, stop, escalate, and improve.
Your first workflow should prove that the company can assign, control, measure, and improve a real unit of work. The primary outcome is a documented choice with a usable baseline, a safe boundary, and a measurable final state.
That choice begins with evidence of what people actually do each day.
Find repeated work before naming a use case
Ask teams to record repeated work for two weeks. Use queue reports, inbox labels, call reasons, task logs, error reports, and employee observation. Look for work that returns every day or every week, not a special project that happened once.
Operators on Reddit describe a useful pattern: people often name the outcome they want but cannot describe the small decisions inside the work. The fastest way to surface those decisions is to watch a person complete recent cases and ask what changed from one case to the next. Treat that discussion as qualitative research, not a performance benchmark.
Write each candidate as a trigger and final state. Replace automate reporting with: when the weekly data closes, gather these approved figures, place them in this template, validate the totals, and route any mismatch.
That exercise turns a wish list into a set of workflows you can compare. It also exposes the candidates that need process work before they need an agent.
Eliminate workflows that are not ready
Some work should not be the first deployment. Set a candidate aside when any of these conditions is true:
- No one owns the current process.
- Two employees cannot agree on the correct final state.
- The policy changes before the team can document it.
- The required data is routinely unavailable or untrusted.
- Most cases require negotiation, discretion, or relationship judgment.
- A mistake would create high impact and cannot be reviewed or reversed.
- The agent would need broad access that cannot be limited.
- The company cannot observe whether the action succeeded.
This is not a permanent rejection. It is a sequencing decision. Improve the process, data, access, or control, then score it again.
Score seven dimensions
1. Outcome value
What changes when the workflow improves? Use a business measure such as response time, completed appointments, cycle time, backlog, error cost, cash collection, service capacity, or staff time returned. Give the highest score to work with a direct link to a result the operating team already tracks.
2. Volume and repetition
Count eligible cases in a defined period. High-volume work creates more opportunities to learn and spreads implementation cost across more outcomes. Repetition also matters. A frequent workflow with a stable pattern is usually a better first role than a lower-volume process with a different path every time.
3. Process stability
Ask how often the policy, steps, systems, fields, and owners change. Automation does not stabilize a moving process. It makes every change a production change. Score workflows higher when the team can explain the current process and expects the basic shape to remain useful.
4. Observability
Can the agent verify its work? A sent message is not the same as a booked appointment. A submitted form is not the same as an accepted record. Score workflows higher when the final state can be read back from the system and compared with the intended result.
5. Access readiness
List every channel, browser application, file, credential, record, and permission the role needs. Score the workflow higher when access can be limited to the job. Lower the score when the role requires shared administrator credentials, unrestricted data, or an unavailable system owner.
6. Reversibility and impact
A draft can be reviewed. A note can be corrected. A payment, deletion, price commitment, or public message may be harder to reverse. A strong first workflow contains routine, low-impact actions or clear approval gates before material actions.
7. Exception burden
Review recent cases and categorize what broke the normal path. Missing data, duplicate records, unavailable staff, customer requests, policy questions, and system errors all count. A workflow can still be valuable with exceptions, but the exception types need owners and a usable handoff.
Together, these seven dimensions show whether a workflow is valuable and ready. The scorecard turns that judgment into a decision the team can explain.
Use a simple weighted scorecard
Score each dimension from one to five using evidence from the same review period. A practical starting weight is:
- Outcome value: 25 percent.
- Volume and repetition: 15 percent.
- Process stability: 15 percent.
- Observability: 15 percent.
- Access readiness: 10 percent.
- Reversibility and impact: 10 percent.
- Exception burden: 10 percent.
Multiply each score by its weight and add the results. The weights are a decision aid, not an industry standard. Change them to reflect the company's risk and operating goals. Document who scored each workflow and the evidence used.
A high total does not automatically win. Apply the readiness exclusions first. The score ranks viable candidates. It does not make an undefined or unsafe workflow ready.
Compare the smallest useful units
Do not score automate lead management against automate the back office. Those units are too broad. Break them into work that one role can own:
- Contact new leads and record a disposition.
- Qualify an inbound lead and book the approved next step.
- Collect missing intake information.
- Classify one approved document type and update the record.
- Follow up on incomplete appointments.
- Prepare a routine reconciliation and route mismatches.
The smallest useful unit should still reach a business outcome. If it only creates another summary or queue, the employee may inherit as much work as the agent removed.
Once the unit is narrow enough to own, measure how it performs today. Without that baseline, the business case will confuse motion with improvement.
Build the baseline before the business case
For the top candidates, measure eligible volume, human minutes, wait time, completion, rework, errors, and exceptions. Use current records and define the period. If a workflow has revenue value, write the attribution rule before projecting the gain.
Avoid borrowed productivity claims. Research on 5,179 customer-support employees found a 14 percent average increase in issues resolved per hour after access to a generative AI assistant. The same study found that results varied by worker experience. That is useful evidence that AI can change work, but it is not a forecast for an autonomous agent, another industry, or your workflow.
Your baseline and pilot should carry the decision.
Value alone does not make a workflow the best place to begin. The first role should also teach the company something it can reuse.
Choose for learning as well as value
The first workflow should teach the company something reusable. A strong pilot may validate identity matching, one browser system, one communication channel, an approval pattern, an escalation route, or a quality measure that the next role can reuse.
Prefer a workflow that an operating leader cares about and frontline employees understand. Give one person authority to decide policy questions during the pilot. Name a backup for exceptions. The technology cannot compensate for missing ownership.
At this point, the choice should no longer depend on enthusiasm for AI. It should rest on a visible problem, a bounded role, an owner, and evidence. Capture those decisions before the build begins.
Turn the winner into a production brief
The final choice should fit on one page:
- Problem and business outcome.
- Trigger, required inputs, and final state.
- Baseline with source and period.
- Allowed actions and approval gates.
- Systems, channels, files, and credentials.
- Known exceptions and named owners.
- Pilot boundary, measures, and stop conditions.
- Decision date and accountable sponsor.
Vida helps teams turn that brief into an outcome-based AI agent with the communication channels, computer access, skills, controls, handoffs, and observability needed to do the work.
Bring us your top three queues. We will help score them and design a paid pilot around the strongest first outcome. Choose Your First Workflow.
Citations
- Brynjolfsson, Erik, Danielle Li, and Lindsey R. Raymond. "Generative AI at Work." NBER Working Paper 31161, 2023. Referenced for the study of 5,179 support employees and its context-specific 14 percent average productivity finding. https://www.nber.org/papers/w31161
- National Institute of Standards and Technology. "AI RMF Core." Referenced for continuous governance, mapping, measurement, and management of AI risk. https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
- Reddit, r/automation. "Those of you who recently started automating things, what was hard in starting it?" 2026. Used as anecdotal operator research on observing real work and narrowing the first unit. https://www.reddit.com/r/automation/comments/1u1zu8f/those_of_you_who_recently_started_automating/
- Reddit, r/automation. "How to decide what to automate." 2026. Used as anecdotal operator research on maintenance, reconciliation, verification, and process stability. https://www.reddit.com/r/automation/comments/1vhzi4h/how_to_decide_what_to_automate/




