TL;DR
To hire an AI systems builder: search four title families (forward deployed engineer, head of AI operations, founding builder/GTM engineer, automation architect), then screen on six signals — shipped unattended systems, scheduled autonomy, failure stories with process fixes, cross-stack fluency, business judgment, and a verification reflex. Respect the anti-signal: if the role's core is SDK internals or distributed systems, hire a software engineer instead. Then skip the panel circuit and run the 48-hour test: hand the candidate a real backlog problem and see if a working, verified pipeline comes back.
I sit on both sides of this hire. I run five brands' worth of agent infrastructure as one person, and I maintain a scored pipeline of the roles companies post when they're trying to buy that capability. That pipeline taught me something uncomfortable for hiring managers: most job descriptions for this role screen for the wrong things, under the wrong titles, with interviews that a fluent talker passes and a real builder finds insulting. So here's the playbook I'd hand any founder or ops lead who wants to hire an AI systems builder — the screens that work, the anti-signal most companies ignore, and the one test that replaces the entire panel.
If you're still deciding whether this is the role you need at all, start with what an AI systems builder actually does — this post assumes you've decided and now have to find one.
Step one: stop searching for the title
Nobody's résumé says "AI systems builder." The work ships under at least four title families, and your sourcing search has to cover all of them:
| Title family | Where it lives | What the work looks like |
|---|---|---|
| Forward deployed engineer / AI solutions engineer | AI labs, AI-native startups | Embed with a customer, ship working agent systems on their real stack |
| Head of AI operations / AI automation lead | Mid-size companies adopting agents | Own internal agent adoption: workflow selection, SOPs, verification |
| Founding builder / GTM engineer | Early-stage startups | Build the revenue systems themselves — content engines, funnels, outbound — with agents as the team |
| Automation architect | Ops-heavy orgs | Starts from n8n/Zapier-class plumbing, grows into agent orchestration |
The titles differ; the four underlying disciplines — workflow decomposition, agent operating procedures, infrastructure glue, verification — are the constant. (If the FDE-versus-operations distinction matters for your org chart, I've broken it down in forward deployed engineer vs. AI operations lead.) When you source, search all four families. When you screen, ignore what the last job was called and ask what's still running from it.
The screening rubric: six signals
Vocabulary is cheap now. Anyone who's spent a month on AI Twitter can talk about agents, orchestration, and evals. Screen for evidence instead — these six, in rough order of how hard they are to fake:
1. Shipped systems, not demos
Ask for something running in production today, unattended, that the candidate built. "I prototyped a chatbot" is a demo. "This engine has published on schedule for months — here are the posts, here's the schedule" is a system. Demos prove the person can start; systems prove the person can finish and then leave. The one-question version, which I use as the core screen: does the work continue when the person stops typing?
2. Scheduled autonomy
A subset of the first signal, worth isolating because it's so easy to check. Can the candidate show automation that fires on cron, launchd, or CI without a human trigger — and, critically, explain how they verified it's actually armed? Anyone who's run scheduled agents for real has a story about a job that silently never fired. If they've never been burned by a scheduler, they haven't run one long enough.
3. Failure stories with process fixes
This is my favorite interview question because it's unfakeable. Ask: "Tell me about a time an agent fabricated something, drifted off-spec, or failed silently — and what you changed." Everyone who has operated agents in production has these stories. The tell is in the fix. "I watch it more closely now" is a weak answer — vigilance doesn't scale. "I encoded the correction into a permanent procedure the fleet follows" is the strong answer, because it means the person manages agents the way a good operator manages a team: with written SOPs, not supervision. If a candidate claims they've never had an agent fail on them, end the interview politely; they either haven't shipped or won't tell you the truth about it.
4. Cross-stack fluency
Agent systems touch everything: payments, hosting, e-commerce platforms, email infrastructure, automation platforms, browser automation. You want breadth over any single deep specialty — someone fluent enough across the stack to wire agents into all of it, not a career specialist in one layer. Probe by walking their proudest system end to end and counting how many distinct services it touches. A real one usually touches five or more.
5. Business judgment
The scarce skill isn't wiring APIs. It's knowing which workflow is worth automating, what "good output" means for your specific business, and where a human must stay in the loop. Look for operating history — has this person actually run a function like content, publishing, or lead handling, with their own name on the outcome? — not just technical history. A builder without operating judgment will automate the wrong things beautifully.
6. A verification reflex
Ask: "How do you know a system's output is correct?" Then be quiet. If the answer doesn't include inspection gates, defined completion criteria, and outputs viewed before they ship, the systems this person builds will drift — and drifting agent systems produce confident garbage at scale. This failure mode is common enough that I wrote it up separately in the verification gap in production agents; it's the most underrated screen on this list because it's the discipline that separates a system that runs from a system you can trust.
The honest anti-signal: sometimes you need a SWE instead
Here's the part most hiring guides skip, because it costs the guide's author work: this role is not a software engineer, and if your role fundamentally is one, hire a software engineer.
If the core of your job is hands-on production software engineering — SDK internals, distributed systems, ML model training, deep API and infrastructure engineering — then a systems builder is the wrong hat, no matter how impressive the agent fleet. I apply this gate to myself, in writing, in my own application pipeline: roles that are fundamentally SWE or ML engineering wearing an FDE label get classified as rejects, and roles that want production-code depth under a solutions title get flagged as long shots rather than fits. My edge is orchestrating agent systems with business judgment; it is not writing production software the way a career SWE does, and pretending otherwise would waste both sides' time.
That boundary cuts both ways for you as the hiring manager. The best candidates for this role will name it unprompted — "that part of your stack needs a real engineer, not me" — and that honesty is itself a hiring signal. The candidates who claim to be both a senior distributed-systems engineer and an agent-fleet operator are usually neither. Write your job description so the two roles can't be confused: if you list Kubernetes internals and agent orchestration in the same requirements block, you'll attract people who are bluffing on one of them.
The 48-hour test: the interview that isn't one
Panels are theater for this role. A fluent talker passes them; a builder resents them; neither outcome tells you what you need to know, which is whether this person leaves running systems behind.
So replace the panel with the thing itself. Pull a real problem from your backlog — the reporting workflow nobody has time for, the content operation running on manual effort, the lead-handling process held together by one overloaded person — and give the candidate 48 hours with it. It's the close I use in my own interviews: give me a real problem from your backlog and 48 hours, and I'll ship a working system before you finish interviewing other candidates.
What you're grading when the clock runs out:
- Did a working pipeline come back — or a deck about one? The deck-versus-system distinction is the whole role, compressed into two days.
- Did it come with SOPs? A real builder ships the operating procedures alongside the pipeline, because they know unattended systems drift without them.
- Did it come with verification gates? Look for defined completion criteria and inspection points, not just "it ran."
- What questions did they ask first? A systems builder's opening questions are about the business — what does good output look like, who consumes it, what must never ship wrong. A tinkerer's opening questions are about tooling.
- What did they refuse to automate? The strongest candidates will fence off part of the problem as "this stays human." That's business judgment showing itself under time pressure.
Pay for the 48 hours if the candidate is senior — it's real work and treating it as free spec labor filters out exactly the experienced people you want. Two days of contractor-rate pay is the cheapest de-risking you will ever buy on a hire whose comp, across these title families' public listings, clusters roughly $150k–$400k.
One caution: the test evaluates the candidate, but it also evaluates you. If you can't produce a well-scoped backlog problem in a day, you're not ready to manage this hire — the role consumes well-scoped problems, and a company that can't supply them will turn a systems builder into a very expensive chatbot consultant.
FAQ
How do I interview an AI systems builder?
Skip most of the panel circuit. Screen on six signals — shipped unattended systems, scheduled autonomy, failure stories with process fixes, cross-stack fluency, business judgment, and a verification reflex — then run a paid 48-hour test on a real backlog problem. Grade the artifact: working pipeline, SOPs included, verification gates defined.
What's the biggest red flag when hiring for this role?
A candidate with no failure stories. Anyone who has operated agents in production has watched one fabricate, drift, or fail silently. Candidates who claim otherwise either haven't shipped or won't be honest about it. The matching green flag is a failure story that ends in a permanent process fix, not increased vigilance.
When should I hire a software engineer instead of an AI systems builder?
When the role's core is hands-on production software: SDK internals, distributed systems, ML model training, deep infrastructure engineering. An AI systems builder orchestrates agent systems with business judgment; that's a different discipline from production software engineering, and the best builders will tell you so themselves.
Should the 48-hour test be paid?
Yes, at least for senior candidates. It's two days of real work, and unpaid spec work filters out experienced people who can afford to decline it. Contractor-rate pay for two days is cheap insurance on a six-figure hire.
What job titles should I post or search to find an AI systems builder?
Post or search forward deployed engineer, AI solutions engineer, head of AI operations, AI automation lead, automation architect, and GTM engineer or founding builder. The same four disciplines ship under all of these labels — screen on what's still running from the last role, not on the title.
RELATED
Everything in this playbook is a standing offer. Bring a real problem from your backlog and I'll ship a working, verified pipeline against it in 48 hours — get in touch.
— Italo Campilii. If you're building something that needs this kind of operator, get in touch.