How to evaluate an AI engineering portfolio
Demos are easy and evaluation is hard. Here is how to tell a shipped AI system from an impressive prototype.
The barrier to building something impressive with a foundation model is now very low. The barrier to knowing whether it works is unchanged. That gap is where portfolio evaluation should focus.
What distinguishes a shipped system from a demo?
A shipped AI system has an evaluation set, documented failure modes, latency and cost figures, and a defined behaviour for when the model is wrong. A demo has a happy path. The fastest way to tell them apart is to ask how the candidate knew the system was working, and listen for whether the answer involves measurement or impression.
What questions surface real experience?
Ask what the system did when the model was confidently wrong. Production AI work is largely about handling that case: fallbacks, confidence thresholds, human review paths, and guardrails. A candidate who has shipped will answer immediately with a specific incident. A candidate who has prototyped will treat it as hypothetical.
Other high-signal questions:
- How did you build your evaluation set, and how did it change over time?
- What did you measure before and after a prompt or retrieval change?
- Where did retrieval quality break down, and how did you diagnose it?
- What did the system cost to run, and what did you do about it?
- What did you deliberately decide not to use a model for?
That last question is quietly one of the best. Engineers with real production experience have usually removed a model from somewhere it was not earning its cost.
How much does framework experience matter?
Framework experience matters less than evaluation discipline, because the frameworks change faster than the underlying skills. An engineer who has built rigorous evaluation for a system using one toolchain will transfer to another quickly. An engineer who knows a specific framework deeply but has never measured output quality will not.
Does research background predict engineering performance?
A research background predicts different strengths, not higher ones. Research experience is valuable when the role involves training or adapting models. For applied roles built on foundation models, product engineering experience with strong measurement habits is usually the better predictor, and requiring a research background narrows the pool without improving the outcome.
How should the loop test this?
Give a practical exercise with a deliberately ambiguous quality bar and see whether the candidate defines one. The strongest applied AI engineers will ask what "good" means and propose how to measure it before writing much. That instinct is the single most transferable skill in the discipline, and it is visible within twenty minutes.
Frequently asked questions
What should I look for in an AI engineer portfolio?
Look for evidence of evaluation, not demonstration. A shipped AI system has an evaluation set, known failure modes, latency and cost figures, and a story about what happened when the model was wrong. A portfolio of impressive demos with none of that describes prototyping ability rather than production experience.
Are side projects useful signal for AI roles?
Side projects are useful when they include measurement. A project with a documented evaluation set and honest limitations is stronger signal than a polished demo, because it shows the habit that production AI work depends on. Polish alone mostly demonstrates familiarity with a framework.
Related reading
ML engineer or applied AI engineer? Hiring for the right role
These two titles attract different candidates and solve different problems. Getting the distinction wrong is the most common reason AI searches stall.
Hiring data platform engineers for games and AI teams
Data platform roles get written as analytics jobs and filled by analytics people. Here is how to brief and screen for the engineering version.
How to hire gameplay engineers without slowing your build
Gameplay engineers are judged on feel, not just code. Here is what to screen for, how to structure the loop, and the signals that separate a shipper from a strong interviewer.
Hiring for a gaming or AI team?
Talentfinders delivers shortlists in three to five days, on a success-based fee. United States, onsite and remote.
Start a search