DAIR.AI 创始人 Elvis Saravia 用 AI 员工 Viktor 解决 Agent 评测结果排查耗时的问题:每晚跑 harness 实验后,Viktor 会检查全部失败日志,把 23 个昨日通过、今日失败的任务追溯到具体改动并建议回滚,由他本人复核决策。Viktor 还会主动标记问题,现可免费试用并获 $100 额度、无需绑卡。
Reading eval results is now the slowest part of building agents.
I'm Elvis, founder of @dair_ai. I lead research, build, and teach about AI agents.
I run harness experiments every night, but reading the results was eating my mornings.
Every change to my harness gets evaluated overnight, whether it touches memory, tool use, or context compaction.
The morning after is the hard part. I check which tasks my agent got right yesterday but wrong today. Then I open the logs for each failure, one by one, to figure out which of my changes caused it.
I tried a dashboard first. It showed the pass rate dropped. It couldn't tell me why.
That is the job Viktor, an AI employee in Slack, is built for. He reviews the results overnight.
Here is how that plays out. Say 23 tasks that passed yesterday fail today. Viktor checks all 23 logs, traces them to the one change that caused them, and suggests undoing it. I check the logs and make the call.
Viktor does the digging. I decide what goes into the harness.
He is also proactive. He flags problems before you ask, which helps you stay on track with complex eval runs and other research tasks.
Harness engineers, do you check every eval run, or only when the pass rate drops?
Try free at @viktor_com. $100 in credits, no card. Full link in my first reply.
Thanks to the team for partnering with me on this post
来源:elvis · x.com