Skip to content
View original post on X: elvis· 18/100AI score18/100

Viktor, a Slack AI employee, reviews overnight agent eval failures

AISummary

Elvis Saravia describes using Viktor, an AI employee in Slack, to review his nightly agent harness evaluation results. Viktor traces tasks that regressed from passing to failing back to the specific harness change that caused them and suggests reverting it, while the human makes the final decision. The post is a sponsored partnership, offering $100 in free credits with no card required.

Post on XView on X
@omarsar0

Reading eval results is now the slowest part of building agents.

I'm Elvis, founder of @dair_ai. I lead research, build, and teach about AI agents.

I run harness experiments every night, but reading the results was eating my mornings.

Every change to my harness gets evaluated overnight, whether it touches memory, tool use, or context compaction.

The morning after is the hard part. I check which tasks my agent got right yesterday but wrong today. Then I open the logs for each failure, one by one, to figure out which of my changes caused it.

I tried a dashboard first. It showed the pass rate dropped. It couldn't tell me why.

That is the job Viktor, an AI employee in Slack, is built for. He reviews the results overnight.

Here is how that plays out. Say 23 tasks that passed yesterday fail today. Viktor checks all 23 logs, traces them to the one change that caused them, and suggests undoing it. I check the logs and make the call.

Viktor does the digging. I decide what goes into the harness.

He is also proactive. He flags problems before you ask, which helps you stay on track with complex eval runs and other research tasks.

Harness engineers, do you check every eval run, or only when the pass rate drops?

Try free at @viktor_com. $100 in credits, no card. Full link in my first reply.

Thanks to the team for partnering with me on this post

Source: elvis · x.comPublished · added here