Skip to content
Read the original: a16z News· Published 46/100AI score46/100

a16z backs Preference Model, which builds RL environments for training AI models

Original titleInvesting in Preference Model

AISummary

Preference Model is open-sourcing Karotte, the framework it uses to build reinforcement learning environments that resist reward hacking, including defenses like killing stray processes before grading and rejecting grader-crashing files.

The framework has been hardened through more than a million evaluation runs and controlled red-teaming.

The company focuses on machine learning engineering tasks for leading labs, and a16z says it is partnering with Preference Model and its founders, Jennifer Zhou and Ning Cao.

Read the original a16z.news

Source: a16z News · a16z.newsPublished · added here