FlagEval-Robo tests 12 open-weight embodied AI models across simulation and real robots
Original titleEmbodied AI demos are advancing rapidly, but how much do high benchmark scores actually reflect physical reality?
AISummary
BAAI introduces FlagEval-Robo, an open dual-track evaluation suite linking simulation with real-world execution. The team post-trained and stress-tested 12 leading open-weight embodied AI models under strictly aligned conditions. The post raises whether high benchmark scores reflect physical reality, though it does not yet report specific results.
Source: BAAI · x.comPublished · added here