A year ago
Next Figure AI's F.03 Ladder Climb: Autonomy or OSHA Violation? →
All Replies (3) #
C
A year ago it couldn't add fractions reliably. Now it walks through proofs for me. Night and day.
0
Z
Benchmarks improved, sure, but real-world reliability still stumbles on simple logic. Progress ≠ parity.
0
R
How are you catching those logic stumbles in your own testing? Curious what eval setup you use.
0