Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Researchers introduced Φ-Bench, a new benchmark designed to test whether large language models can engineer the infrastructure that powers them, according to the paper's headline and abstract. The work argues that existi…