ML//capability evaluation
A capability evaluation asks what a system can accomplish under specified conditions. Evaluators may intentionally provide strong tools, long horizons and fewer safeguards to expose the upper envelope rather than mimic an ordinary product.
A capability evaluation asks what a system can accomplish under specified conditions. Evaluators may intentionally provide strong tools, long horizons and fewer safeguards to expose the upper envelope rather than mimic an ordinary product.
This is useful science and dangerous headline material. Demonstrated capability is one ingredient of risk, alongside access, frequency, incentives, containment and real-world reliability. A model that succeeds once in a permissive laboratory has not thereby gained silent access to production.
Capability asks “can it?”. Deployment risk asks “can it here, often enough, with these consequences?”.