Our latest work investigates conditions under which large language models attempt to mislead ...
Our latest work investigates conditions under which large language models attempt to mislead users or evaluators — and what that means for safe deployment at scale.
We work on technical problems in AI safety — including robustness, interpretability, and eval...
We work on technical problems in AI safety — including robustness, interpretability, and evaluation — alongside a global network of researchers and partner organizations.
FAR AI investigates concrete problems: how AI models behave under adversarial conditions, how...
FAR AI investigates concrete problems: how AI models behave under adversarial conditions, how deception manifests, and how to evaluate systems reliably before deployment.