A team of AI researchers gave six language models a task to design a human study, and the result was an unexpected survey question about the last digit of your birth year.
Automated evaluation tools like Scorable can generate custom test suites and metrics from a simple description, making it easier to ensure AI systems work as intended without manual setup.
When AI models read information that contradicts their own knowledge, they often ignore it. This counter-intuitive behavior reveals a fundamental aspect of how these systems work.
A compact 1.5-billion-parameter model outperforms many larger models in competitive math and coding tasks, proving that size isn't everything when it comes to AI performance.
A new study reveals how minimal adversarial changes to text can dramatically degrade AI-based security systems, offering insights for developers and security experts.