Page 5: Research news on AI alignment

AI alignment examines how artificial systems acquire, represent, and act on goals, values, and social norms, and why their behavior often diverges from human expectations. Work in this area studies systematic failures such as bias, sycophancy, hallucinations, deceptive or selfish reasoning, and cultural or linguistic inequities, as well as limitations in commonsense, emotion, and social understanding. It also develops methods for preference learning, norm-following, interpretability, and reliability guarantees to better align AI behavior with human values and societal constraints.

Machine learning & AI

Q&A: Neural transparency and the future of AI design

Millions of people are now designing their own personalized artificial intelligence companions, yet most have little idea how those creations will actually behave. In a new paper, MIT Media Lab Assistant Professor Pat Pataranutaporn ...

Computer Sciences

Testing the limits of what's possible (and what isn't) with AI

When can we trust the results we get from AI, and when is learning impossible? Researchers have shown that there are some problems that even the most powerful AI cannot reliably solve, no matter how much data it is given.

Consumer & Gadgets

AI job rejections felt least fair when avatars shared just one trait

Companies are increasingly using artificial intelligence in their hiring processes. It's not just CVs that are evaluated automatically. AI tools can also conduct job interviews—usually in the form of avatars, which are animated ...

Machine learning & AI

Is recursive self‑improvement the dawning of AI superintelligence?

The US AI research company Anthropic has become known for building powerful AI models while simultaneously warning about their dangers. Most recently, its executives wrote about the threat posed by "recursive self-improvement." ...

page 5 from 40