Page 2: Research news on AI alignment

AI alignment examines how artificial systems acquire, represent, and act on goals, values, and social norms, and why their behavior often diverges from human expectations. Work in this area studies systematic failures such as bias, sycophancy, hallucinations, deceptive or selfish reasoning, and cultural or linguistic inequities, as well as limitations in commonsense, emotion, and social understanding. It also develops methods for preference learning, norm-following, interpretability, and reliability guarantees to better align AI behavior with human values and societal constraints.

Security

How attackers persuade AI agents to break the rules

Today, most of us interact with AI assistants—reactive bots that wait for human instructions. Yet AI assistants are rapidly being replaced by agentic AI agents that can interact with external tools, browse the web, generate ...

page 2 from 40