Turing Seminar – An introduction to AGI Safety
L. PILLAUD-VIVIEN, L.DANA
Opening

Objectif du cours

The rapid advancements in artificial intelligence are showing no signs of slowing down. From GPT-2 to GPT-5, we remain uncertain about where this race for performance is leading us, though many experts warn of potentially catastrophic risks. The highly publicized launches of powerful chatbots are only the tip of the iceberg. In the first part of the course, we will review these advancements and explore where they might lead us. Informed by scaling laws and expert forecasts, we will see that Artificial General Intelligence (AGI) could emerge within this decade.

While these technologies are enabling breakthroughs in biomedical research, breaking down language barriers, and reducing administrative burdens, significant technical challenges remain in building general-purpose models that are both reliable and safe. The second part of the course will discuss the benefits and risks of future AGIs.

These risks stem from both our limited understanding of the models we create and the lack of international coordination to avoid competitive races and ensure the safe deployment of such powerful systems. Assessing which dangerous capabilities a model has—and how they can be elicited—requires careful evaluation, often of closed-source systems. The final part of the course will examine solutions for controlling AIs in a robust way: understanding, evaluating, and regulating them. We will explicitly discuss how engineers can contribute through research and machine learning expertise.

Organisation des séances

7x2hours sessions + Project. Before each session, students must read pedagogical resources listed on the syllabus [Link].

Activities during sessions: Presentations of papers, expert interventions, research projects, and sometimes debates and discussions.

Mode de validation

Grading 100% Project. (See the Syllabus)

Références

This course will be based on several general resources presented here:

  • The AI Safety Atlas is a textbook developed by Markov Grey & Charbel-Raphël Segerie for the French Center for AI Safety, and used around the world for teaching the risks associated with the development of safe AIs. Most sessions of the Seminar are based on the textbook’s chapters.
  • The Introduction to AI Safety, Ethics, and Society is the textbook developed by Dan Hendrycks for the Center for AI Safety to be used by teachers around the world. It gives the core principles of the security of AI systems.
  • The BlueDot Impact courses on AI Alignment and Governance compile key introductory resources on GPAI for students around the globe.
  • The Artificial Intelligence Index Report 2025, a resource from Stanford University’s Human-Centered Artificial Intelligence, which compiles key metrics and trends on the development of AI.
  • The International AI Safety Report is a report commissioned by the UK AI safety summit to Prof. Yoshua Bengio’s team to scientifically ground the risks and advances of AI systems.

 

You can also view the recorded past Seminars on youtube: 2024 edition.

Les intervenants

Loucas PILLAUD-VIVIEN

(ENPC Cermics)

Léo DANA

(INRIA Sierra)

voir les autres cours du 1er semestre