Source · Tier A · Peer-reviewed

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

U. Anwar, A. Saparov, J. Rando, D. Paleka, M. Turpin, P. Hase, E. S. Lubana, E. Jenner, S. Casper, O. Sourbut, B. L. Edelman, Z. Zhang, M. Günther, A. Korinek, J. Hernandez-Orallo, L. Hammond, E. Bigelow, A. Pan, L. Langosco, T. Korbak, H. Zhang, R. Zhong, S. Ó hÉigeartaigh, G. Recchia, G. Corsi, A. Chan, M. Anderljung, L. Edwards, A. Petrov, C. Schroeder de Witt, S. R. Motwani, Y. Bengio, D. Chen, P. H. S. Torr, S. Albanie, T. Maharaj, J. Foerster, F. Tramèr, H. He, A. Kasirzadeh, Y. Choi, D. Krueger. 2024. Transactions on Machine Learning Research.

Linkhttps://openreview.net/forum?id=oVTkOs8Pka
arXiv2404.09932
VersionPublished in TMLR (2024); arXiv preprint 2404.09932. Venue and author list checked 2026-09-24 against the ML Anthology record of the TMLR paper, which gives "Sumeet Ramesh Motwani" (the arXiv metadata reads "Motwan").
Accessed2026-09-23
Imported fromhodgkins-ai-verification-papers@c71e59ff0e8e
NoteListed under "Motivations and policy proposals" in the Hodgkins bibliography (CC BY 4.0).

Cited by

Not yet cited by any record. It is part of the bibliography.