
SPONSOR MESSAGES: *** Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich. Dr. Max Bartolo from Cohere discusses the gap between model capabilities and genuine robustness: why next-token prediction can produce impressive results yet still fail on slightly reformulated… --- TIMESTAMPS: 00:00:00 Model Reasoning and Consistency Verification 00:03:25 Influence Functions and Distributed Knowledge Analysis 00:10:28 AI Application Development and Model Deployment 00:14:24 AI Alignment and Human Feedback Limitations 00:20:15 Human Evaluation Challenges and Factuality Assessment 00:27:15 Cultural and Demographic Influences on Model Behavior 00:32:43 Adversarial Examples and Model Robustness 00:41:54 DynaBench and Dynamic Benchmarking Approaches 00:50:02 Benchmarking Challenges and Data-Centric Evaluation 00:55:15 Cohere Command A Development Process 01:00:26 Model Quantization and Performance Evaluation 01:05:18 Reasoning Capabilities and Training Progression 01:13:48 Context Windows and Enterprise Applications --- REFERENCES: person: [00:00:00] Max Bartolo Website https://www.maxbartolo.com/ company: [00:00:00] Cohere https://cohere.com/command [00:12:10] Command A Model https://huggingface.co/CohereForAI/c4ai-command-a-03-2025 paper: [00:03:25] Procedural Knowledge in Pretraining Drives Reasoning in LLMs https://cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20 [00:04:15] Influence Functions in Machine Learning https://arxiv.org/abs/1703.04730 [00:08:05] Studying Large Language Model Generalization with Influence Functions https://arxiv.org/abs/2308.03296 [00:16:15] Human Feedback is not Gold Standard https://arxiv.org/abs/2309.16349 [00:27:15] The PRISM Alignment Dataset https://arxiv.org/abs/2404.16019 [00:32:50] Adversarial Examples Are Not Bugs, They Are Features https://arxiv.org/abs/1905.02175 [00:43:00] DynaBench: Rethinking Benchmarking in NLP https://aclanthology.org/2021.naacl-main.324.pdf [00:50:15] Sara Hooker on Compute Limitations https://arxiv.org/html/2407.05694v1 [00:53:25] DataPerf: Benchmarks for Data-Centric AI https://arxiv.org/abs/2207.10062 [01:04:35] DROP: A Reading Comprehension Benchmark https://arxiv.org/abs/1903.00161 [01:07:05] GSM8k https://paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k [01:09:30] ARC-AGI Challenge https://github.com/fchollet/ARC-AGI --- LINKS: Full Transcript: https://app.rescript.info/share/163bf5e7338685f635fc0b8d6920005c Download PDF transcript: https://app.rescript.info/api/public/sessions/7167a679366f97f8/pdf REFS: [00:03:10] Research at Cohere with Laura Ruis et al., Max Bartolo, Laura Ruis et al. https://cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20 [00:04:15] Influence functions in machine learning, Koh & Liang https://arxiv.org/abs/1703.04730 [00:08:05] Studying Large Language Model Generalization with Influence Functions, Roger Grosse et al. https://storage.prod.researchhub.com/uploads/papers/2023/08/08/2308.03296.pdf [00:11:10] The LLM ARChitect: Solving ARC-AGI Is A Matter of Perspective, Daniel Franzen, Jan Disselhoff, and David Hartmann https://github.com/da-fr/arc-prize-2024/blob/main/the_architects.pdf [00:12:10] Hugging Face model repo for C4AI Command A, Cohere and Cohere For AI https://huggingface.co/CohereForAI/c4ai-command-a-03-2025 [00:13:30] OpenInterpreter https://github.com/KillianLucas/open-interpreter [00:16:15] Human Feedback is not Gold Standard, Tom Hosking, Max Bartolo, Phil Blunsom https://arxiv.org/abs/2309.16349 [00:27:15] The PRISM Alignment Dataset, Hannah Kirk et al. https://arxiv.org/abs/2404.16019 [00:32:50] How adversarial examples arise, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry https://arxiv.org/abs/1905.02175 [00:43:00] DynaBench platform paper, Douwe Kiela et al. https://aclanthology.org/2021.naacl-main.324.pdf [00:50:15] Sara Hooker's work on compute limitations, Sara Hooker https://arxiv.org/html/2407.05694v1 [00:53:25] DataPerf: Community-led benchmark suite, Mazumder et al. https://arxiv.org/abs/2207.10062 [01:04:35] DROP, Dheeru Dua et al. https://arxiv.org/abs/1903.00161 [01:07:05] GSM8k, Cobbe et al. https://paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k [01:09:30] ARC, François Chollet https://github.com/fchollet/ARC-AGI [01:15:50] Command A, Cohere https://cohere.com/blog/command-a [01:22:55] Enterprise search using LLMs, Cohere https://cohere.com/blog/commonly-asked-questions-about-search-from-coheres-enterprise-customers