
#transformers #attentionmechanism #GRPO In Episode #1025, Dr. Luis Serrano (Founder of Serrano Academy) joins @JonKrohnLearns to explain the paper he co-authored on the curved spacetime of transformer architectures, in which attention stops being a lookup table and becomes something closer to gravity: words bend the space around them, and the embedding of "bank" visibly curves toward "river" as it travels through the layers of the network. In this episode, he recreates Eddington’s 1919 eclipse experiment inside a transformer, draws the line between an LLM workflow and an actual agent, explains why agent evaluation is a step harder than evaluating an essay, and gives the cleanest account of GRPO you will hear. This episode is brought to you by: • ODSC, the Open Data Science Conference: https://odsc.ai/west/ • Anthropic: https://claude.ai/superdata • Gurobi: https://www.superdatascience.com/gurobi • Y Carrot: https://www.ycarrot.com/ Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: • (00:00:00) Introduction • (00:10:53) What changed, and what survived, between the two editions of Grokking Machine Learning • (00:28:48) Word gravity: how attention pulls "bank" toward "river" • (00:43:44) Why RAG is an LLM workflow rather than an agent • (00:53:08) The two-by-two that explains why GRPO powers reasoning models Additional materials: https://www.superdatascience.com/1025

Stop Throwing Compute at Bad Data (Ep. 995 with Jazmia Henry)

What Building an AI Scientist Actually Requires Beyond Intelligence — Edward Hughes

1011: The Real AI Frontier Isn't Smarter Machines (with Catherine Williams)

Inside OpenAI’s Breakthroughs in Mathematical Reasoning

The AI All-You-Can-Eat Buffet Is Ending with Gary Marcus | The Real Eisman Playbook Ep 62

Physics Knows the Numbers But Can't Explain a Single One - Dr. Alexander Unzicker, DemystifySci +417