Why IBM Is Betting Everything on Small AI Models
In this episode of Eye on AI, Craig Smith sits down with Sriram Raghavan, Vice President of AI at IBM Research, to explore one of the most important debates in enterprise AI right now. Do you actually need a massive model to get world class results? IBM's answer is no, and Sriram breaks down exactly why.
Sriram explains why IBM chose to train its Granite models directly using reinforcement learning rather than distilling from larger models like most of the industry. The reason goes beyond performance. It comes down to data lineage, safety alignment, and a belief that small, efficient models are the only sustainable path for enterprises running AI across hybrid cloud environments.
We get into the full technical stack behind that bet. How data quality has replaced model size as the real competitive advantage. Why parameter count is becoming the wrong metric entirely. How IBM's inference time scaling techniques allow an 8 billion parameter model to match the performance of GPT-4o and Claude 3.5 on code and math benchmarks. And why IBM is pioneering a new concept called Generative Computing, which treats AI models not as prompt receivers but as programmable computing elements with runtimes, modular LoRA adapters, and proper programming abstractions.
Sriram also shares where IBM Research is headed next, including breakthroughs in continuous learning, agent orchestration, and making unstructured enterprise data actually usable at scale.
Subscribe for more conversations with the people building the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Why IBM Skips Distillation and Trains Small Models Directly
(04:50) Did We Even Need Giant AI Models in the First Place?
(08:12) How Data Quality Became the New Competitive Moat
(11:54) Why Parameter Count Is the Wrong Way to Measure a Model
(15:36) Reinforcement Learning Without Losing Broad Capabilities
(22:05) Inference Time Scaling: Getting Big Model Results From Small Models
(28:12) Generative Computing: Treating AI as a Programming Element
(36:40) Why IBM Open Sources and How Small Models Make It Sustainable
(41:25) The Path to Continuous Learning Without Rewriting Weights
(51:00) IBM's Full Roadmap: Models, Data, and Agents
Major breakthroughs in artificial intelligence research often reshape the design and utility of Al in both business and society. In this special rebroadcast episode of Smart Talks with IBM, Malcolm
Gladwell and Jacob Goldstein explore the conceptual underpinnings of modern Al with Dr. David Cox, VP of Al models at IBM Research. They talk foundation models, self-supervised machine learning, and the practical applications of Al and data platforms like watsonx in business and technology.
When we first aired this episode last year, the concept of foundation models was just beginning to capture our attention. Since then, this technology has evolved and redefined the boundaries of what's possible. Businesses are becoming more savvy about selecting the right models and understanding how they can drive revenue and efficiency.
This is a paid advertisement from IBM. The conversations on this podcast don't necessarily represent IBM's positions, strategies or opinions.
Visit us at ibm.com/smarttalks
See omnystudio.com/listener for privacy information.
Major breakthroughs in artificial intelligence research often reshape the design and utility of AI in both business and society. In this episode of Smart Talks with IBM, Malcolm Gladwell and Jacob Goldstein explore the conceptual underpinnings of modern AI with Dr. David Cox, VP of AI Models at IBM Research. They talk foundation models, self-supervised machine learning, and the practical applications of AI and data platforms like watsonx in business and technology.
Visit us at: https://www.ibm.com/thought-leadership/smart/talks/
Learn more about watsonx: https://www.ibm.com/watsonx
This is a paid advertisement from IBM.
See omnystudio.com/listener for privacy information.
Our show is all about heroes making great strides in technology. But in InfoSec, not every hero expects to ride off into the sunset. In our series finale, we tackle vulnerability scans, how sharing information can be a powerful tool against cyber crime, and why it’s more important than ever for cybersecurity to have more people, more eyes, and more voices, in the fight.
Wietse Venema gives us the story of SATAN, and how it didn’t destroy the world as expected. Maitreyi Sistla tells us how representation helps coders build things that work for everyone. And Mary Chaney shines a light on how hiring for a new generation can prepare us for a bold and brighter future.
If you want to read up on some of our research on the InfoSec community, you can check out all our bonus material over at redhat.com/commandlineheroes. Follow along with the episode transcript.
Today we’re joined by Doug Burdick, a principal research staff member at IBM Research. In a recent interview, Doug’s colleague Yunyao Li joined us to talk through some of the broader enterprise NLP problems she’s working on. One of those problems is making documents machine consumable, especially with the traditionally archival file type, the PDF. That’s where Doug and his team come in.
In our conversation, we discuss the multimodal approach they’ve taken to identify, interpret, contextualize and extract things like tables from a document, the challenges they’ve faced when dealing with the tables and how they evaluate the performance of models on tables. We also explore how he’s handled generalizing across different formats, how fine-tuning has to be in order to be effective, the problems that appear on the NLP side of things, and how deep learning models are being leveraged within the group.
The complete show notes for this episode can be found at twimlai.com/go/541
Today we’re joined by Yunyao Li, a senior research manager at IBM Research.
Yunyao is in a somewhat unique position at IBM, addressing the challenges of enterprise NLP in a traditional research environment, while also having customer engagement responsibilities. In our conversation with Yunyao, we explore the challenges associated with productizing NLP in the enterprise, and if she focuses on solving these problems independent of one another, or through a more unified approach.
We then ground the conversation with real-world examples of these enterprise challenges, including enabling level document discovery at scale using combinations of techniques like deep neural networks and supervised and/or unsupervised learning, and entity extraction and semantic parsing to identify text. Finally, we talk through data augmentation in the context of NLP, and how we enable the humans in-the-loop to generate high-quality data.
The complete show notes for this episode can be found at twimlai.com/go/537
Today we’re joined by Noam Slonim, the principal investigator of Project Debater at IBM Research.
In our conversation with Noam, we explore the history of Project Debater, the first AI system that can “debate” humans on complex topics. We also dig into the evolution of the project, which is the culmination of 7 years and over 50 research papers, and eventually becoming a Nature cover paper, “An Autonomous Debating System,” which details the system in its entirety.
Finally, Noam details many of the underlying capabilities of Debater, including the relationship between systems preparation and training, evidence detection, detecting the quality of arguments, narrative generation, the use of conventional NLP methods like entity linking, and much more.
The complete show notes for this episode can be found at twimlai.com/go/495.
Blake has a PhD in physics from Yale and is the quantum platform lead. You can find him on Twitter here and read some of his recent writing here.
Robert is VP of IBM Quantum Ecosystem Development, IBM Research. He's the author of Dancing with Qubits and has put together a great list of tutorial videos on his website.
No Lifeboat badge winner today, but if you're a fan of Schrödinger's cat, be sure to check out this question from our Quantum Computing Stack Exchange.
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
There has been a debate in the past few years between the symbolists and the connectionists about the future of artificial intelligence. The symbolists say that traditional, explainable, logic-based approaches still hold tremendous promise while the connectionists say that the power of deep learning, for all its current opacity and narrow application, holds the key to more general forms of machine intelligence. This week, I speak with David Cox, IBM Director of the MIT-IBM Watson AI Lab, which is blending the two traditions in what they call neuro-symbolic AI in hopes to move AI forward.
Today we’re joined by Jannis Born, Ph.D. student at ETH & IBM Research Zurich, to discuss his “PaccMann^RL” research. Jannis details how his background in computational neuroscience applies to this research, how RL fits into the goal of anticancer drug discovery, the effect DL has had on his research, and of course, a step-by-step walkthrough of how the framework works to predict the sensitivity of cancer drugs on a cell and then discover new anticancer drugs.
This week, I talk to Ken Church, a pioneer in Natural Language Processing, whose use of statistical models on part of speech tagging revolutionized the field and is what makes automatic dictation and machine translation so popular today. We talked about his early days at MIT, about explainable AI and about how the Holy See played a role in his probabilistic approach to NLP.
In this episode of the SuperDataScience Podcast, I chat with Dr. Guillermo Cecchi about the role of data science in medical research and maybe even the future of artificial intelligence. You will learn how data science and artificial intelligence are pushing the boundaries of mental healthcare. You will hear some very interesting approaches about getting insights from audio samples of patients’ voices and their speech and you will also learn about the development of some fascinating techniques, like transferring intuitive knowledge from professionals in the healthcare field into algorithms.
If you enjoyed this episode, check out show notes, resources, and more at www.superdatascience.com/241
Ajay Royyuru and Guillermo Cecchi from IBM Healthcare join Chris and Daniel to discuss the emerging field of computational psychiatry. They talk about how researchers at IBM are applying AI to measure mental and neurological health based on speech, and they give us their perspectives on things like bias in healthcare data, AI augmentation for doctors, and encodings of language structure.
Sponsors:
Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com.
Rollbar – We catch our errors before our users do because of Rollbar. Resolve errors in minutes, and deploy your code with confidence. Learn more at rollbar.com/changelog.
Linode – Our cloud server of choice. Deploy a fast, efficient, native SSD cloud server for only $5/month. Get 4 months free using the code changelog2018. Start your server - head to linode.com/changelog
Algolia – Our search partner. Algolia’s full suite search APIs enable teams to develop unique search and discovery experiences across all platforms and devices. We’re using Algolia to power our site search here at Changelog.com. Get started for free and learn more at algolia.com.
Featuring:
Ajay Royyuru – Website
Guillermo Cecchi –
Chris Benson – Website, GitHub, LinkedIn, X
Daniel Whitenack – Website, GitHub, X
Show Notes:
IBM 5 in 5: With AI, our words will be a window into our mental health
Predicting Cognitive Impairments with a Mobile Application
Automated analysis of recent-onset and prodromal schizophrenia
Prediction of psychosis across protocols and risk cohorts using automated language analysis
Upcoming Events:
Register for upcoming webinars here!
How can we make AI that people actually want to interact with? Raphael Arar suggests we start by making art. He shares interactive projects that help AI explore complex ideas like nostalgia, intuition and conversation -- all working towards the goal of making our future technology just as much human as it is artificial.
Hosted on Acast. See acast.com/privacy for more information.
"Hold your breath," says inventor Tom Zimmerman. "This is the world without plankton." These tiny organisms produce two-thirds of our planet's oxygen -- without them, life as we know it wouldn't exist. In this talk and tech demo, Zimmerman and cell engineer Simone Bianco hook up a 3D microscope to a drop of water and take you scuba diving with plankton. Learn more about these mesmerizing creatures and get inspired to protect them against ongoing threats from climate change.
Hosted on Acast. See acast.com/privacy for more information.
Driving in Johannesburg one day, Tapiwa Chiwewe noticed an enormous cloud of air pollution hanging over the city. He was curious and concerned but not an environmental expert -- so he did some research and discovered that nearly 14 percent of all deaths worldwide in 2012 were caused by household and ambient air pollution. With this knowledge and an urge to do something about it, Chiwewe and his colleagues developed a platform that uncovers trends in pollution and helps city planners make better decisions. "Sometimes just one fresh perspective, one new skill set, can make the conditions right for something remarkable to happen," Chiwewe says. "But you need to be bold enough to try."
Hosted on Acast. See acast.com/privacy for more information.
Nearly every other year the transistors that power silicon computer chip shrink in size by half and double in performance, enabling our devices to become more mobile and accessible. But what happens when these components can't get any smaller? George Tulevski researches the unseen and untapped world of nanomaterials. His current work: developing chemical processes to compel billions of carbon nanotubes to assemble themselves into the patterns needed to build circuits, much the same way natural organisms build intricate, diverse and elegant structures. Could they hold the secret to the next generation of computing?
Hosted on Acast. See acast.com/privacy for more information.