1000: Ten Years of the Super Data Science Podcast, with Jon, Kirill and Special Guests
For this landmark 1,000th episode and the show’s 10-year anniversary, host Jon Krohn is joined by SuperDataScience founder Kirill Eremenko, who hosted the podcast for its first 400-plus episodes before handing over the reins. In a first for the show, the episode was recorded live with the audience invited to join on air, alongside surprise appearances from the team, longtime guests, and even Jon’s family. Together, Jon Krohn and Kirill look back on a decade of the podcast and field listener questions on AI’s biggest opportunities, the build-versus-buy dilemma, how to break into the field today, and how to stay grounded amid the relentless pace of AI.
Additional materials: www.superdatascience.com/1000
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
917: 8 Steps to Becoming an AI Engineer, with Kirill Eremenko
Founder of SuperDataScience, Kirill Eremenko, talks to Jon Krohn about how he found the best tools and approaches to help launch his 8-week AI engineering bootcamp. He breaks down the topics participants cover each week, and he also shares his tips with listeners who might want to start their own tech bootcamp or sign up for SuperDataScience’s September 2025 cohort.
This episode is brought to you by the Dell AI Factory with NVIDIA and by ODSC, the Open Data Science Conference
Additional materials: www.superdatascience.com/917
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(10:58) Weeks 1-4 of the SuperDataScience bootcamp
(37:52) How to use AI to drive the bottom line in business
(47:50) Weeks 5-8 of the SuperDataScience bootcamp
(54:50) How to convert LLMs to agents
(1:09:33) Jon’s feedback on the SuperDataSciencebootcamp
899: Landing $200k+ AI Roles: Real Cases from the SuperDataScience Community, with Kirill Eremenko
Data science skills, a data science bootcamp, and why Python and SQL still reign supreme: In this episode, Kirill Eremenko returns to the podcast to speak to Jon Krohn about SuperDataScience subscriber success stories, where to focus in a field that is evolving incredibly quickly, and why in-person working and networking might give you the edge over other candidates in landing a top AI role.
Additional materials: www.superdatascience.com/899
This episode is brought to you by Adverity, the conversational analytics platform and by the Dell AI Factory with NVIDIA.
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(04:35) Stories from five SuperDataScience subscribers
(27:32) How to secure a career in a fast-paced industry
(44:19) How to stand out against huge competition in data science
(1:01:40) The importance of communication in data science
(1:16:41) Where to focus your skills in AI engineering
853: Generative AI for Business, with Kirill Eremenko and Hadelin de Ponteves
Kirill Eremenko and Hadelin de Ponteves AI educators, whose courses have been taken by over 3 Million students, sit down with Jon Krohn to talk about how foundation models are transforming businesses. From real-world examples to clever customization techniques and powerful AWS tools, they cover it all.
bravotech.ai - Partner with Kirill & Hadelin for GenAI implementation and training in your business. Mention the “SDS Podcast” in your inquiry to start with 3 complimentary hours of consulting.
This episode is brought to you by ODSC, the Open Data Science Conference. Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(07:00) What are foundation models?
(15:45) Overview of the foundation model lifecycle: 8 main steps.
(29:11) Criteria for selecting the right foundation model for business use.
(41:35) Exploring methods to customize foundation models.
(53:04) Techniques to modify foundation models during deployment or inference.
(01:11:00) Introduction to AWS generative AI tools like Amazon Q, Bedrock, and SageMaker.
Additional materials: www.superdatascience.com/853
786: The Six Keys to Data Scientists' Success, with Kirill Eremenko
Learn about the six keys to data science success as host Jon Krohn welcomes back Kirill Eremenko, the mastermind behind SuperDataScience. Kirill shares his top insights on data science careers, from building strong portfolios to leveraging mentors and hands-on labs. With over 2.7 million students, his advice is a must-hear for aspiring and experienced data scientists alike.
Additional materials: www.superdatascience.com/786
Interested in sponsoring a SuperDataScience Podcast episode? Visit passionfroot.me/superdatascience for sponsorship information.
771: Gradient Boosting: XGBoost, LightGBM and CatBoost, with Kirill Eremenko
Kirill Eremenko joins Jon Krohn for another exclusive, in-depth teaser for a new course just released on the SuperDataScience platform, “Machine Learning Level 2”. Kirill walks listeners through why decision trees and random forests are fruitful for businesses, and he offers hands-on walkthroughs for the three leading gradient-boosting algorithms today: XGBoost, LightGBM, and CatBoost.
This episode is brought to you by Ready Tensor, where innovation meets reproducibility, and by Data Universe, the out-of-this-world data conference. Interested in sponsoring a SuperDataScience Podcast episode? Visit passionfroot.me/superdatascience for sponsorship information.
In this episode you will learn:
• All about decision trees [09:17]
• All about ensemble models [21:43]
• All about AdaBoost [36:47]
• All about gradient boosting [45:52]
• Gradient boosting for classification problems [59:54]
• Advantages of XGBoost [1:03:51]
• LightGBM [1:17:06]
• CatBoost [1:32:07]
Additional materials: www.superdatascience.com/771
759: Full Encoder-Decoder Transformers Fully Explained, with Kirill Eremenko
Encoders, cross attention and masking for LLMs: SuperDataScience Founder Kirill Eremenko returns to the SuperDataScience podcast, where he speaks with Jon Krohn about transformer architectures and why they are a new frontier for generative AI. If you’re interested in applying LLMs to your business portfolio, you’ll want to pay close attention to this episode!
This episode is brought to you by Ready Tensor, where innovation meets reproducibility, by Oracle NetSuite business software, and by Intel and HPE Ezmeral Software Solutions. Interested in sponsoring a SuperDataScience Podcast episode? Visit passionfroot.me/superdatascience for sponsorship information.
In this episode you will learn:
• How decoder-only transformers work [15:51]
• How cross-attention works in transformers [41:05]
• How encoders and decoders work together (an example) [52:46]
• How encoder-only architectures excel at understanding natural language [1:20:34]
• The importance of masking during self-attention [1:27:08]
Additional materials: www.superdatascience.com/759
747: Technical Intro to Transformers and LLMs, with Kirill Eremenko
Attention and transformers in LLMs, the five stages of data processing, and a brand-new Large Language Models A-Z course: Kirill Eremenko joins host Jon Krohn to explore what goes into well-crafted LLMs, what makes Transformers so powerful, and how to succeed as a data scientist in this new age of generative AI.
This episode is brought to you by Intel and HPE Ezmeral Software Solutions, and by Prophets of AI, the leading agency for AI experts. Interested in sponsoring a SuperDataScience Podcast episode? Visit passionfroot.me/superdatascience for sponsorship information.
In this episode you will learn:
• Supply and demand in AI recruitment [08:30]
• Kirill and Hadelin's new course on LLMs, “Large Language Models (LLMs), Transformers & GPT A-Z” [15:37]
• The learning difficulty in understanding LLMs [19:46]
• The basics of LLMs [22:00]
• The five building blocks of transformer architecture [36:29]
- 1: Input embedding [44:10]
- 2: Positional encoding [50:46]
- 3: Attention mechanism [54:04]
- 4: Feedforward neural network [1:16:17]
- 5: Linear transformation and softmax [1:19:16]
• Inference vs training time [1:29:12]
• Why transformers are so powerful [1:49:22]
Additional materials: www.superdatascience.com/747
671: Cloud Machine Learning
Get to grips with AWS, Azure, Google Cloud Platform on this week’s episode. Host Jon Krohn speaks with Kirill Eremenko and Hadelin de Ponteves about CloudWolf, a cloud computing educational platform that prepares students for certification in AWS (Amazon Web Services). Find out why an accreditation in cloud computing could be the safest investment for your data science career.
This episode is brought to you by Posit, the open-source data science company, and by AWS Inferentia. Interested in sponsoring a SuperDataScience Podcast episode? Visit JonKrohn.com/podcast for sponsorship information.
In this episode you will learn:
• About CloudWolf [07:04]
• Why learning the cloud is important for data scientists [09:12]
• Is learning cloud computing complex? [22:30]
• Essential AWS services [28:31]
• Database options on AWS [33:47]
• How to run analytics on AWS [40:58]
• Why an AWS certification is so helpful [56:35]
Additional materials: www.superdatascience.com/671
649: Introduction to Machine Learning
Looking for a short primer on Machine Learning concepts? SDS Founder Kirill Eremenko and AI expert Hadelin de Ponteves are back, joining Jon Krohn to review essential ML concepts. From classification errors to logistic regression, feature scaling, the elbow method and more. The popular data science instructors also introduce their latest course: Machine Learning in Python: Level 1.
In this episode you will learn:
• Kirill and Hadelin's new course [17:34]
• Supervised vs unsupervised learning [26:23]
• False positives and false negatives [31:21]
• Logistic regression [43:00]
• Holding out a set of test data [46:39]
• Feature scaling [52:45]
• The Adjusted R-Squared metric [59:44]
• The five assumptions of linear regression [1:05:12]
• The Elbow Method [1:11:41]
Additional materials: www.superdatascience.com/649
Interested in sponsoring a SuperDataScience Podcast episode? Visit JonKrohn.com/podcast for sponsorship information.
471: 99 Days to Your First Data Science Job
Kirill Eremenko returns to the SDS podcast as a guest to debunk common myths you may believe about getting a data science job.
In this episode you will learn:
What has Kirill been up to? [3:48]
The genesis of the 99-days challenge [5:27]
5 myths about pursuing a data science career [15:49]
First data science jobs [1:00:53]
5 components for success [1:08:19]
Additional materials: www.superdatascience.com/471