MCP, Agents and What AI Engineers Are Thinking About Right Now feat. Swyx
What's at the top of AI engineers' minds? Swyx, organizer of the AI Engineer Summit and host of Latent Space, joins us to discuss MCP (Model Context Protocol), the rise of AI agents, and how the role of AI engineers is evolving.
Find our guest online:
https://x.com/swyx
https://www.ai.engineer/
https://www.latent.space/
Get Ad Free AI Daily Brief: https://patreon.com/AIDailyBrief
Brought to you by:
KPMG – Go to https://kpmg.com/ai to learn more about how KPMG can help you drive value with our AI solutions.
Vanta - Simplify compliance - https://vanta.com/nlw
Plumb - The Automation Platform for AI Experts - https://useplumb.com/nlw
The Agent Readiness Audit from Superintelligent - Go to https://besuper.ai/ to request your company's agent readiness score.
The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Subscribe to the newsletter: https://aidailybrief.beehiiv.com/Join our Discord: https://bit.ly/aibreakdown
Scaling Airbyte: Challenges and Milestones on the Road to 1.0
Summary
Airbyte is one of the most prominent platforms for data movement. Over the past 4 years they have invested heavily in solutions for scaling the self-hosted and cloud operations, as well as the quality and stability of their connectors. As a result of that hard work, they have declared their commitment to the future of the platform with a 1.0 release. In this episode Michel Tricot shares the highlights of their journey and the exciting new capabilities that are coming next.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
Your host is Tobias Macey and today I'm interviewing Michel Tricot about the journey to the 1.0 launch of Airbyte and what that means for the project
Interview
Introduction
How did you get involved in the area of data management?
Can you describe what Airbyte is and the story behind it?
What are some of the notable milestones that you have traversed on your path to the 1.0 release?
The ecosystem has gone through some significant shifts since you first launched Airbyte. How have trends such as generative AI, the rise and fall of the "modern data stack", and the shifts in investment impacted your overall product and business strategies?
What are some of the hard-won lessons that you have learned about the realities of data movement and integration?What are some of the most interesting/challenging/surprising edge cases or performance bottlenecks that you have had to address?
What are the core architectural decisions that have proven to be effective?How has the architecture had to change as you progressed to the 1.0 release?
A 1.0 version signals a degree of stability and commitment. Can you describe the decision process that you went through in committing to a 1.0 version?
What are the most interesting, innovative, or unexpected ways that you have seen Airbyte used?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on Airbyte?
When is Airbyte the wrong choice?
What do you have planned for the future of Airbyte after the 1.0 launch?
Contact Info
LinkedIn
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
Links
AirbytePodcast Episode
Airbyte Cloud
Airbyte Connector Builder
Singer Protocol
Airbyte Protocol
Airbyte CDK
Modern Data Stack
ELT
Vector Database
dbt
FivetranPodcast Episode
MeltanoPodcast Episode
dlt
Reverse ETL
GraphRAGAI Engineering Podcast Episode
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
AI Engineers, Pendants, and Competition Between OpenAI and Developers with Swyx of Latent Space
In this episode, Nathan sits down with Swyx of Latent Space to chat about AI engineers and tools to check out, competitive dynamics between OpenAI, other foundation model providers, and developers, and if Swyx would wear an AI Horcrux. We also learn that Nathan lived in the same dorm as Mark Zuckerberg back in his college days (and was a late adopter to Facebook…). If you're looking for an ERP platform, check out our sponsor, NetSuite: http://netsuite.com/cognitive
Subscribe to the show Substack to join future conversations and submit questions to upcoming guests! https://cognitiverevolution.substack.com/
We're hiring across the board at Turpentine and for Erik's personal team on other projects he's incubating. He's hiring a Chief of Staff, EA, Head of Special Projects, Investment Associate, and more. For a list of JDs, check out: eriktorenberg.com.
SPONSORS: Shopify | NetSuite | Omneky
Shopify is the global commerce platform that helps you sell at every stage of your business. Shopify powers 10% of ALL eCommerce in the US. And Shopify's the global force behind Allbirds, Rothy's, and Brooklinen, and 1,000,000s of other entrepreneurs across 175 countries.From their all-in-one e-commerce platform, to their in-person POS system – wherever and whatever you're selling, Shopify's got you covered. With free Shopify Magic, sell more with less effort by whipping up captivating content that converts – from blog posts to product descriptions using AI. Sign up for $1/month trial period: https://shopify.com/cognitive
NetSuite has 25 years of providing financial software for all your business needs. More than 36,000 businesses have already upgraded to NetSuite by Oracle, gaining visibility and control over their financials, inventory, HR, eCommerce, and more. If you're looking for an ERP platform ✅ head to NetSuite: http://netsuite.com/cognitive and download your own customized KPI checklist.
Omneky is an omnichannel creative generation platform that lets you launch hundreds of thousands of ad iterations that actually work customized across all platforms, with a click of a button. Omneky combines generative AI and real-time advertising data. Mention "Cog Rev" for 10% off.
TIMESTAMPS:
(00:00) Episode Preview
(00:00:49) AI Nathan’s intro
(00:03:14) What is an AI engineer?
(00:05:56) What backgrounds do AI engineers typically have?
(00:15:51) Sponsors: Netsuite | Omneky
(00:17:13) Swyx’s Discord AI project
(00:20:41) Key tools for AI engineers
(00:23:42) HumanLoop, Guardrails, Langchain
(00:27:01) Criteria for identifying capable AI engineers when hiring
(00:30:59) Skepticism around AI being a fad and doubts about contributing to AI
(00:34:03) AI Engineer Conference speaker lineup
(00:41:14) AI agents and two years to AGI
(00:46:04) Expectations and disagreement around what AI agent capabilities will work soon
(00:50:12) Swyx’s OpenAI thesis
(00:53:03) AI safety considerations and the role of AI engineers
(00:56:24) Disagreement on whether AI will soon be able to generate code pull requests
(01:01:07) AI helping non-technical people to code
(01:01:49) Multi-modal Chat-GPT and the future implications
(01:03:33) Nathan living in the same dorm as Mark Zuckerberg
(01:04:44) Competitive dynamics between OpenAI and other AI model developers
(01:05:39) Play.ht vs ElevenLabs
(01:09:20) The tension between platforms and developers building on top of them
(01:11:40) The best thing startups can do to compete with foundation model providers
(01:16:26) User identity/authentication services like Login with OpenAI
(01:19:20) Google vs the other live players
(01:20:46) AI Horcruxes / Pendants
(01:22:05) The concept of an AI app bundle for consumers and developers
Code Interpreter is GPT-4.5: A Summer AI Technical Roundup [feat. Swyx and Alessio of Latent Space]
Today NLW is joined by Swyx and Alessio, the hosts of the Latent Space podcast to discuss the key technical developments from the last month of AI, including code interpreter; llama 2; the latest in AI agents; growing interest in AI companions, and more.
Latent Space podcast -https://www.latent.space/podcast / https://twitter.com/latentspacepod
Swyx - https://twitter.com/swyx
Alessio Fanelli - https://twitter.com/FanaHOVA
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
Twitter: https://twitter.com/nlw / https://twitter.com/AIBreakdownPod
Learning in Public with swyx
About swyx
swyx has worked on React and serverless JavaScript at Two Sigma, Netlify and AWS, and now serves as Head of Developer Experience at Airbyte. He has started and run communities for hundreds of thousands of developers, like Svelte Society, /r/reactjs, and the React TypeScript Cheatsheet. His nontechnical writing was recently published in the Coding Career Handbook for Junior to Senior developers.
Links Referenced:
“Learning Gears” blog post: https://www.swyx.io/learning-gears
The Coding Career Handbook: https://learninpublic.org
Personal Website: https://swyx.io
Twitter: https://twitter.com/swyx
Self Service Open Source Data Integration With AirByte
Summary
Data integration is a critical piece of every data pipeline, yet it is still far from being a solved problem. There are a number of managed platforms available, but the list of options for an open source system that supports a large variety of sources and destinations is still embarrasingly short. The team at Airbyte is adding a new entry to that list with the goal of making robust and easy to use data integration more accessible to teams who want or need to maintain full control of their data. In this episode co-founders John Lafleur and Michel Tricot share the story of how and why they created Airbyte, discuss the project’s design and architecture, and explain their vision of what an open soure data integration platform should offer. If you are struggling to maintain your extract and load pipelines or spending time on integrating with a new system when you would prefer to be working on other projects then this is definitely a conversation worth listening to.
Announcements
Hello and welcome to the Data Engineering Podcast, the show about modern data management
When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With their managed Kubernetes platform it’s now even easier to deploy and scale your workflows, or try out the latest Helm charts from tools like Pulsar and Pachyderm. With simple pricing, fast networking, object storage, and worldwide data centers, you’ve got everything you need to run a bulletproof data platform. Go to dataengineeringpodcast.com/linode today and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
Modern Data teams are dealing with a lot of complexity in their data pipelines and analytical code. Monitoring data quality, tracing incidents, and testing changes can be daunting and often takes hours to days. Datafold helps Data teams gain visibility and confidence in the quality of their analytical data through data profiling, column-level lineage and intelligent anomaly detection. Datafold also helps automate regression testing of ETL code with its Data Diff feature that instantly shows how a change in ETL or BI code affects the produced data, both on a statistical level and down to individual rows and values. Datafold integrates with all major data warehouses as well as frameworks such as Airflow & dbt and seamlessly plugs into CI workflows. Go to dataengineeringpodcast.com/datafold today to start a 30-day trial of Datafold. Once you sign up and create an alert in Datafold for your company data, they will send you a cool water flask.
RudderStack’s smart customer data pipeline is warehouse-first. It builds your customer data warehouse and your identity graph on your data warehouse, with support for Snowflake, Google BigQuery, Amazon Redshift, and more. Their SDKs and plugins make event streaming easy, and their integrations with cloud applications like Salesforce and ZenDesk help you go beyond event streaming. With RudderStack you can use all of your customer data to answer more difficult questions and then send those insights to your whole customer data stack. Sign up free at dataengineeringpodcast.com/rudder today.
Your host is Tobias Macey and today I’m interviewing Michel Tricot and John Lafleur about Airbyte, an open source framework for building data integration pipelines.
Interview
Introduction
How did you get involved in the area of data management?
Can you start by explaining what Airbyte is and the story behind it?
Businesses and data engineers have a variety of options for how to manage their data integration. How would you characterize the overall landscape and how does Airbyte distinguish itself in that space?
How would you characterize your target users?
How have those personas instructed the priorities and design of Airbyte?
What do you see as the benefits and tradeoffs of a UI oriented data integration platform as compared to a code first approach?
what are the complex/challenging elements of data integration that makes it such a slippery problem?
motivation for creating open source ELT as a business
Can you describe how the Airbyte platform is implemented?
What was your motivation for choosing Java as the primary language?
incidental complexity of forcing all connectors to be packaged as containers
shortcomings of the Singer specification/motivation for creating a backwards incompatible interface
perceived potential for community adoption of Airbyte specification
tradeoffs of using JSON as interchange format vs. e.g. protobuf/gRPC/Avro/etc.
information lost when converting records to JSON types/how to preserve that information (e.g. field constraints, valid enums, etc.)
interfaces/extension points for integrating with other tools, e.g. Dagster
abstraction layers for simplifying implementation of new connectors
tradeoffs of storing all connectors in a monorepo with the Airbyte core
impact of community adoption/contributions
What is involved in setting up an Airbyte installation?
What are the available axes for scaling an Airbyte deployment?
challenges of setting up and maintaining CI environment for Airbyte
How are you managing governance and long term sustainability of the project?
What are some of the most interesting, unexpected, or innovative ways that you have seen Airbyte used?
What are the most interesting, unexpected, or challenging lessons that you have learned while building Airbyte?
When is Airbyte the wrong choice?
What do you have planned for the future of the project?
Contact Info
Michel
LinkedIn
@MichelTricot on Twitter
michel-tricot on GitHub
John
LinkedIn
@JeanLafleur on Twitter
johnlafleur on GitHub
Parting Question
From your perspective, what is the biggest gap in the tooling or technology for data management today?
Links
Airbyte
Liveramp
Fivetran
Podcast Episode
Stitch Data
Matillion
DataCoral
Podcast Episode
Singer
Meltano
Podcast Episode
Airflow
Podcast.__init__ Episode
Kotlin
Docker
Monorepo
Airbyte Specification
Great Expectations
Podcast Episode
Dagster
Data Engineering Podcast Episode
Podcast.__init__ Episode
Prefect
Podcast Episode
DBT
Podcast Episode
Kubernetes
Snowflake
Podcast Episode
Redshift
Presto
Spark
Parquet
Podcast Episode
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast