Please welcome Director of Advocacy and Developer Experience Engineering, Adi Polak. Good morning, Current NOLA, New Orleans 2025. Welcome to the second day. I hope you had a lot of fun yesterday. You got a chance to learn new things as well as have fun in the parties, and as we learn new things and propel ourselves to the future, there are a lot of things history can actually teach us, so I want you to sit back, relax, and join me on a journey back in time to the 1600s, all the way in Western Europe. During that time, Spain and England were continuously battling each other. Spain was the most powerful empire in the world, massive, industrial-scale navy, global wealth. England, on the other side, scrappy, unimpressive, way, way smaller, very much the early-stage startup of Europe. One day, Spain decided to send 141 ships containing almost 30,000 men to crush England. This was the Spanish Armada. The Armada sailed out of Lisbon to link up with a larger army in Belgium. Spain, back in the days, had a superior force, but England had one thing Spain didn't have. They had what we all know today in tech, information velocity. When the Armada appeared off the coast of Cornwall, England decided to light a chain of beacons, a chain of lights. You can imagine it's kind of like an optical pub-sub system with shockingly low latency, and the system was able to inform London just within a couple of hours. To give you context, the distance from London to Cornwall is about 250 miles, while back in 1600, it used to take a good 24 hours for any horse rider, assuming they'll take zero breaks, so the beacon system was super innovative for the time, and it really helped London get informed fast. England started mobilizing their commanders. They coordinated the fleet, and they aligned the whole nation extremely fast. 10 days, just 10 days later, the Spanish Armada was defeated by a much weaker force. From that moment on, that trajectory of these two nations flipped for 400 years. That means that faster information won. Fast forward a couple of centuries, and here we are, architects of real-time systems. In the last decade, we've evolved our systems completely and developed new approaches to how we run software. If you remember the story about microservices versus monolith, we discovered that microservices didn't ruin everything. They actually saved us from the gravitational pull of the monolith. What was back in the day, deploying anything from a monolith, definitely felt like defusing a bomb with sweaty hands and outdated docs. Microservices gave us agility. They gave us independent teams, and they gave us faster delivery. The boundaries of the microservices actually map today to business realities. Yes, they also gave us distributed complexity, observability stack, and probably 19 ways to spell time out in a YAML file, but still net positive. As an industry, we all leveled up. And just as microservices emerged, Kafka showed up as the central nervous system of everything that we do. Kafka basically said, "Hey, you know, everyone, stop panicking. Here's the log. Please put things in this log and read from that log." And for a second, we had clarity. Now, I don't know if you all remember that, but we entered an era with a great data swamp. Essentially, we all thought, you know, just dump it in S3 and pray future us will make sense of it. This started the whole separation of compute and storage, led by Hadoop and HDFS. HDFS is cheap, scalable. But querying HDFS felt like watching paint dry in cold weather. Painful. Back in the days, batch was a king. Latency was a pure rumor. Then Spark arrived and started giving us more agility. Suddenly, we could transform data a little bit faster, slice it, dice it, wrangle it. But in the world of software, things still go wrong. So what happened there? What we started to do is taking everything, all the data that we had, and put it in Parquet files. But Parquet files is not exactly a system. It became a junk drawer. No transactions, no governance, no guarantees. Every job had a different interpretation of the data and the truth. So we entered the great era of table formats. We introduced solutions like Iceberg, like Hudi, like Fluss, like Delta Lake. And now we finally have things that we all know and love from the data warehouse. We have snapshots, we have metadata, we have versioning, we have ACID guarantees, and we have time travel. All was built for the cloud and the new world of separation of storage and compute. Now, this was the moment that data lake turned from being a swamp and became something we could really trust because it gave us solutions like reliable tables, reproducible reads, point-in-time correctness. And that mattered because we know business doesn't always happen once a quarter. It happens day to day. And our analytical system began shifting too, from eventually accurate to consistently accurate, and from stale snapshot to living data sets. Meanwhile, we start streaming by default without even thinking about it. And Neil reached kind of like a similar conclusion that if events arrive now, we want the compute to run now too. Not batch, not later. We want it right now. And so we start building solutions. Kafka Streams, Flink, SQL on Streams, all of them were variations of the same idea. Do the work when reality happens for us. And then we're not done. AI arrived. And suddenly, every architecture that we're building grew three extra boxes, two new databases, and a mysterious agentic layer nobody can quite explain. We went from microservices are hard to let's orchestrate a swarm of autonomous decision-making processes driven by probabilistic models we cannot debug. You can think of it like a single human trying to reason about an AI swarm, is like one SRE babysitting 10,000 microservices at 2:00 A.M. during an outage. Sure, it might work, but if nothing ever goes wrong in the world of software? So we kept on saying as an industry, what we need is better foundation models. We just need better models. And while this is true, there's another layer that we need in order to actually take these codes and move it into production. What we need, we need a layer of auditable systems. We need logs. We need replays. We need guardrails. We need version prompts. We need context engine. We need deterministic fallbacks. And we need real-time truth that we'll be able to trust. Without that trustworthy system, up-to-date context, all AI produces for us is actually a slop. Inconsistent decisions, stale outcomes, things we just cannot explain with a straight face. So that brings us to the real problem that we have in tech, is that principle that kind of governs everything that we do, is that things happen. It happens in the world. It happens in the market. It happens in the system. It happens in our data. AI, unfortunately, is not going to replace that reality for us. It's going to amplify it. Every trigger, every event, every action, all of them are going to trigger another action, and another event, and another outcome, and at the end, it's going to be a complete chain where we cannot know and we cannot understand what is happening as it happens, so if you think about it for a second, microservices taught us to react. Kafka taught us to think in event-driven. Stream processing taught us to compute in the moment. AI, if it's going to be useful, it must be built on top of the same principles. Because when things happen, and they always do, the winner is not the big system. It's the system that is ready. If systems have to be ready for what's next, so do we. Right now, I want you all to get your phone out of your pockets and scan this QR code. This is going to take you to a mobile site. Keep this side handy. Don't close it. Keep it open. For the rest of the time, things are going to happen in this room. Your job is to capture them. Was that a cow? Okay. Capture it, friends. I trust you. Write it down. Explain what happens. Explain what you see. Explain what you hear. Behind the scenes, in our backend system, we have an AI pipeline that is going to capture all of your inputs, and it's going to show and summarize it right in the dashboard. So keep your eyes ready and be open, and listen, friends. I'm trusting you on this one. We need to know what is happening in the room as it happens. Next, we're going to move into a late-night fireside chat. It's going to be a lot of fun, and again, a lot of things are going to happen. So please, I trust you. Good morning. Good morning. Wonderful to see you all here. How are you doing this morning? Great. That's what I want to hear. I mean, Adi was talking about being ready for anything, and imagine actually being a late-night talk show host, which I'm not, but we're going to work together on this. You know, they got a whole room of writers who write the day's jokes, and then some big thing happens, some news item, you know, some economic thing, or one of the late-night hosts starts beefing with the president, and then they have to rewrite all their stuff just right on the spot. Things change. Things happen, and they have to be ready for that, and that is exactly like what we do, like the systems we build, and I'm curious. I could just barely see the lights. See a little bit. But I want to see who consider yourself like an application developer, and that's kind of my background. Any hands? Okay. Not many hands. Data engineer. All right. Good. Analytics person. Yeah. All right. Executive. Okay. Don't be afraid. It's okay. Don't like raising your hand, well, that's by definition, you won't do that. Yeah. So there's a lot of us. We are collectively data streaming engineers, those of us who are practitioners here. You build systems that are always ready for what's next, right? That's the deep structure of the systems we build. They're designed with change in mind. And change hits us. Think about what AI is doing. Now, the stuff that we talked about on stage yesterday, I'm really excited about that kind of stuff, like a real-time context engine. I've been wanting that for years, glued on to Kafka in my favorite managed service. I think that's fantastic. But AI's impact is not just about product lines that folks are developing, however wonderful those products are. There's new patterns. There's new architectures. There's new acronyms, new ways of thinking about software, software that processes now unstructured inputs. So much of our time is spent on creating the structure of the data that we move around and doing predictable things with them. And now we're building AI-based systems that are kind of, if I could use a word, like more stochastic, more probabilistic in the sorts of results they generate. That's a very different way of thinking about software. I mean, to me, that's the province of electrical engineering and things. It's not computer science. But that is a way that these things are happening in our world. And we're having to really change and adapt to new ways of doing things. And when it comes to AI, yeah, you talk to folks. We have no small amount of collective anxiety about the profession, about our job market. And getting into software used to be like the surest of sure bets if you had what it took, right? And now a lot of us know people who are trying to get started as junior developers. It's a little unsure, right? Things are happening. Changes are happening. We're having to adapt to those changes. Frankly, I think AI is creating a tremendous amount of opportunity ahead of us. I'm not at all skeptical there. But it's a changing and it's a different world. We've got some guests here who are going to join us and help us think through that. What does it mean to be a data streaming engineer in the age of AI? What do you need to be ready for what's next? What can we give you from this stage that can help you prepare for that? We've got guests lined up who are going to help figure that out. But before I call them up, I just yesterday. Did you enjoy yesterday? That's what I'm talking about. Yeah. Round of applause for yesterday's speakers. And the party. Who went to the party? Did you hold the alligator? I saw pictures. I know some of you did. The boa constrictor? I didn't. But you know what? I sure had a ball there at the party. It was great. Now, remember the app. I hope you have it open. I just need to help you a little bit here. Weird things happen in the room. Things happen. You have to be ready for those things. Type in a description of what's going on. Every time a weird thing happens, just be on the lookout. Our first guest today, I didn't see her at the party last night, but we're about to see her here on stage. Ladies and gentlemen, Anna McDonald. Yeah. Oh, you missed the back end. Hello. I guess I'm not that cool. Oh, gosh. Now I'm torn. I don't like goats, but I love Neil. I didn't know the goat would have Neil's name on it. Okay. I was going to say, is the goat going to perform? I ate goat last night. Why do you hate goats? When I was little, my great uncle had a farm, and he thought it was hilarious to put me to feed the goats because goats are known for headbutting you in various areas of your body. So I eat them as vengeance. Okay. All right. My neighbor, when I was little, had a baby goat. And he was out of town for a while. I had to nurse it with this bottle. And they jump. Yeah. They're vindictive. Yeah. I didn't know that. Anyway, yeah, this is Neil Buesing. Where's Neil? Neil? I can't. Woo. Round of applause. Neil Buesing. He is the speedrun champion of the Data Streaming Engineer Certification. He's got the fastest time. Now, for the record, he's the only one that's going to be on record. It's not like a metric that we don't rush necessarily. Neil just happens to have it like that. He did. He does. All right. Anna? Are you willing to answer some questions that are not about goats? Yeah. I'm willing to answer the following questions. Let's go. Okay. Let's do it. You're a CISO? No. But close. I am the Distinguished Technical Voice of the Customer, which I did on purpose like a longest title. It really is. So what do you do? What would you say you do here? What would I say I do here? That's a great question. I work directly with engineering and product to make sure that the use cases, because I am all about alright, fly. There's a giant fly. Cheese and rice. To make sure that our use cases and all of our products work for things like bond risk, everything from that, A/B testing, because I am addicted to learning about the world. I love it. I love it. And you have a reputation as being a Kafka Streams person. I know you've done a lot of Flink in the last little bit. What do people do with Flink? You actually talk see, in my role, I talk about the happy path and what could be. You're actually with the people who are building the things. So what are the things that you're? So I think, and I've always said this, systems are kind of like humans in which you can find someone you feel comfortable with immediately. And when that happens, it's beautiful. And if that person and you happen to be going to the same place, you get along very well. And I think it's the same thing. So Flink, and I've always said this, Kafka Streams is Kafka Streams for a reason. It's not just streams. It inherits all the good from Kafka and some of the limitations. So if you have a gigantic key space, Flink is great. You can reshard outside of a partition count. And what we've noticed since we've made our Flink cloud offering is, believe it or not, people don't sometimes want to manage complex distributed systems. It's not something that they enjoy. And one of the things that we've seen be incredibly popular is data pipelines. It's really, really a great way to kind of move from a batch and give much, much faster updates, data streams for your product. So I've enjoyed it. I enjoy my Flink friends. You're in. This message brought to you by Confluent Cloud. Confluent Cloud. That's right. If you call now, only pay shipping and handling. We should have arranged for an 800 number to blink at the bottom of the screen that goes to your phone. Yeah. That sounds pretty cool. Now, and you've been doing some AI things. Yeah. I'm super fascinated by this. Right. So I'll be honest with everyone in this audience. I am not somebody who normally jumps on buzzwords. I know it's shocking, right? You would think that I loved that, but I do not. I've been waiting for frameworks to come out for real vertical use cases, things that we could use. And so Stanford just released a paper on the ACE Framework. If you haven't read it, please go read it. It was dropped. It's like a giveaway, but that's the way I treat white papers. I'm like, "Did you see what just dropped?" It's amazing. I should get like a top 10 books. Some people drop mixtapes. Yeah, exactly. Of white papers. It's really great. And it's about personas where we designate. And it's actually training without having to adjust weights. So yeah. And it combats things like brevity bias, context collapse, which are an absolute no-go for legal and financial use cases. So working that into something like Flink and into kind of like an agent aspect of Flink with the new Flink, I think, is a very natural fit. And I'm excited about it. I'm just thinking. Yes. Real-time sort of streaming retraining. Right. Correct. And so it's not offline. It's online training, right? Basically, you're refining that model. And also, again, please read the white paper because it kicked the crap out of every speed for some of the weight training models and FiNER and all the other kind of custom vertical-specific ones and 87% more performance efficient. Would this be like traditional ML rather than? Kind of, yeah. It's really about adjusting the model in real time, which we really have not been able to do efficiently without some of the side effects I talked about, like biodegradability and stuff like that. Yeah. I want to link to that. We'll have to. I will share one with you. As I say on the podcast, put that in the show notes. We're not going to have show notes here. But yeah, that would be good to, yeah, that would be, I'd like it. Yeah. I'm watching the dashboard that's summarizing people's options. You're doing great. Keep it going. Yeah. I got an observation. Go Bills. The fly stole the show. Someone put go bills up there? From the squirrel. I thought the squirrel was going to hit. And it was the, see, things have to be ready for what happens. Do people just put in observations and it shows up there? Yeah. We're going to talk about it later. Yeah. Okay. Yeah. That's actually, that's a real-time. I'd like to see go bills if anyone's listening. Let's go buff. Yeah, that's right. There's two people. Go bills. Yeah. We should get Simon out here. Bills Mafia. Oh, yeah. Simon's amazing. Let's do it. Let's do it. All right. Hey, let's get Simon Aubury out here. Come on, Simon. Good to see you. Simon. Oh, absolutely. Wow. What a room. Isn't it great? Yeah. Simon, I've got a few resources here that we were able to uncover. First of all, shameless, or is it maybe a shameful plug? Oh. We zoom in on that. Simon Aubury, co-author of the book Getting Started with DuckDB, shown here with a printout of the cover. You're getting tight on that. Well, okay. There we are. There we are. I don't know which camera that is. That's important. Simon's an author. I was able in our archives to uncover this picture of early Simon. Where are we there? Okay. Simon, is that an Amiga 500? Yes, that is an Amiga 500. And what else is different between you here and? Look, I think time matches on, but some things never change. And in my mind, I look exactly the same as you. Every fair from fair sometime declines by chance or nature's changing course untrimmed, as the bard would say. But I also was an Amiga 500 guy. Oh, excellent. Excellent. I learned C. I did. The Lattice C compiler. Yes. That's what I'm talking about. Yep. Your former employer. The white book, these common roots, I didn't know this. Anyway, Simon, you are a principal engineer based in Sydney. Sydney, Australia. Yep. I always hear Texas in your accent, but. Fantastic. I hope I can speak slowly enough. We can have a bit of a conversation. We can make it, yes. Yeah. Two people separated by common language. We can always put subtitles up. We need the 800 number for Anna, the subtitles for Simon. It's all good. It's all good. The lessons you learn. We're among friends here. It's all fine. Okay. So you've been working in, I mean, I think doing cutting-edge things with streaming and Kafka for a long time. I have been, I think, personally most interested in your work with data products. Data products, yes. Could you tell them what I mean by that? Just sort of go for a little bit and I'll talk about it. Yeah. So we're probably in a room of engineers. We all live. We breathe data. But sometimes there's sort of friction. And actually, I might use that beautiful photo that you were holding for us. Oh, would you? Actually, do you remember the time that we got these beautiful computers, that Commodore Amiga 500? It came in a big box. It said, "Commodore down the side." I love it. Yeah. And you took it out of the box and it had everything you needed. There was the processing unit. There was a monitor, the mouse, the keyboard. You took it out of the box. It just worked. You didn't need to register. You didn't need to plug it into the internet. It had everything you needed. And when I think about data products, this is the same kind of concept that I'm thinking about. It's got everything you need to make that data useful. It's addressable. You understand what it is. And if there's anything wrong with it, there's a name on the box so you know who to go talk to. There's a name on the box. Name on the box, yes. And so concretely, this data product is its data from an application developer's perspective, which is my default, that I emit. Let's say it's a message I produce to a Kafka topic that's some unit of work, the result of some unit of work. Here you go. It's in its topic. But the schema is defined and I've got my name on it. Yes. Well, that creates some interesting incentives. Yes. Yes. How has that gone for you? Because now there's some pressure on me. Yeah. Yeah. So having a product owner gives you a level of ownership, but also a level of respect around the data. You're going to make some trade-offs around the quality of the data, maybe the ownership of the data. I'm not sure why Anna's giggling on my right here, but I'm a bit worried. Just having that sort of ownership drives sort of a lot of incentives around quality, SLAs, and all of that kind of good stuff. There's a slight distraction over there. I don't know what's going over there, Tim. Just a disruptive audience. Things are. I don't know how that guy got here. I kind of like him. Yeah. So Anna, were you laughing about data products or about. I was laughing because I got a thing for mascots. I like Gritty, and I was just watching over there. Okay. Oh, people are throwing things at him. Watch your head there. Yeah. All right. I think he's being escorted out by security by the looks of it. Yeah, but with data products that are managed, there's a schema. There's an owner. Now, downstream consumers, because this is a story I've been telling for a long time. You produce data into durable logs instead of ephemeral RPC calls. Now you have a record of what's happened in the business, and things can grow. It's like soil. But it doesn't really grow if you don't have the system that you've built. So you get these great outcomes. You get this evolvable architecture. You get great business outcomes. Business leaders are going to be happy about it. But how do you get me? I'm an application developer. I'm just worried about a crazy business stakeholder who wants a crazy feature that I want to get done. Where are the levers to build incentives for me to want to do that? Oh, yeah. So this is a fantastic question. So you actually want to make sure everyone is aligned on essentially the value of the data product. So technologists might think about this very much from the technical sense. Do you have schemas? Do you have SLAs? Do you have a Slack channel that's going to be responsible for managing it? But you also need to apply that business lens and the organizational lens. So the business lens is a single source of truth. And that's very, very powerful. You have one definition of customer or product or whatever. And having that ownership construct really sort of drives out making sure that you've got a level of consistency on the business side. And maybe the last thing is making sure you've got some organizational ownership around it as well. Okay. And that's it. Make sure that you've got sort of some cross-functional teams and maybe a bit of a feedback loop because not every product's perfect. So you also want to make sure that from an organizational perspective, you've got a kind of concept that, yeah, if it's not right, how are we collectively going to make it better? And hopefully that drives everyone in the same direction, including technical folk, to own, to manage, and make things sort of living artifacts going forward. Yeah. I love it. And I love what you've, again, having talked to you about the stuff you've built out. And in conversation with me, I've seen you really draw out the cultural change that has to take place and the incentives that have to change at the hands-on keyboard level of people who have to begin caring about the craft of their data. And with that kind of system in place, regardless of what you're doing, whether it's we're re-architecting because we want two seconds of latency instead of two hours, or we're rolling out AI and we need access to all of this state of the business all at once, the data product path gets you there. And you're just, again, I think one of the thinkers in the space that's done some of the most interesting work that I've seen. So yeah. I appreciate what you do. Thank you. And you do cool things with Raspberry Pis, and you're an Amiga guy. This is very nice. Sound of applause. Yeah. Simon Aubury. Thank you, Tim. Should we get Alex out here, you think? I think so. Okay. Awesome idea. Alex Merced, please join us. Good to see you. Sir. Alex, you are, I understand, head of DevRel at Dremio. That is correct. Yeah. If I asked you to describe what you do, how do you? It's a tough one when you're in DevRel. You're at a party with non-technical people, and they say, "What do you do?" And you wish you could just say, "Sales," and they'd know, and they'd leave the conversation, right? But I mean, the way I think about it is I wake up every day. I talk about data lakehouses. I talk about Apache Iceberg. I talk about Apache Arrow, Apache Polaris. Then I go to sleep, and then I generally do it again with also some blog writing, educating, just basically being out there and getting people excited about these technologies and excited about this architecture. And yeah, also just being excited myself about it. It's funny. You're walking through your day, and that felt like you were up till about 11:00 A.M., and then you went back to sleep. But you spend the whole day doing those things, and you sleep. Oh, pretty much. I even have post-scheduled even when I'm sleeping. Nice. Got perfect. Okay. I get some pictures also. If we could zoom in here. Co-author of Apache Iceberg: The Definitive Guide. Are we picking that up? There we are. Yes, with a couple of co-authors. Yes. Great book. I have read this book. Love it, and I have a picture of you too. This is when you were. Yeah, it was me, I think. Yeah, sophomore year of college over there in Bowling Green State University, going right before we were going to like an '80s party. So I was going glammed up for. Okay. Good times. Okay. I don't want to. Don't take this wrong, but Alex, you were really cool. Once upon a time. Yeah, right? That's the best, and okay. '80s. Okay. I was there, Gandalf, 3,000 years ago. Yes. Me too. At a younger time. Right. I mean, you talk about Iceberg. Tell us about Iceberg. Tell us about the lakehouse. Why has this thing emerged? What's going on in the world? What do we need to know as data streaming engineers? Well, you have data lakehouse, and you got Apache Iceberg. And the data lakehouse basically did the separation, or basically says, "Hey, how about we take your data lake and treat it more like a data warehouse?" Not just being able to run analytics on your data lake storage, but being able to actually not look at it as a bunch of files, but look at it as tables. And this allowed it where you didn't have to have as much duplication of data, reducing costs because you have one set of data that can work with multiple tools. You have also just the flexibility and the ability to deliver your data quicker to your AI use cases, your BI use cases, because you're not having to spend as much time moving the data from multiple systems. Now, Apache Iceberg accelerated that even further because it created multiple standards that allow that interoperability to flourish. One being the actual standard, the specification of how you write metadata for a table, but two, also the catalog standard, the Iceberg REST standard that allows multiple catalogs to have massive interoperability across the entire ecosystem of tools that want to work with your data, allowing us to basically work with our data faster and deliver to those use cases faster at a lower cost, which is essentially maximizing the value of those workflows for the business. All right. The interoperability, and it's a political statement, but if we could say Iceberg is sort of the emerging de facto open-source standard, to me, that points to if there were open table format wars, and those have been resolved, it seems like that is a setup for the catalog wars. Yes. And who's going to control the schema? Seems like a thing. What are your thoughts on that? To me, the catalog is definitely the center of where I spend my nights thinking about. Because at the end of the day, once you have the standard for how the data is represented as a table, you need to have a way to track those tables, to govern those tables. But at the end of the day, the reason you're going to eventually want a standard catalog is because the world is not just Iceberg tables. So whichever catalog becomes a standard is also going to determine what are the APIs for your unstructured data, for other types of data formats. So it becomes interesting seeing the landscape of catalogs and the different ways they're approaching their APIs for other things and kind of saying, "Okay, what is that?" I'm always one taking a look at history and seeing in the last couple sort of forays of competing standards, typically sort of like the open-source community standard kind of wins. But that's one approach where one catalog wins. But thinking about it, I thought of sort of like two other paths where sort of this unification of catalogs can happen. One, sort of in the Iceberg REST spec, for example, there's what's called the scan planning endpoint that basically abstracts a lot of the working with the metadata to the catalog. Now, if you also had something equivalent on the right side, theoretically, then you could create the catalog could be the center of, or any catalog could be the center of that sort of massive universal interoperability, and then the third way is actually saying that Iceberg metadata could work with all sorts of other different file formats, and then that's something that's actually already underway with the file. There's a file API proposal to make it easier for Iceberg to work with different file formats, which could open up the door to all sorts of different possibilities. So I'm pretty excited seeing all these different efforts that could really break open to, "Hey, a catalog can govern my entire data world," making it all interoperable, making it all accessible, which then just means that all the technologies that I use on top of that can deliver what I need to deliver faster and in a governed way. All right. What does this do with the rest of the ecosystem if there's this sort of normalized way? Anybody can get at the data that's in the data lake in a normalized way, put differently. Compute and storage are kind of unbundled. What do you think is going to happen? How are vendors going to rebundle in ways that create value? I mean, I do think systems go through the cycle of bundling and rebundling because there are benefits to integrating things, and then there's benefits to not integrating things. And the reason, for example, we had databases and data warehouses that have integrated the way we store data, the way you track tables, the way you govern the tables, and the way you actually process queries. And they were bundled systems. But now that we're seeing this, the data lakehouse was this sort of trend of unbundling all these components and making them separate things. And the beauty of it is now we have separate people optimizing each layer. How do we optimize storage? How do we optimize the metadata for the table? How do we optimize the tracking and governing the catalog? How do we optimize the processing, the query engines? But after a while, once we kind of build in all these new innovations and we discover sort of what is the better way to do each piece, then the next step is sort of like rebundling it and sort of like, "Okay, how do we know that we've discovered the better new way, how do we repackage it again to get the benefits of integration?" And you're starting to see that. You're starting to see different layers of this starting to be put together. Generally, when I think of an Iceberg lakehouse, I think five things: storage, ingest, catalog, basically a semantic layer, and the table format. And you're starting to see different vendors take maybe one or two of these things and bundle them together again. That's just going to kind of continue to expand until someone sort of discovers the, or you start seeing massive sort of bundles again. Then we'll go through that cycle again. Eventually, there'll be a new paradigm. We'll unbundle it again, and it'll be like data lakehouse 2.0, and then we'll just kind of keep cycling through that to allow that innovation to accelerate. So it's pretty exciting. Yeah. I like it. I like it too. Next, we want to hear from some folks who are building things out in the world, out in the actual world. A few companies we want to recognize with our annual data streaming awards. So the four of us get to step aside, clear the stage for the data streaming awards. Ladies and gentlemen, thank you. I like it. Please welcome Senior Developer Advocate Sandon Jacobs and Staff Developer Advocate Olena Kutsenko. Good morning, Nolans. Nobody said it like that yet. Hello. Yeah. That's the way you really say it is, Nolans, by the way. It took me some time to figure out how exactly to pronounce the name of the city, but I'm proud that I can do it. New Orleans. I want to hear you try it though. New Orleans, New Orleans. We'll try again later. We're here to present this year's Data Streaming Awards, and people ask me all the time, "Sandon, what are Data Streaming Awards?" well, here we go. It's an industry-wide awards program that recognizes organizations in data streaming to use data streaming to transform their business and deliver outstanding value to their customers and to the community. It's designed to bring this community together to highlight those innovators and their use cases from around the world. So let's get right into it, Olena. So our first data streaming award is for data strategy and contribution, and it goes to Slack. I'll do this. I will shake your hand, but my hands are busy. It's for you. Capturing data from a one petabyte database. Taking 600,000 writes every second, it's no small task. This is extreme scale in action. And to manage that, Slack built a real-time streaming pipeline, which cut data latency from 48 hours into less than 10 minutes, saving the company millions and unlocking faster decisions. And they didn't just adopt open source. They improved it. Their contribution to Debezium is now helping developers across the whole world to stream data more reliably at scale. Congratulations to Slack for showing that even the busiest data can keep up with conversation. Congrats to Slack. All right. Our next award is for data modernization and contribution. The winner is Michelin. This guy's a hugger, man. We got to bring it in. Bring it in. There you go. There you go. Now, a lot of us are here today because of Michelin products, tires, and of course, that iconic Michelin guide. Who says you can't teach an old dog new tricks? Because this 130-year-old company back in 2018 embraced event-driven architecture to modernize their software, to empower their teams, and to strengthen their technical culture. They give back to the community through open-source contributions like Kstreamplify, excuse me, Ns4Kafka, and several KIPs, especially around the Kafka Streams area. The result is a higher quality of service, a resilient decoupled architecture across over 60 factories worldwide. They've lowered their run rate costs by 3x, and they de liver 5x faster delivery, all with zero downtime releases. Congratulations to Michelin. Our next award is for innovation and contribution, and it goes to Uber Technologies. Are we hugging? Thank you. It's you. Uber runs one of the largest real-time data platforms in the world, powering instant decisions for rides, deliveries, and more. They built a complete streaming ecosystem, real-time quality checks, rock-solid financial data handling, and fast analytics. And Uber is also a major open-source leader, helping the whole industry to move forward. Congratulations to Uber. Our next award is for enterprise scale and contribution, and the winner is LinkedIn. We all know LinkedIn, the world's largest professional network connecting professionals through jobs, learning, and community. Their managed stream processing platform on Flink, running on Kubernetes, powers thousands of mission-critical pipelines, including notifications, search, jobs, ads, AI, and abuse detection. As scale and reliability demands surge, they built this self-service production-grade Flink platform that abstracts complexity, and large-scale and large-state jobs are more reliable, and they cut costs. The results? Well, they can iterate faster on their features, and they achieved major infrastructure savings of up to 80% on some targeted workloads. So congratulations to LinkedIn. And our last award, but not the least, is for mission-critical application, and the winner is Raft. Raft seemed to create a CBC2, a cloud-based command and control system that brings together over 800 real-time data feeds into one clear view. CBC2 supports quick decisions to help protect civilian airspace and critical infrastructure. Kafka Streams and Apache Kafka keep the data flowing when every second matters. Raft successfully replaced a 16-year-old system with a modern streaming architecture to support the most demanding mission-critical operations. Congratulations to the team for delivering innovation where reliability matters the most. If you are part of the team or know the team doing incredible things with data streaming technology, you can already submit nominations for 2026. All right. One more time, everybody, and Nolans, put your hands together for our data streaming award winners. All right. So there's been some interesting things happening through the course of this keynote on this demo. You've been seeing some of the stuff? I've seen it. My favorite's been about Scala. I kind of seeded that one. But anyway, no better way to finish things up here than by going to the man who actually built this himself, and I'm actually dressed like him for Halloween. I think he inspired you. Yeah, yeah. His spirit is on the stage right now. But anyway, I would like to introduce you to the man behind our demo here this afternoon. He is the self-proclaimed Chief Vibe Coding Aficionado at Confluent. Please welcome Viktor Gamov. Welcome to Day Two Developer Keynote, the moment that you were waiting for, right? Because how many developers? I didn't see hands. So how many developers do we have in the room? All right. At least a few developers. All right. There's a very old movie that I really like, and there's a cool phrase saying that you need to have a maniac to catch one. Anyone get the reference? You can write this in an application. So you need AI to write AI. And this is exactly how I built this demo. I used modern tooling that is available for your disposal, modern data infrastructure that's available for your disposal to write applications in 2025. Who's with me? Who's ready? Who's ready to see some code? Make some noise. Come on. All right. So hopefully my system, as always, things are not working as we discussed. Looks like the battery died. That's totally cool. So while the computer is loading, this was totally unexpected. As you expected, this is a Kafka event, and the Kafka events need to go somewhere. So essentially, the application that you were using all day, or like this morning, or at least like a couple of hours, this application is Spring Boot applications, pretty cool microservices that take your messages and write data into Kafka. There's nothing fancy about this, and we know that microservices is something that is much easier for us to reason about. Can we have a dashboard for a second on the screen? So the dashboard that you were using today, or at least you were seeing today, it's another microservice that is written also in Spring Boot. And let me see. Yep. Perfect. We have my computer back. So this is also a microservice written in Spring Boot that, as you expected, also will be reading Kafka topic in real time, and we're doing some of the real-time data checks. Who is the Swift Zebra? You know who you are because you have this name on the screen. Who is the smarty pants who probably used data generator to generate some load? Maybe it's a Scala people write the data generators using Scala. No? Anyone? No one wants to reveal themselves. So all the messages, as you know, they land in the Kafka topic. In the Confluent Cloud, we do have an infinite retention, meaning that all your messages that you submitted, they will be here like, I mean, forever or whatnot, and the cool thing about this, now my cursor is not working. Okay. What is happening? That's what we not expected. That's a live demo. Let's see if the thingy will work. No, so we're going to use a different laptop. That's why we have live demos, and we have backups. Always come with backup, always come prepared for this type of situations. There we go. There we go. Now, first of all, let me show you this funny little IDE that I use called Kiro. Anyone heard about Cursor? Anyone heard about Windsurf? So Kiro, this is something that you should use in 2025 because everyone is tired about vibe coding. I don't know if you kind of like to try this yourself, but vibe coding is not the way how you write the scalable applications. The vibe coding is something that you kind of like to play around. Real developers would use what we called spec-driven development, and I took my developer hat and actually wrote a lot of specs instead of a lot of code. So I wrote the specs that let me see if I can switch theme. Oh yeah, I can, and we go with dark. That's what I like more. I wrote a lot of specs, and those specs will define some of the features. You saw this probably multiple times when you need to describe your intent. What are you doing with this and stuff like that, so this is what I'm doing. I'm describing my intent. I want this agent to go and write this application for me. I'm describing this in a very well-detailed fashion. And after that, this tool will take this and break this recommendation and break it down into a design document where we have actual implementations, and we will have a task list where we can follow along what kind of implementation things happened on the screen. All right. So the cool thing about this tool is that it is Visual Studio Code-based, meaning that we can use all these tools together with Confluent Extension for VS Code. And I can show you how my Kafka infrastructure would look like. And so we're going with the current demo. We're going to do the Maestro because you were all Ensemble, and the Kafka is my Maestro where all the messages will go. And this is your user messages. So as I said, these messages might include something smart. Let's see if anyone tried to do something smart. All right. Who is this smarty pants? Who is doing this? Drops. Okay. I was expecting something more smarter, like Drop Table. Let's see if we have the SQL injection specialists. Oh, Drop Table, all tables. Nice try, guys. So again, as I said, infinite retention, meaning that all these messages that you wrote can go directly to your work HR if you wrote something. But if you write something good, it's fine, right? It's fine. Now, so since we're doing this in Confluent Cloud, and this is something that I was personally excited about, yesterday's announcement is the Confluent Intelligence and how you can write the agentic applications these days. So normally what I would do, I would use Confluent MCP Server and some of the existing, obviously Java-based, not the Scala-based agentic framework, and I'll connect my things together. So I will ask this tool to go to Kafka topic, read the messages, and push it to there. Now, we don't need to do this because inside the Confluent, we already have, say, I'll go with this Flink Compute Pool, which is in my, and it runs my SQL here that will be doing our processing. So some of the messages that you already sent, they might include some additional information. We also want to have a heartbeat because we want to advance some of the Flink watermarks. And now we also run this against what I'm looking at is something like model. I don't have it here, unfortunately, for some reasons. Let me see if I have an actual code here. I don't see it here, which is fine probably because it is not the right laptop, right? So I'm wrapping up here. You can find me and talk about this demo. Yeah, because I'm looking at the wrong place. That happens on the stage. So we can connect to any model. So in this case, we're connecting to Bedrock. We're actually submitting two things because I knew, and we knew, and Robin, who helped to build this demo, we knew that you will try something funky. So first query, or first model, will actually generate summary. And after that, we will use another AI to clean up this. So first one, we get all these messages that you sent. After that, we got to sanitize it so it will be presentable on the stage. So how cool is that? No? So the tooling, again, you need to embrace AI. You don't need to be afraid about AI. You still can use your skills. And the most important skill that you have these days is the way how you can express your intent. And this is something very important. And for prompt engineering, context engineering, for writing specifications, as long as you clearly can describe your intent and write this in the words, you will be successful with this small adventure, what we call Agentic AI. With this, my name is Viktor Gamov. And as always, live long and prosper, and I'll see you in streams. That was a little more meta than I was expecting. Things happen, and you have to be ready. In the case of Viktor, he's ready with a spare laptop because, hey, you know. A program like this involves dozens of people turning hundreds or thousands of cranks and keeping track of so many different tasks. It's a tremendously complex operation to put on. The program itself, the sessions that you spent all day yesterday in, you're going to spend the rest of today in, that takes a lot of work, and that's these people. Round of applause for the Current 2025 Program Committee led by, you saw her at the very beginning, Adi Polak. Adi is the Program Committee chair. She takes the votes of this august body and munches them and crushes them and turns them into the program that you have. Trust me, it's a lot of work, I know, because I've done it before, and she does a tremendous job. So it's great to have that and just everybody making this event happen. I mentioned being a data streaming engineer. I mentioned certification. You can, for free here at Confluent, become a current, become a Certified Data Streaming Engineer. You can do it online, actually. Scan that QR code, do it from your phone if you don't like being around people. If you don't like being around people, that's a little awkward because you're around a lot of them right now. But you can do it online. It is free. You can go to the Certification Lounge, which is in Room 297. There's coffee. Yesterday, there was king's cake there. I suspect there is today. So if you want some sugary snacks and caffeine and free certification, honestly, do it. I don't really know that I'm going to take no for an answer. So check it out. You need to become you get swag too. Thank you very much. For example, here, I don't think I'm supposed to go over here right now, kind of breaking the rules, but one example would be, oh dear, this isn't set to the biggest setting. And if it's going to fit on my head, it's going to have to be, yeah, you get a hat. So we'll stop there. There's more than just that, but do it. Get certified. And you've seen this unfold on stage today. I've seen it unfold because I was kind of watching the dashboard, and y'all were actually responding to things that I wasn't even expecting you to respond to. I orchestrated that some weird things would be going on in the room, and then a giant fly starts buzzing around, but he makes it up there. And the Buffalo Bills somehow make it up into the summary. But see, there's the one. There's two. Now, Anna had to step out. She's in a meeting. There's the other one. It's a terrible, terrible cliché to say the only constant is change, right? We all know that. But a static environment is not what we have. And if that's what you wanted in your life, this wouldn't have been the career path to take. We build systems. Again, the deep structure of the systems we build is based on the notion that something is going to happen next. There's always another event coming. We always have to be ready for that. And we're living out changes in application architecture, changes in data architecture that we've talked about today, the tremendous impact of AI on the way we work and the systems that we build. How do you prepare for what's next? Fine. I'm saying be ready. How do you do that? Well, you be here. You're already doing it. As the old TV commercial said, you're soaking in it. You learn. You go to sessions. You talk. You meet people. You tell your story. You listen to theirs. You take a selfie with them and post it, your favorite social media, hashtag streaming selfie. You talk publicly about what you've been doing here. That's how you be ready. You learn and you be a part of a community like this, which you, by choosing to be here, you're already doing it. So go to your first session. Have a fantastic day. Track me down. I'll be around today. I would love to meet you if I haven't already. Thank you.
Loading workspace