Please welcome to the stage GitLab's Chief Executive Officer, Bill Staples. Welcome to GitLab Transcend. We are broadcasting live from a packed house here in London to more than 15,000 registered people around the world. No matter where you're tuning in from, thank you for spending time with us today. He doesn't know I'm going to do this, and he doesn't crave the spotlight very often, but I'd be remiss if I didn't recognize a very special guest in the audience today, because none of us would be here without him. He is our co-founder, our exec chair, and he is healthy and cancer-free. Please join me in welcoming Sid Sijbrandij. The company Sid has built is truly amazing. It's a platform in every sense of the word. We just surpassed $1 billion in annual revenue last quarter, serving more than 50 million users and hundreds of thousands of organizations around the world. More than 50% of the Fortune 100 trust GitLab to build their software to serve their customers. These are iconic companies in every industry vertical that I know every one of us, in our consumer lives and in our professional lives, do business with. We use the software that they use, GitLab, to build every single day. It is such a privilege to be part of this community. What's really remarkable, though, is despite all of that success, inside GitLab, coming to work every day with 2,000 teammates, is the passion we have for solving your problems, for innovating new ways of building software. In this era, reinventing how software is built. It's an incredible community, and the community is growing. Just last quarter, we added 30% more new paying customers than the same time last year. Thanks to you, we've seen 100% double the number of code contributions into GitLab from our customers and our community from just one year ago. Developers are choosing GitLab to build software more than they ever have. Over last year, 250% more user namespaces created. Isn't that amazing? What are you all doing with GitLab? Let me tell you. Platform usage is surging. CI/CD pipelines have grown 40% one year, 50% increase in code pushes, 60% increase in secure repos. Some of your code bases over one year ago have grown 500%. What is going on that's driving more value from GitLab than ever before? Can you guess? In a word, AI. AI is transforming software engineering. We launched our Duo Agent Platform at the start of this year, and in its first quarter, we took more bookings in our first quarter with that platform than any prior quarter with Duo Pro and Duo Enterprise combined. Clearly, agentic engineering is in demand. In fact, since the beta went general availability, we've seen 1,000% growth in weekly active users on that platform. Agentic engineering is here. We see about it in the news, we read about it in the tech press, we talk about it in the hallways and the virtual hallways, around the water coolers, everywhere. It's amazing, isn't it? It's been fueled by our partners from Anthropic, Google, who will be with us on stage today, as well as others who've made coding incredibly fast. In fact, non-technical users who've never written a line of code can now generate working code 10 times faster than a professional developer did one year ago. Professional developers are harnessing these tools to take ideas all the way to production in minutes. That is what's igniting software engineering on fire. Speed can also come with a downside. Just like a race car, it doesn't matter how fast you can go if you can't stay in control. It doesn't matter how fast you can go if you can't trust the steering wheel to get you where you want to be or trust the brakes when you need to slow down around the curves. Speed without control is chaos. We also see this everywhere we look today. I'm guessing your social media feeds look a lot like mine. With the explosion of code, we see the explosion of bugs and quality issues. We see code review queues get longer and longer. We see more security issues than ever. We see infrastructure and major services that can't maintain reliability. Yes, we see costs explode as well. Speed without control is chaos. We've been anticipating this problem, and we've been thinking about it for a while, and we think we know the answer. Together with agentic coding, GitLab will bring agentic infrastructure to help you harness speed with control. That's our theme today, helping you maintain speed with control. How do we do that? A few weeks ago, I shared a letter to all of our customers and investors titled Act Two. In that letter, I shared five architectural bets that we have been making and that you're going to see today, which are going to change the world, change the way software is built. Let me give a brief overview before we dive in and get a chance to see them. Number one, machine-scale infrastructure. You see, the DevSecOps infrastructure that was built today, over the last decade, was built for human-scale engineering. Machines, agents, they work 24/7. They don't take coffee breaks, they don't get sick. In fact, they work in parallel to software engineers, sometimes multiple agents per engineer. The infrastructure on the other side has to scale at machine scale. We're building that. Agents are also great at, and getting better at, performing tasks. You don't need more tasks, you need quality software that meets all of your engineering standards and regulatory standards, going to your consumers and driving your business. Orchestration is what takes agentic tasks, connects them, passes context to them, and gets you working software certified for your customers on the other end. Context is a superpower of GitLab. We've always been good at capturing all of the software lifecycle data and providing that to your human engineers. We're now doing the same in an all-new way for agents. You're going to see GitLab Orbit today, because the difference between hallucinations and reality, the difference between false confidence and real confidence, is really good context. With GitLab Ultimate, we've been your trusted partner to make sure that you meet the quality standards, the security standards, and the compliance and regulatory standards of your business. We're extending that now to agents as well to ensure that you can govern and audit every single action from every entity building software in your team, whether that's a human or an agent. You're going to see that today as well. Finally, we're delivering this in one platform for all the ways that software engineers will work. I'm guessing your engineering teams look a lot like our engineering team. We have teams and projects that work with human-led software engineering as they have for the past decade. We have engineers who are using agents like Duo Agent Platform to do agentic assist, and they're going two to four times faster than they were just one year ago. And we have bleeding-edge teams that are using advanced agentic techniques to do autonomous engineering. What's incredible is they're building software at roughly 20 times the rate that those same engineers were one year ago. We're learning from all of those teams. We're capturing their problems, we're bringing solutions, and we're sharing that with all of you. In fact, you'll meet some of the engineers in all three modes today. What is incredible about GitLab, unlike the cloud era where you were forced to decide a technology stack in the cloud or on-prem, where you were building different ways of working in the public cloud versus your own private data centers, is with GitLab, you can stay in one place for all modes of engineering, have one set of engineering standards, have one security boundary to manage and audit, and to serve your customers no matter how your teams want to work. We deliver it in a cloud-neutral and model-neutral way. That's the promise of GitLab. All right. Let's go ahead and dive in now. It is my pleasure to introduce our Chief Product and Marketing Officer, Manav Khurana. Okay, Bill. Thank you. You're welcome. Hey, everybody. Welcome to Transcend. We're going to do a few demos so you can see how you get speed with control. To anchor these demos, come with me on a four-part exploration of the new GitLab, which is your agentic infrastructure. You see, whether it's your team or your team's agents, when they are building and shipping software, they need a motor system, as in the arms and legs or the execution layer, to build and ship code fast. That human and agentic brain also needs a nervous system, as in the context to make better and faster decisions. That human and agentic brain also needs an immune system to ship software safely. Finally, that human and agentic brain also needs an orchestration system so that they can coordinate all the tasks that need to happen across the software lifecycle. Let's start with the motor system. You all use GitLab today because you get all the tools you need in one platform, whether that's planning, source code management, continuous integration, artifact management, continuous deployment. All the tools you need are stitched together in one platform so your teams don't have the friction of putting everything together and can do their job a lot faster. Let's dive into source code management. It's been a hot topic recently. Literally hot. Literally burning hot because the Git platforms, in fact, the most popular Git platforms in the world, are buckling under the load of not just your teams cloning, branching, and merging code, but also dozens, in some cases, hundreds of agents working simultaneously and putting a lot of pressure on those systems. You've all seen the same headlines I have seen. Today, I am beyond excited to introduce the next generation of source code management. Internally and lovingly, we call this project Project Switch. Well, because switches are better than hubs, if you're a networking geek like me. Really, it is the same Git protocol for backward compatibility, but a completely redesigned backend, new interfaces for agents to work blazingly fast and do so at scale without any disruption. To show you how this works, please welcome Nick from Anthropic and Kranti from GitLab. Hey, Nick. Hello. Hey, Kranti. Nick, I'll start with you. Thank you for being a design partner on this initiative. You have the unenviable job of managing the coding infrastructure for perhaps one of the most demanding software engineering teams in the world. Tell us what you're seeing. Yeah. Thank you for having me. At Anthropic, we're seeing development accelerate to the point that a lot of tools that people take for granted, like source control, really can't keep up with the load. The same thing you were describing. We put out a blog post about a week ago that showed that we had an eight times increase in the amount of developer output since a year ago, and I really don't see why this is slowing down. What are we running into? It's a good question. We're running into two major categories of problems. The first are sort of traditional big repo problems, so things like getting your checkouts into your CI jobs, or even just supporting a large commit rate. We see that a lot of people have these problems, but people don't necessarily talk about them. It's kind of a shame that everyone has to reinvent the wheel here. We think this should just work out of the box. The second category of problem that we're seeing is kind of unique to agents. We want to run a lot of checkouts, a lot of developers basically, that need the full repo context, so not just CI jobs, not a single commit, but the history, the blame, the logs. This is an even bigger problem than just checking out code for CI. Git doesn't really handle this very well, dealing with partial checkouts of the repo, dealing with slices of context, Git does not support very well. We imagine that there's new ways of agents interacting with the repo, maybe via rich source control APIs, to solve this problem at a bigger scale. Yeah. This is not just a problem that is an Anthropic problem. All of you, as you scale your coding efforts, your development efforts in the agentic era, these are problems that either you're already running into or will be running into very soon. Yeah. Kranti, you've been tackling this problem now. I gather you have something to show us. Absolutely. Hello, everybody. The challenges that Manav spoke about are precisely the ones that we have been set out to solve in the new architecture. Let me demonstrate that with a harness that lets us compare our current generation system with our next generation system side by side. On your left, what you'd see is a cluster that is running our current generation system, which is Community Edition 18.0. On your right side, you'd see a cluster that is running the next generation, and both of them are provisioned with the same amount of memory and CPU resources. For the first scenario, let me touch upon one of the common problems that we have right now, which is doing clones at scale. As you see here, the next generation system is going to get to the act of doing the clone pretty fast, while the current generation system feels a little sluggish. When a clone happens on the server side, the server has to compute a packfile from all the files in your repo and all of its history. If your repo has a lot of files and a lot of history, it's going to take seconds, in some cases even minutes, to get to that packfile. The next generation system is efficient at doing that because it's going to convert that into a manifest pointer towards a pre-computed packfile. If one packfile does not exist prior, and your server is getting 1,000 requests at the time, instead of doing this per request, it's going to coalesce the request and compute the packfile once, and let the clients stream the output from the object store directly, and the clients can scale because the object store behind the scenes is very, very scalable. As you can see on the right side, the next generation system is already done. It has finished 100 clone operations within a second on a modest-sized repo of our own, Gitaly. This is going to take a little while on the current generation, so let me show you a run from the prior run to show the statistics. All right. As you can see in this case, it's 42 times faster on raw wall clock time. That means your agents and your clients can wait 42 times lesser. Not only that, it consumes way less CPU and memory on your server side. That means if you were to run this on your premises, it's going to be very cost-effective in addition to being very, very fast. Amazing. Now let's look at a write scenario. Let's see how the writes perform at scale. Again, what I kicked off here is a comparison of performing 100 write operations on both the clusters. Behind the scenes, it has to create a fork, read a file, make some changes, commit the change, and for good measure, it'll go and even verify the commit is successful and stuff like that. Now, as you can see here again, it is pretty stupendously fast on the right side. Because creation of a workspace in the next generation is super quick. It does this with a clever use of manifests and pointers on a shared pool of objects across all the forks in your repository. Effectively managing that when people are creating forks on smaller repos, massive repos, it doesn't really scale with the size of the repo anymore. It just lets you get to the fork and start working on it right away. Now we have looked at both reads at scale and writes at scale. Let me show you an agentic use case as well. Before that, let me show how the writes would perform on wall clock time. The current generation is going to take a little while, let me show a prior run for you to get a quick understanding. In this particular case, it finished both 100 tasks. It's 17 times faster. That means again, your agents and your clients can get to doing the actual work much faster and not be slowed down by the underlying source control system. For the third scenario, I'm going to let loose Claude on performing a task. The task is this. The task is to go and look at some undocumented code files in Gitaly repo, then go and create the documentation and check it in. As you can see here, the next generation system is chugging along super fast. It's just going after, reading the files, understanding it, incorporating what has to be documented, and it is also going and checking it in. The current generation system looks a little sluggish, but it's going to get there. Actually, the next generation system already finished it within a wall clock time of 30 seconds. Mind you, this 30 seconds is being measured from the client side. That means it has the agent inference time. getting the data all over the wire, all of it. On the server side, the stats are even more fabulous. The current generation is going to take a little while. Let me show you a prior run just to kind of give you a taste of how it looks like. On the wall clock time, you are getting a massive benefit as it is. It is 22 times faster. It is moving much lesser data on the network because it doesn't need to. I would like to bring your attention to a couple more interesting stats here. Look at how few tokens it has used. On the right side, our next generation system ended up using just 500,000 tokens to perform this task. With our current generation, it ended up using 1.4 million tokens to get to the same outcome. That is almost three times cheaper, and it is going to translate to lesser cost for your agents to do your thing. Behind the scenes, what's happening is the full architectural advantage is coming to fruition. The new access patterns that we created are going to enable the agents to really interact with the code base in more effective ways to get to the act of doing the task much faster. How do we do all of this? You want to see the slide? Yes, I would love to kind of show you the architecture slide. All right. With the next generation architecture, we got in three important advancements in the architecture. Number 1. We are letting the compute and storage be separated and allow them to scale horizontally on their own. Number 2, we put a layer of intelligence in between that can bring together the benefits of this distributed compute and storage together, and does a lot of hard work behind the scenes, like routing the request to the right place, caching what is important, partitioning your objects as your repo size grows, creating PAC files, updating bitmaps. A lot of heavy lifting is done by that intelligence layer. The third, rather more important advancement that we are bringing in, is we are allowing the clients and agents to interact with the source code system using newer access patterns. Nick, you've obviously seen this work in progress over the last several weeks and months. What's your take? Does this scratch the itch? Truthfully speaking, I was very impressed when I first saw the demo for this. It sort of worked better than I even imagined it would. We've been playing with stuff like this at Anthropic as well. I think it's going to be increasingly important to actually scale how these agents work. We're also pretty excited that we now have Fable available in a duo agent platform, we expect this to be even more impactful for these large models. Amazing. Thank you, Nick. Really appreciate the partnership. Thank you. Yeah. Thank you, Kranti. Nice work. Absolutely. Yeah. Very good. What you just saw there is the next generation of source code management that is available today in private beta. It is the same Git protocol that you and your teams are used to, but with a redesigned motor underneath for agents to work a lot faster and for your overall system to scale a lot better. You saw things like less than half the number of tokens. In Kranti's example, it was three times. In our tests, we've seen 50 times faster wall clock time already and over 1,000 times lesser network traffic required. Really incredible. Can't wait for all of you to use the product. All right. Let's move on to the next part of the agentic infrastructure, which is the nervous system. Today, all of you, when you use GitLab, one of the amazing things about the platform is that you have a common data store under everything that you do in GitLab. Whether it's your code, whether it's your pipelines, your merge requests, your tests, your security scans, all of those data points are stitched together for you and your teams in one data platform, so it makes it easy operationally to correlate what's happening across the software lifecycle. Turns out that's also quite useful for agents because they need that context. Here's the thing with context. When you are working on a small project, a bounded project, agents can very quickly get the information they need and deliver a magical experience like we've all seen, where we ask a question and we get a fantastic response in seconds. I'm sure you've also tried to use agents in a large monorepo, where there are tens of thousands of files that are in your code repository. When you use agents in that setup, you'll see that agents are constantly iterating. They're constantly trying to ping the backend to get the right information to do the task that you have asked, and this goes back and forth. Each time it goes back and forth, it takes more time, it takes more tokens, and agents reach a point where they give up. They don't have complete information, yet they give you a response which honestly is more artificial confidence than artificial intelligence, and you are left the bag to fix what the agent told you. Worse, if you're working across multiple repositories in your business, across different teams, and you need not only the code information, but you also need information across the software lifecycle, all the related metadata, that's where agents just flat out fail and cannot succeed in doing the job that you have asked them to do. That's why today I'm excited to share that we are introducing GitLab Orbit. It is a context graph for the entire software lifecycle, where all the context your agents need, whether that's in a monorepo or across repositories with all the lifecycle data, is available with a single query so that your agents work faster, are more accurate, require fewer tokens, and more importantly, you can answer questions that you could never answer before with agents. To show you how all this works, I'm going to invite the Orbit team, Angelo and Meg, to stage to give you a quick demo. Hi, Meg. Thank you. How are you? Hey, Angelo. Thank you. Great to see you. Great to see you. You have something to show us? I think we have a few things to show you. We've been cooking. Manav's right. For a single small local repo, agents shine, but that's not the stack our enterprise customers are working on. With large monorepos or multirepos, agents break down because they're trying to chain together context by calling dozens of tools and running thousands of API requests, but the data quality suffers. We had to reimagine context at scale. Angelo, how about you tell us how we solved this by building Orbit? Thank you, Meg. Just to touch on that, GitLab itself is a classic example of a mega monolith repo and thousands of repositories within our own organization. That's where we had a pretty crazy question. What if we took all of GitLab's data and turned it into a graph that agents could query directly? We did that by building a highly distributed system in just three months that essentially is an indexing engine, and as I show you here in the schema, what we can do is. Actually, just the other week, we've been able to index 160,000 repositories into code graphs, as you can see on the right here, we take those graphs, and we marry that data to the rest of the software development life cycle. That unlocks a variety of different questions, which we'll jump into. One thing to touch on here is we built this for scale. All of this can be indexed, namely 500 million nodes and 2 billion edges- Amazing in just 15 minutes. To touch on that a little bit more and how this improves agentic outcomes, I'd first like to ask everybody, has anyone ever gone through the painful experience of refactoring a code base? I don't know, maybe show of hands. Me too. It's not fun. With that, let's pull up a little bit of an example of what Orbit can do for you. I'm going to kick off this run here for what we call the Orbit benchmark. As you can see, we'll go into a little bit about what Orbit does, but we've got this Orbit benchmark here, and we've asked a very relatable question to GitLab itself. At GitLab, we're actually currently exploring decoupling authentication authorization within GitLab itself. As anyone knows, if you have a mega monolith, decoupling such a critical service is a huge pain, and if you're doing that manually, that'll take months, and if you're doing that with agents, there's high risk involved, right? In this repo, how many files are there? There's around 50,000 files, it's millions of lines of Ruby code. If we're going to make any change, we want to know what's going to be affected and how. If we go back to the benchmark here that's already running, we've asked a prompt saying, "Can you get the complete authorization class hierarchy and all of its front-end consumers within the GitLab monolith?" Remember, this is 50,000 files. Let's jump into a little bit what's going on here. Claude Code without Orbit is doing the normal thing that you would expect any agent to do. It's searching through all 50,000 of those files. It's using the classic tools like grep and text search and essentially assembling the entire picture from scratch. On the right, the agent has access to Orbit. What we've done is built a universal code indexing engine that indexes over 11 languages into a unified graph that essentially acts as a pre-built map for your agents to query directly. What that means is that agents are able to write their own queries here, as you can see, and get back the entire authorization tree in just a few hundred milliseconds. That effectively allows you to pretty much answer most of the questions that you would normally ask an agent, but get that answer back a lot quicker with a lot higher quality. As you can see, the agent on the left is still running. Here's the output report with all the authorization classes. To increase accuracy, we've leveraged a lot of different techniques, like SSA and various compiler techniques. Angelo, I think what I saw when you scrolled up, the agent on the right with GitLab Orbit is already done cooking at just one minute and 15 seconds. We're still chugging along on the left. I have the most important question of all, which is tell us about the data quality. That is the most interesting part about this whole experiment that we've been running. As you said, it's still running, we won't bore everybody with the results there. With this previous run, you can see that in just one minute, we completed the results, and we just saw that with the previous run. It took over 11 minutes. This is, Manav, you and I were just talking about this. This is you and I, if we're coding, we can stay in the loop on the right here, and on the left, as a developer, maybe I'll go get some coffee. It could break your flow. Yes. We don't want to lose our flow. With that, the last thing I wanted to touch on that's very interesting about this benchmark is the accuracy. We did something to measure, in a deterministic way, the output of this report. What we did is we took the actual Ruby on Rails runtime. We generated a script to get that same class hierarchy. As you can imagine, both of these Claude Codes don't have access to that runtime. What we're doing is comparing the classes that come from the actual runtime itself with the results from the agent. We've instructed the agent to output those classes. What's very interesting is it's not as complete in the amount of rules that it found from the authorization classes. Claude Code without Orbit actually hallucinated 1,000 more rules as compared to Claude Code with Orbit. That's 1,000 more rules that you have to manually go make sure that Exactly you got right. Yep. That's what we're talking about with agent trust. With that, we can imagine how this is super useful for a developer like myself on your local machine. We didn't want to just stop there. We wanted to empower all of organizations and enterprises with this technology. With that, we built a service. Meg, why don't you show us what else we've been cooking? Yes. As Angelo's alluding to, we didn't just stop at indexing the code base. We indexed your entire GitLab instance, which means you have access to all of your rich GitLab SDLC data. One of the perfect examples is a pipeline analysis. For all the DevOps and platform teams in the room, you might have heard that one in three CI pipelines fail, which can really add up at scale. In the paradigm before Orbit, you could really only analyze your kind of pipeline health at a single project at a time. Now with Orbit, because we have these traversal and aggregation abilities, you can understand thousands of projects and their pipelines at once. Another one of our goals with Orbit is to make Duo Agent Platform even more capable. We're jumping in here to agentic chat, and we're going to ask a pretty heavy-hitting question here. We're going to ask the agents to deep research our most recurrent failing pipelines and their jobs over the last 60 days. In GitLab org, that's 8,000 projects and it's 12 million pipelines. This is a huge question that we're asking. We have two instances. We have GitLab Duo Agent in the old paradigm on the left with access to just the GitLab API, and we have the same GitLab Duo Agent on the right with access to Orbit. I'm going to fast-forward into a completed output and show you what we're looking at here as well. On the left, let's zoom in for just a moment. The agent says to us, "I'll be straight with you, I can't do it." This is a limitation not of GitLab Duo, but of all agents today because they don't have access to Orbit. To achieve an analysis like this, it would have to make tens of thousands of API calls, and the API would just time out. We have Orbit now. Our agent on the right is actually traversing the whole graph and aggregating the job failures. It's understanding the projects, the failed pipelines, and the common failed jobs underneath them. It's giving us a CI compute cost attribution and looking at those common failed pipelines, and then ultimately taking us to the most important question of all, which is, what do I change to resolve this? As we continue to see the agent with Orbit, it gives us the shared CI template hotspots and points us exactly to what I'm going to work on today, which is resolve these. With this resolution, we're probably going to save developers a few headaches and maybe a few dollars on CI compute. Amazing. This is just one of the possibilities that Orbit makes possible. We've been doing a lot of cool stuff, Angelo, I know our engineers have been asking some crazy questions because Orbit can go further than any agent ever could. How about you tell us some of those wild things our engineers have been looking at? Thanks, Meg, that's a great question. We as a team sat down, and we've been just playing around with it and asking some of the most wild engineering questions that you would ask when you can ask anything about your GitLab instance. Just some examples here that we pulled up off the cuff, one of them is find all the dead repos across GitLab org. Find all the critical services within the fulfillment department and who maintains them and who is the expert in those services. Then one of the craziest ones was we actually took the call graph technology that we built and indexed all five years' worth of repositories, then used the graph that's available for the SDLC and did a cross comparison, and we were able to basically get a trend line of various security fixes throughout GitLab. over the past five years. Basically, whatever you can do with your imagination is the limit. Like Manav said, if the nervous system is what lets your body act coherently, then Orbit is that for your entire software organization. Until today, your agents have been definitely flying blind without one, so Agents in the GitLab Orbit just work a lot better as a result. Yep. We're sending them to Orbit, so. Amazing. Amazing. Hey, thank you. Great job, Angelo. Thank you. Thank you. Great job. What you saw there with GitLab Orbit that is now available in public beta for all of you to use in your GitLab instance is your agents will work just a lot faster. Getting a response in a few seconds instead of a few minutes, using up to four and a half times fewer tokens, and most importantly, you will see up to fewer than 45x hallucinations, which is a real important thing. That's why today we are also kicking off a community hackathon where all of you here, as well as throughout the world, can join for the next two weeks and see what you can do with Orbit and all the different agents you use inside and outside GitLab and see what you can build with that. All right. Let's get into the immune system. This is about making sure that you, your teams, and their agents can build and ship code safely. With GitLab Ultimate, you already can be proactive with security and compliance because every tool you need is already built into your workflow, whether that's Security Scanning, Secret Detection, Software Composition Analysis, Vulnerability Management, policy enforcement, making sure you're meeting all your compliance frameworks. All of that's built into the same platform that you use to build and ship code to make sure you're always secure and always compliant. Here's the thing, in the agentic era, the security and compliance exposure is only multiplying. That's because agents, just like they are great at writing code, they're also great at exposing vulnerabilities, and they can do that faster than we can react. The typical cycle goes like this: A new vulnerability is found, there is all this excitement inside a company to decide, "Hey, do we have that vulnerability?" As you go search for that, it's very common that we find that there are coverage gaps in our testing, and we are not testing all of our code repositories to find if that vulnerability exists. When we set up that security testing and cover the coverage gaps, we find that there are a lot more vulnerabilities to address. We have to triage them, we have to fix them, that takes weeks of coordination, expanding the risk window for you. All of us are also using agents across the software lifecycle. The compliance team also wants to know if those agents are acting with the right rules, with the right setup, and everything is in compliance. We invariably discover that we need more controls to make sure that anything agents do from this point onwards will always be compliant and meet our regulatory requirements. That's why we have expanded on top of GitLab Ultimate recently by bringing agents to security to automate a lot of these tasks for you so you don't have to wait weeks, you can address these problems in minutes. Today, we're introducing new governance capabilities for agents so you can always remain compliant. To show you how this all works, please welcome members of our security team, Alan and Michael. Hey, Michael. Hi. How are you? Hey, Alan. Amazing. Great. Amazing. Good to see you. Thank you. Awesome. Let me pick up from the first problem Manav earlier called out. Security can't keep up. Let's walk through a common scenario that we're all familiar with. You wake up on a Monday morning, a new vulnerability dropped in. Our team thinks it's in production. This first question isn't how to fix it's whether we know where it exists. Alan, what can we do about this? Sure. We all know the page, vulnerability report page for the project where you know you have your scans enabled. You can quickly go and solve it, either by using resolve with AI or by using one of your specialized agents like Security Analyst Agent or one of your own you built for your own organization. That's not really a problem I would like to solve today because you see, the biggest gap in security is rather related with the lack of scans running for your project. You don't know if they're running or not. Let's go to Security Inventory and check that. You see, this is the problem I was talking about. We have scans enabled only on two projects where we have this vulnerability found. Let me quickly fix that. I know I would like to enable those scans, but only for most critical projects. Let me do that by selecting Business Impact and then choose Business Critical projects. Now I have a list of projects where I would like to enable those scans. Just by going through a few clicks, you can just enable those scanners one after another, fast Secret Detection and Dependency Scanning. Now, you remember we had those builds. Those builds that you see were wide here, so no scans were running. Now for every single of this project, now scans are running whenever you push changes to those projects. On top of that, let's also ensure that we have enabled our Duo workflows like fast false positive detection, and vulnerability resolution workflows. These are all enabled for all of those projects from this moment. We got one vulnerability, now we have hundreds. Months of coordination, hundreds of vulnerabilities, all replaced by one critical action. Your critical estate are covered. The agents are watching. Wait a minute, like you said, we had one critical vulnerability, and now we have over 200. Yeah. How do we know which one we really care about? Sure. I mentioned agents that are running in the background and understanding if vulnerabilities that were found are false positive or not. You notice there is this icon associated with each vulnerability saying if that vulnerability is a false positive or not. I can use filter to filter out the noise. Let me do that. Just like that, we have reduced the noise. Also, we just heard about Orbit. Security Analyst Agent was also integrated with Orbit to help you understand everything about your vulnerabilities and how about the code is being exposed in other projects as well. That's amazing. We just brought down the vulnerabilities from over 200 down to a little bit over 20. These are still real vulnerabilities, and we still need to fix them. Typically, before the pre-agentic era, you have to book your meetings with your OPSEC team, a lot of coordinations. You need to discuss about what is the best possible path to actually resolve these vulnerabilities. That would take a long time. Alan, I know we just shipped something that could make this even faster. Yeah, I love the challenge. I mentioned those activity icons next to each vulnerability. I mentioned they are telling you if there is a false positive or not. There's also new icons here added. You see, we have information about for each of those vulnerabilities, AI agents already created merge request to fix it. I can go and immediately go to that merge request, talk with the team, let them review it, and merge that change immediately. Isn't that amazing? Instead of months of coordination, instead of having these meetings with the AppSec team to discover these vulnerabilities, to think about what's the best path forward, the meetings you're having with this AppSec team is to decide if this MR is good to go. The fix is right there waiting for the developers even before they open their laptop. Amazing. That would save everybody a lot of time to get that. What about the second problem? Invariably we all have compliance teams in our companies as well, and they want to know if these agents that we are using are doing things the right way, and are we exposing new risk? Are we meeting our compliance regulations? What's the story there? Manav, that's a big problem. Because of the EU AI Act, regulators and auditors are starting to require that most teams and organizations can prove that their AI agent acted within predefined boundaries. When auditors walk into the room, most teams have no answer for them. Alan, do we even know what our agents are doing? Sure. Let me go to the part of the GitLab that we're building, AI governance, where you can have all agent artifacts, all interactions agents did, and all tools that they were calling, and whenever they had approval or not. I'm in a session, and I can view more details about each action that agent did, and I can quickly go to the session details to learn more about what happened. It looks like we have dismissed vulnerability without human approval. Let's see what we can do about this. That's a problem. Yeah. How do we make sure that we can control these agents from now on? Yeah. That is why we are working as well on the second part of the AI governance called tool management. Within tool management, you're able to decide how your agents are interacting with you, and then there's tools. Either you would like them to write, read, and if you'd like to allow them to do it, if you would like to decide to rather they should always ask or always deny, they will not be able to do that action. In this particular case, I was talking about dismissing vulnerability. Let me switch that option from always allow that was previously configured to always ask. Now whenever I would like to dismiss a vulnerability, the agents will still provide me a helpful guidance, but at the end, it will ask me for my approval. Agents can still help, but I just need to approve it. That's awesome. Let's see what the agent just did. It caught the violation, fixed the policy, proved it worked within one platform, GitLab. This is exactly what the EU AI Act are asking for. When auditors want, they want this today. Yeah. That is just a start. Security policy store that we're working on will also include more capabilities, like allowing you to decide about triggers, rules, and actions that should be taken based on the situations that are happening in your code, either related to AI, security, or compliance. We fixed the coverage in seconds, then we had remediation solved in minutes, and compliance already built into GitLab. Awesome. Speed without control is more risk. GitLab gives you both, native security with governed agents. Let's recap a little bit. First, we brought agents to security to improve security coverage, detect false positives, and accelerate resolutions. Part one, done. Part two, we brought governance to agent. We both saw what agent can do and the risks of doing that. I would want you to be able to control what those agents can do in the future. Over to you, Manav. Great. Nice job, team. Very good. Thank you. Great job. Thank you. Thank you. Can't wait for you all to try the new agents for security and the new governance for agents. All right. Let's move on to the last part of the agentic infrastructure, which is the orchestration system. In January this year, we introduced in general availability GitLab Duo Agent Platform that Bill had referenced earlier. GitLab Duo Agent Platform brings agents, specialized agents, and agentic workflows for you and your team at every stage of the software life cycle, so you and your team can be a lot more productive. Since GA in January, we've been busy making GitLab Duo Agent Platform even better. Now, when you go log into GitLab, you'll see several specialist agents available for you that are task-tuned to handle specific goals for you and your team right out of the box, whether that is helping you plan what you want to work on next or fix security issues like what Alan and Michael had just shown, and many others. You'll also see built-in agentic workflows that automate complex tasks by chaining agents together in a predetermined workflow that we know works. For example, you can now, with one invocation, with one click, or one CLI command, go from an issue to working software, and agents take care of everything else in the middle, and many other agentic flows that are now available out of the box. Recently, we have introduced several agentic triggers, you and your team can automatically invoke agents when new code is introduced or when things happen in your environment. We've also introduced several manual triggers across different surfaces that you and your team work in, not only in GitLab, but in your IDE, in your CLI, and many other places that your team works. Let's see the power of Duo Agent Platform. Please welcome Shekhar, our distinguished engineer. Hey, Shekhar. How are you? Thanks, Manav. It's great to be here. Manav spoke about the platform. What I want to talk to you about today is how our customers and our internal developers are using the platform. I typically start my day by looking at my backlog. I've got this little work item assigned to me in the backlog, which talks about adding product search functionality to the homepage. Now, I could use the UI to do this, I prefer the surfaces that I use. I like to use my IDE, I like to use my CLI. Shekhar, maybe we can get the demo up on the screen first. There we go. There we go. Yeah. What I'm going to do is I'm going to quickly copy this issue. This issue's well-written, by the way. It's been written by Duo Planner, it's got all the details I need to actually start iterating on this. I'm going to go ahead and copy this, switch to my terminal. In the terminal, I'm going to go invoke the new Duo CLI. The new Duo CLI, now available. All I'm going to do is go ahead and ask it to implement this issue. Not even going to say issue. It's going to go ahead and because it has our organizational context, it's got all the rules that we need, it's going to go ahead and actually start implementing this issue. In the interest of time, it's going to take a bit of time. I'm going to go ahead and switch back and show you an issue that I actually implemented this morning. This is an issue I had asked Duo CLI to implement. What it did, it went ahead and created an MR. Exactly what we expect, it made a number of commits, it went ahead and made the changes that I needed. It went ahead and tested the MR as well to make sure it works. What's really interesting is that as soon as the MR was cleared, automatically, Duo Code Review went ahead and actually started reviewing this MR. It went ahead and made several recommendations. It said, "Hey, your styles are not in order. You can make a few changes as far as hard coding are concerned. You can go ahead and actually make some security changes as well." It's doing this because it understands my organizational context. It's doing this because it saw the review instructions that we've given it. We can have these review instructions at various levels. We can have it at the project level. We can have it at the group level, so that it's applicable to all your projects within that group. That's really powerful. That's really powerful because you can provide exactly what you want from an organizational perspective. You can do things like provide it exactly with CSS refactoring rules you want. You can provide the style of code you want. You can provide all the security things that are important to you, it'll do that. These review instructions can be applicable to certain files. We've been working hard on making the code review agent as good as it can be. We've been working hard on trying to make this code review agent as good as it can be, and we've been making constant improvements. In our own benchmark. You may have to go back to your podium, Shekhar. In our own benchmarks, as well as third-party benchmarks, we have made tremendous improvements. Now we are top five in the BigCodeBench, and we are extremely proud of that. Amazing. Now I'm going to switch back, I could have gone ahead and actually gone ahead and looked at these changes. What I'm going to do instead is I am going to ask Duo Developer to implement these changes. I don't need to do anything. I just go ahead and ask Duo Developer, "Can you please make these changes based on these recommendations, which are great?" Duo Developer actually went ahead and did this. It went and implemented the security fixes. It went ahead and implemented all the style changes that I needed it to. It went ahead and did all these things for me. That's really the power of automation. We've been doing a lot in terms of automation. As Manav mentioned, we now have triggers. Triggers, based on different GitLab events, can automatically invoke the agents that you have in your catalog. The triggers are really useful. For example, when a review is mentioned, it can go ahead and invoke an agent. When there's a merge conflict, it can go ahead and invoke an agent to actually fix the merge conflict. My favorite trigger is the pipeline events trigger. The pipeline event trigger, what it does, it goes ahead and fixes pipelines. As a dev, it's been a constant source of annoyance to actually go, whenever a pipeline fails, I have to switch context. I have to break out of my flow, and I have to go and look at what the pipeline is doing. I need to go and push a new change, figure out the logs, all of that. That takes me away from the flow. It stops me from what I'm doing right now, and I need to then switch context. The pipeline event trigger is something we've rolled out internally, and our teams love it. It's automatically going in, and every time there's a failed pipeline, it goes fixes it automatically. This is an example of a real project. This is our CLI project, and the pipeline fix trigger here is going ahead and running an agent which goes and says, "Hey, there's a race condition here. I'm going to go and fix that race condition." Or in this case, it looks at it and realizes it's a flaky test, and it knows that because of the organization context. It says, "Is this a flaky test? I'm just going to restart the pipeline to fix this." That's really powerful. To recap, I could have done all of these things. I know how to fix code. I know how to fix a pipeline, look at the logs, all of that, but every time I do so, it takes me away from the work that I want to do. This is where the Duo Agent Platform seamlessly fits in and slots in. It is able to fit into your existing workflow and help you automate the parts that you are interested in automating. Shekhar, it's not just you and your flow, because it's very common that if I'm writing code, I'll ask somebody else to review it, so I'm taking them out of their flow. Or if I'm running into a pipeline problem, it's common for me to go ask the DevOps or lead engineer to help me fix that pipeline, and I'm taking them out of their flow as well. This is all about giving you the right productivity for you and your teams to do what they do best. Exactly. There's a cascading effect. Duo Agent Platform lets you automate the way you want it to, and it gives you speed with control. Back to you. Amazing. Nice work, Shekhar. Thank you. What Shekhar shared there was really about helping you and your teams be a lot more productive. We've been looking at how our early customers over the last several months have been using Duo Agent Platform, and here are the top five use cases that we see across our entire customer base with Duo Agent Platform. Some of these ROI numbers are staggering. For example, with code review, we've seen our customers, on a per person basis, save a minimum of 20 minutes because agents are doing the code review for them as opposed to they themselves doing that code review. When you take that and the labor cost of doing that code review and the fact that a code review only costs $0.25 per run, that's 100x ROI. The rest of the ROI numbers are calculated similarly, and I can't wait. If you haven't used Duo Agent Platform already, check it out, see how these ROI numbers play out for you in your particular environment. All right. Let's move on to something else. This is not a technical challenge. I want to share with you a new commercial challenge that's showing up in the agentic era. You see, the way you buy software, any software, and GitLab for that matter as well, you buy software through fixed contracts. In the agentic era, what you need is constantly changing, and your fixed contracts force you to define what you need for the next year, in some cases multiple years, up ahead. In the agentic era, I've heard many of you say, "Hey, I may need more people in my company access GitLab because I want to give product managers and designers access to GitLab so they can contribute and code and help with the various projects that we are doing." I've heard many of you say that you don't know how much AI usage you will have a few months from now or a couple of years from now because AI technology is evolving. How your teams use AI is changing. How each person is enabled to use AI is changing. It's really hard to predict how much AI usage, and therefore, how many credits you need. With all of the new innovation that we just introduced today, and many more coming in the months, they're also going to be billed on a credits basis, and it is going to be really hard for you to predict how much you should budget for the credits you need for those new capabilities. The net is that the fixed contract model and the agentic era were not built for each other. That is why today we are introducing GitLab Flex. It is a new buying program that allows you to commit once, just like you do today, and then decide how you use the dollars you spend on GitLab, on which product and how much of which product, at any time. You can make that change as your needs change. To show you how Flex works, please welcome the Flex team, Courtney and Jerome, to give you a quick demo. Hi, Courtney. Hey, Manav. How are you? Good. Hey, Jerome. Hi. How is it going? Good to see you, Manav. Good to see you. Hey, everyone. I'm Jerome, Director of Engineering. I'm Courtney, Group Product Manager. Courtney, for this demo, how about I be Acme Inc.'s billing account manager? That way you can do most of the talking, and I'll just click the buttons. That sounds good. Hey, don't undersell it. You did build most of those buttons, after all. At Acme, we are on a Flex contract. We've signed a $1.2 million annual commitment, and we're currently six months in. Let's take a look at our setup. We currently have 500 Ultimate seats along with 50,000 Duo Agent Platform credits, both of which are reserved. That reserved part is part of what makes Flex really powerful. For the products where Acme has a good sense of what they'll need, they can reserve spend upfront and lock in a volume discount. How does that sound? That sounds great. I love saving money. For the products where maybe you're not quite as sure about your spend, you can enable per use and pay as you go, drawing from the same pre-committed pool. Flex offers you both discounted economics on what's predictable and the flexibility to spin up new things as your needs evolve. Speaking of which, how are Acme's needs evolving, Jerome? Let's take a look at the seats side first. For seats, we have 500 Ultimate seats reserved, and we're using pretty much that amount. The credit side tells a slightly different story. For Duo Agent Platform, we've reserved 50,000 credits, but we've already exceeded our allocation. Teams are really leaning into Duo Agent Platform, and AI usage is growing quickly. Hey, that's a great problem to have. That's the kind of split that a lot of customers find themselves in mid-contract. The shape of what you committed to in January is not necessarily how you're pacing come June. Let's take another example. A contracting team rolls off a project, and suddenly the 50 seats they were using last month are no longer needed next month. Now finance is asking, "We have all of this budget locked up in unused seats, but we're getting so many requests for additional AI spend. What do we do?" Under a normal contract, nothing. You would wait until renewal when you could readjust. Since Acme is on Flex, maybe we can see how that would work. On GitLab Flex, I can change our upcoming month's reservations. Here you can see I've got 500 seats right now. Let's bump that down to, say, 450 to account for the 50 contractors that are rolling off. For the GitLab Duo Agent Platform credits, let's increase this from 50,000 up to, say, 60,000, just to account for the increased AI usage. You can see that the overall commitment has stayed the same, it's been reshaped to match our needs. Wow, that seems really simple. I have to ask, what about budgeting guardrails? I know that's top of mind for a lot of customers. Our usage caps actually live here as well. I can set a usage cap of, say, 70,000 credits. This is slightly above our reservation amount, it puts a ceiling in case usage spikes. Okay, great. In finance, you can get predictability at the contract level. In 1811, GitLab added per user credit controls, which means that GitLab admins can allocate additional credits to power users while making sure no one user blows through your entire AI budget. As you all saw today, we've launched GitLab Orbit. Let's give that a try as well. I'll allocate some credits there, say 5,000 credits. Jerome, 5,000 credits? Did you not see the Orbit demo? Let's bump it up a bit. Okay. Let's do 10,000 credits. Better. With Flex, Orbit lives on the same rate card. I can just lock it in, and the commercial side is all handled. Acme's commitment didn't have to change. Their overall contract didn't have to change. What they're getting from GitLab evolved real time with their needs, without Jerome having to go through another procurement cycle to get there. That's Flex, the buying program that evolves with your needs. With Flex, you commit once and then adjust as your year unfolds. You get volume discounts on what you know and flexibility on what you don't. Whether you're running GitLab in our multi-tenant cloud, in a self-managed instance, or in a dedicated tenant, GitLab Flex is available today. Customers can now request orders, and your sales rep is eager to get you on board. Amazing. Nice work. Great job. Great job, Courtney. Thank you. As Courtney mentioned, you can use GitLab Flex today, whether you are a new customer, you have an existing contract, or have an upcoming renewal. If you go to that URL, you can request a quote and move your contract to GitLab Flex today and take advantage of everything that we talked about. All right. Let's recap what we saw today. GitLab, the DevSecOps platform you know, is now the agentic infrastructure. The motor system got a lot better with the next generation of Git, built for machine scale. The nervous system now has GitLab Orbit, so your agents work better, faster, cheaper, but more importantly, you can answer questions that you never could before. The immune system now brings agents to security and governance to agents, so you can stay compliant. You now have GitLab Duo Agent Platform that has gotten a lot better with new agents, new agentic flows, and new triggers available to you. Finally, you have GitLab Flex, where you can commit once and shape your GitLab usage as things change for you. That is the new GitLab. Thank you. Now to show you how all this innovation turns into value for you, our customers, please welcome our Chief Customer Officer, Sherrod Patching. Hey, Sherrod. Hey, Manav. Thank you. Yeah. Well, we've just shown you some of the latest innovations from agentic infrastructure in action. Agent actions that are able to go through next-generation source code management with rich context, all with the level of visibility and governance that you need. I'd like to tell you a little bit more about a study that was led recently by Forrester, the total economic impact study on the Duo Agent Platform. We're revealing these results today, and I am thrilled to tell you about what we found. 40% faster time to remediation. You saw some of what we talked about today, being able to bring context and potential remediation into developer flow, and a 40% faster time on average. 80% faster time for developer onboarding. I know many of you in this room look at time to first commit as one of your key metrics. Whether it's a new developer coming onto your team or whether changing applications, being able to find context there within flow, we saw on average an 80% faster time. And overall, we saw a 400% return on investment for these customers. They were expecting somewhere between 20%-40%, but as a result of Duo Agent Platform, we're excited to tell you that they saw 400%. Joining me on stage today, I'd like to tell you a little bit more about a customer story, showing you this in action. I'd like to welcome onto stage Mercedes-Benz. Mercedes is one of our customers that was able to actually transform how they think about the software development life cycle using GitLab. They were able to see across thousands of developers, the ability to actually go from what was the previous implementation, all the way through to the net new generation using software development life cycle in a highly regulated company with GitLab. To tell you more about this, I'd like to welcome onto stage Bastian from Mercedes-Benz. Welcome, Bastian. Hi, Sherrod. Hey, thank you for coming. All right. Have a seat here. Thanks. Thanks for having me. All right. I have a few questions for you. You recently launched the CLA, C-Class, and GLC on MB.os, Mercedes' in-house vehicle operating system. What does that innovation approach look like, and how is GitLab helping your 20,000 plus engineers on that journey? Well, first of all, you're right. We launched the CLA last year, followed by the C-Class and the GLC, and all run our latest version of MB.os, our in-house developed operating system. Not so long ago, we had several dozens of suppliers who equipped us with control units and the software, and integrating all of them was quite difficult. We moved towards fewer control units with more capable ECUs, where we developed a bigger part of the software. We control the crucial parts when it comes to autonomous driving, infotainment, powertrains, et cetera. Of course, this comes with some challenges. Different to maybe web or app development, embedded software development is kind of quirky, so we have to deal with embedded tool chains, some tools when it comes to SAST and test, which were never meant to run in a CI, but rather on Windows. Luckily, with tools like Fleeting runners, we can run also those tools at scale, running several million jobs regularly, moving several petabytes of data, of artifacts on the platform. We run the platform in different flavors where we have a greater degree of flexibility when it comes to our back end and web services for the connected vehicle. We can use GitLab Dedicated to use all the benefits of a SaaS product, but where privacy is key and we have the highest level of control of our data also being present in different regions of the world where required, we go with GitLab self-hosted to have this greater degree. When we started in-house software development, it wasn't always like that. When we started, we had different instances of Jenkins, Bamboo, Bitbucket, Azure DevOps, et cetera. Over the last years, we all brought this together to GitLab because GitLab gives us this unified user experience. We can govern in one place, can share best practices across all those domains, and are pretty happy now having more than 20,000 users on our platforms. Awesome. I love that. The different flexibility that you're able to bring and that consolidation over time. I think I've been there since the beginning of the journey- Yes With you all. It's been fun to see. Okay. Like every technology leader here, you are navigating AI for software engineering. With automotive software, there comes a level of safety, accountability, and review burdens that most applications teams don't ever face. Yeah. How are you approaching the use of agentic AI for your software delivery? We are doing a lot of different approaches. I think the key is an agent can only be as good as the context and semantics which are fed to them. Therefore, I'm also super excited about what we saw already. Context and semantics in our domain, besides the source code, of course, itself means the functional requirements, but also the non-functional ones, safety constraints, architecture patterns, et cetera. Key is that we move this data out of proprietary tool silos we maybe had in the past, but make it accessible within Git or in graphs so that the agents can operate on that to have the right input. At the same time, also the validation, what we do in CI becomes even more crucial, that we tighten the loop, where we see what the agents did, but also agents ideally can correct themselves, also what we saw just a couple of minutes ago. Yeah. What we really also like is that with GitLab, we have this flexibility. We can use the GitLab Duo feature set, but we can also connect other harnesses like Claude Code or so, because, it is still a young thing. We're exploring a lot of options, and having this flexibility without over-committing into one locked ecosystem is a great advantage. For Duo itself, I can only echo what we also just saw, chat and code reviews are amongst the most loved features on our end. Yeah. Fantastic. Thank you. I know we've been talking a little bit before this also on just the context that GitLab Orbit will bring and the ability to be able to make the decisions within the MR and the time savings that you'll see there. Exactly. Great. Okay. You have an internal framework on AI native engineering. As we launch GitLab Orbit public beta today, how does a richer SDLC context fit into that picture? Well, I think it fits perfectly. With AI, we can get at such a high pace, but also, as we heard already, we need speed with control, and we have automotive regulator standards like ASPICE, which require traceability and human accountability. Accountability is especially important for us. We put people in our cars who trust their life to our cars. We, as Mercedes, have to stay accountable that the software is safe and sound. I think mastering this human in the loop approach, letting agents freely where they can run freely, but then again, having the humans in the loop, reviewing the results, approving, giving consent to what the agent did is right. I think this is key, and there we see GitLab is very well set up as this control plane where we integrate the software and take that accountability. Awesome. Thank you. All right. Last but not least, a favorite question. How are you thinking about measuring whether or not AI is actually improving productivity? Yeah. Well, I think first of all the DORA metrics we did in the past in measuring productivity like DORA metrics, et cetera, are more valid than ever because ultimately, AI is a means to an end in becoming more productive. Also at the same time, we have to justify the spendings, of course. I think everybody who uses this knows there are good and bad patterns how to use AI. What we would like to see is a good integration of the DORA metrics or productivity metrics with AI consumption and insights, how our users are using that, seeing best practices, where are we efficient, and I think you have also a nice guest coming up. We dive into that. Looking forward. Yes. We'll show him in a moment, too. We're excited to have Gene here. Great. Well, Bastian, that is the last of my questions. Thank you so much for joining us today. My pleasure Mercedes-Benz on stage. Thank you. Thank you. All right. As you can tell, GitLab has the potential to transform not just software development, but also, too, what you see on the road and the experience that you have. I love the fact that Mercedes is just a fantastic example of speed with control. There's one more element as you think about what needs to really happen for rich software development life cycle to be true for customers like Mercedes and others, and that is our ecosystem partners and their presence in bringing all of our customers from agentic, essentially agentic testing all the way through to agentic engineering. One of those key partners for us is Google Cloud. Here to tell you more, I'd like to welcome back onto stage Bill and Daniël Rood from Google Cloud. Good job, Sherrod. Thanks. Thank you. Thanks. Oh, hi, Daniël. Hi. How are you? Well done. All right. Daniël, thank you so much for joining us today. Google and GitLab have been partnering for years. We have. In fact, it's, I think, maybe the best-kept secret. I don't know if people realize, but thousands of customers benefit from the partnership every day because gitlab.com runs on Google. It does. I know many of our strategic customers also choose Google as their cloud infrastructure provider for their self-managed instances. Tell us a little bit more about what new options we're introducing today. I'm really excited to announce that for providers, GitLab certified managed providers, there's now an option to deploy GitLab on Google Cloud with sovereign deployment options in EMEA. I think this is really important, especially for regulated organizations, as if you think about those workloads that need to be compliant within certain regions or against certain regulations, we now offer you the controls in order to do so and for you to be compliant with your auditors. I think that is really big news for today. Customers are going to be really excited about that new option. Get out of the toil of managing your own instance on your own infrastructure and take advantage of Google's managed service providers. Amazing. You're not only a cloud infrastructure provider, you're also a model provider, and we've been proud to offer the Gemma and Gemini model support inside Duo Agent Platform. What else do you have for us on that front? Maybe as an introduction, if you think about the Google AI family, there's a number of models that we offer for your customers. We have Gemini Flash, which is our working horse model for your thousands of times a day workflows, and it's cost-efficient, it's token efficient, and it's still a state-of-the-art model. We have Gemini Pro. Pro is really a powerful model for your most complex workloads. Equally exciting is our open weights model, Gemma 4. Gemma is an excellent model for those who want to run these capable models on the edge, maybe even on device or in air-gapped solutions. All of these models come together in Gemini Enterprise Agent Platform, which is tightly integrated with Duo Agent Platform. We are offering that today. The news for today is that we announced just a few weeks ago, Gemini 3.5 Flash, which is now also available in Duo Agent Platform. Woo. Awesome. I know Gemma 4 as well. Gemma 4 for our self-hosted customers in air-gapped environments, because we have many regulated customers who are required to be air-gapped, is going to be a really powerful option as well. Thank you. Speaking of the Gemini Enterprise Agent Platform, we were talking earlier about cost and how everyone's talking about the cost of AI and the ROI and being able to understand where the cost is going. Duo Agent Platform provides the visibility into the cost of agents running within GitLab and the use cases that are under action, what does Google provide in that front? Yeah. If you think about the partnership, there's probably a couple of elements there. First of all, we talked about cost-efficient and token-efficient models. Like a Flash or a Gemma 4, they will help you make the right decisions for your AI workload. I think that's an important one. For those who have a commitment with Google Cloud through the Google Cloud Marketplace, you're now able to draw down against that commitment with GitLab. I think that is really important because what that also does is it gives you the flexibility not entering into new budget cycles. It gives you one view of your costs all the way from GitLab platform through inference and infrastructure with one bill, the other Bill, and also one view of everything you're doing. What is really important there is a lot of the tech leaders here in the room and online as well as I'm sure your CFOs, they all are interested to understand how much are we actually spending on AI this quarter, or even how much value are we getting out of our AI usage. Those are all questions that we can now start to answer with that setup. I think that is really important. If you think about as a customer of GitLab, this go-to-market integration between the Duo Agent Platform and Gemini Enterprise Agent Platform as part of the marketplace, you now can manage your AI workloads end to end. Amazing. Really we've talked about three things today. GitLab as a managed service on Google Cloud, now available through managed service partners. Gemini 3.5 and Gemma 4 in GitLab Duo Agent Platform available today. Use your Google commitments to buy GitLab licenses and credits. Yes. An amazing partnership that continues to get better and better every single year. We're excited. Thank you so much, Daniël. Really appreciate it. Thanks very much, Bill. I'm definitely also really excited about all the things that are coming up in our roadmap in the coming weeks and months. Yeah. Look forward to that. Thanks again, Daniël. Thank you. All right. The agentic engineering is here. Hopefully, you're starting to see how we're bringing agentic infrastructure together with coding agents to deliver speed with control. You've heard amazing customer stories from Mercedes already, and more to come, about the proven ROI of what we do, both in GitLab and now with agents and Duo Agent Platform, and amazing partnerships like with Google and earlier Anthropic. It's an incredible time to be part of GitLab. I hope you're starting to see our new mission in action, which is to unlock every team to ship trusted software at the speed of imagination. I think the most exciting part of Transcend though is often not all of the technology, the amazing demos, but what I hear most often from customers is they love hearing from our partners and our customers about how they use GitLab and the benefits they're seeing. It's my pleasure to welcome back on stage Sherrod and our panelists. We'll just go in order and then we'll go from there, I think. In fact, no, Gene, I'll start with you. To start, when do you think about software engineer or innovation in your organization or in those you advise, what has changed the most in the last 12-18 months because of AI? Oh my gosh. What hasn't changed? This morning is probably proof of that. Having studied high-performing technology organizations for 27 years, I've had a lot of fun in my career, but I think like so many of you, I've never had as much fun as I'm having right now. It's so strange where we're entering this era where coding for planning purposes is becoming free and instantaneous. Well, I don't know about free, but it's certainly pretty close to instantaneous, and that means like every process we've created is now wildly insufficient. Budgeting, procurement, prioritization, getting access to customers. We heard from Nick, from Anthropic mentioning how their teams are generating eight times more output than a year ago. Some of you might say it's just the frontier AI labs, we heard from Angelo from the GitLab Orbit team saying that in one month, I talked to him this morning and they made 3,000 merge requests. That's tens of thousands of commits in a month. Yeah. Some of you will be excited by that, some of you will be scared by it, and some of you will just say, "Oh, that's slop." I think as leaders, we all have to get ready for this era where that's going to be, I think, increasingly commonplace. Yeah, I agree. Thank you. Matteo, maybe we'll have you go next. What are you broadly seeing with AWS? Sure. For me, when I speak with enterprises, I've seen that in the last couple of years the economics of the coding kind of flipped. It used to take maybe one day to build a feature and maybe a few hours of code reviews to actually get it to a measurable state. Today we see that with AI, we can actually have code produced in minutes, but maybe still needing hours to actually steer it back to actually get to measurable state. What I see a lot changing In the last year or so in enterprise, developers are learning that context engineering, elaborating their intents better, and using AI to prioritize helps getting a result faster that actually looks closer to our intended outcome. This is basically, for me, a signal that models are rewarding discipline rather than speed. Fantastic. Thank you. Ryan, what about at Compare the Market? I think for me, the most obvious change, especially for a group of engineers, is where people are spending their time. You'll have seen you're able to produce and write code significantly faster than we were a couple of years ago. That's forced people's time to move either side of the code writing process in the delivery pipeline. People are spending more of their time writing and refining specifications and how work will be conducted, and then also reviewing the output, right? This is kind of interesting because the bit that we really like about engineering is the code writing part, but it's forced us to go either side of that. For us, that's looking at how we address the changing nature of the role, and the impact that that has on a whole bunch of people on what they find valuable and what they're doing. Yeah, of course. Thank you. Mans, what about for you at Cube? Yes. From my perspective, what I see, it's not about what is changing, but how fast it is changing all right now. Indeed. One to two years ago, we were talking about how are we going to implement code suggestion kind of tools into our software development lifecycle, right now we are running multiple agents in every stage of our development lifecycle within GitLab, also within Cloud. I think it's not only a developer tool or tools anymore, but it's more like a full organizational shift where we are in right now, where we see that the way we think about even building software is completely changing. Yeah, completely. I think we have a case study that came out just today with you. Yeah. Some exciting stories there. Ryan, next question is for you. Cool. Compare the Market has been doing some of the most concrete work we've seen on how context changes AI outcomes in software engineering. Tell us more about what problems you look to solve and what you learned in the process. Okay, cool. We saw the same as everyone else, that the volume of code that we were outputting was growing significantly, that again forces the impact on people's time into the code review process. We have an agent that does code review on every single change we have, that's in our GitLab pipelines. Yeah. One of the things for us that we were really interested in is how can we arm the person who's still doing the code review with as much context and meaningful feedback on the code change that's being suggested. The agent was performing these code reviews, we wanted to make sure it had as deep context and meaningful impact as possible. The team, by the way, a couple of them are in the audience, have a chat with them, Marina and Wenye. They're awesome. We wanted to have a look at how do we make the code review that the human gets from the agent as meaningful as possible. We did a study within the team on how we could do that, we took a bunch of different approaches. The first was using Knowledge Graph and Orbit. Yeah With an agent, the other one was using an agent with a RAG tool, the other one was using an agent just on its own, no tools. What was really interesting for us was firstly, using Orbit, using Knowledge Graph significantly outperformed anything else. Kudos to you guys. The sort of evaluations we've done showed that comment accuracy was 21% higher than any other option. Most interestingly for us was that using an agent on its own with no tools outperformed an agent with RAG. Yeah. Sort of digging into this was quite interesting because we found that using a RAG tool with the agent, we were dragging in semantically similar code into the context window, but that was causing a little bit of confusion within the agents. That's why we see that underperformance for RAG. For us, it's been a quite significant unlock. We're getting through, let's say 1,000 MRs per week, and if you're shaving an hour off of review time because you're providing significant context to the human, that adds up over a quarter. Yeah. Fantastic. Maybe just as a quick follow-up, speed with control, we've been talking about the speed with governance. How do you think about that one? Yeah, that is interesting. We were actually talking about this last night too, that the position for us is quite interesting that coming from a sort of regulated, governed industry, that we find ourselves not in the position where we're having to adapt to AI and tag risk and governance onto the side. Because we've operated in that space, we find ourselves in a fairly advantageous position that the guardrails and risk and compliance checks that we perform as part of normal pre-agentic software delivery are kind of already there. Sure. Sure, they're changing. Sure, they'll adapt and the risks that we address, there'll be new ones that we haven't seen before. The way we think about software delivery is that those things are already part of what we do pre-deployment going through our pipelines. Yeah, things are changing. The nature of how we do deployments and how we think about delivery will change. We're in a fairly good position where we can go fast because we already have those guardrails in place. That's great, and the discipline is there. Yeah. Fantastic. Gene, this one's for you. We talked about some of the study that Compare the Market did, and the surprising result around agents with no context being better performing than that with RAG. Maybe, as you think about someone who has studied, as yourself, who studied these system problems for decades, what do these findings tell you overall? Oh my gosh. I think one of the things, I love the work that Ryan and team did at Compare the Market. I think one, it's just fantastic primary research about what makes these tools more productive. Secondly is, we're all learning together. It's what an incredible opportunity in an era where nobody knows what the new practice will actually look like. Here's an opportunity to actually define those patterns, and I think Ryan and team absolutely did that. The fact that they did that in a regulated environment, I think is fantastic. It reminds me about what happened when DevOps in 2010, where organizations, especially in regulated industries, they were just so scared to even say the word CI/CD or DevOps because they were afraid that the regulators would crack down on them. As we got case studies like Capital One, one of the largest card issuers in the U.S., we've sort of normalized that and say it's actually better, you're faster and more secure, and more in control. I'm eager for when we get those case studies down where we can actually make it safe for organizations to say what they're really doing. Yeah. I ran a conference called the Enterprise AI Summit in April, where we had Block, Netflix, Skypoint Healthcare, where companies in regulated spaces are actually sharing that they're working on regulated mission-critical code using AI. That's an exciting time to be in the game. It very much is. Maybe one more for you. You've written extensively about how organizations create flow. How does context fit into that picture? I think it's everything. I guess the thing that really amazed and amused me this morning is how intolerable it is when something takes two minutes. Like, oh my gosh, two minutes. That used to be considered fast, but now if you have to wait to get information from your repo, and it takes two minutes, it is intolerable. It's just an exciting time where, what does it take to get not just the right context, but get it quickly? It's just exhilarating. Absolutely. Thank you. Mans, this one's for you. We've heard about context quality, and we've been talking about that now. You described a model where Claude handles generation and GitLab does the orchestration of everything around and across the software development life cycle. Maybe can you walk us through that use case and the impact that richer context had for you? Yes, of course. I think context and quality of the context is the most important thing in nowadays software development. At Cube, we have been running GitLab for over eight years right now. Our full software development life cycle is managed within GitLab. From issue creation to the actual deployment, it's all within the GitLab environment. Yeah. When we started adopting AI over two to three years ago, it wasn't a question of where is our context at, but more how are we going to implement AI agents within our existing GitLab environment. We are doing that in two different flows right now. One of them is that we are using Claude Code as our daily coding agent for our developers, and we connect that through the MCP and API connection with GitLab. Yeah. Yeah, we keep in control of our software development life cycle, and when we are doing the actual building of the software, we pull it from GitLab context to in our Claude coding agent. There we are building the software, putting it back into our software development life cycle within GitLab, and that's how we are currently building our software with our teams. Besides that, we are also using the Duo Agent Platform, where we are building custom agents within GitLab. For example, when we want to have an agent which is gathering context before the actual development starts, we are implementing that in the first stage of our GitLab flow to get the context in our issue before it gets into our Claude Code development environment. Yeah, what we see is GitLab is orchestrating everything for us, and we are plugging in Claude Code now to do our actual development work. For example, it can also be an auto coding agent. Yeah Near future. Awesome. Thank you. I know a number of our customers are interested to hear more about how these work together. Thank you for sharing. Yeah. Maybe one more for you. As you're shipping faster, what does that mean for your business and for your customers? Yeah. What we see is that we can ship a lot faster, but also the quality is increasing. We are delivering more and higher quality software. Security is also getting better and better. From there on, we can deliver faster, for example, prototypes, where we, earlier needed for months to weeks to develop the first prototype. Yeah. We are now ready in days to weeks, we can show them the value that we can deliver with our software. What comes with that is that we see a shift from the hourly-based software development, where we are shifting to more value-based software delivery. It's not only about the hours anymore that the developer spends to develop the software, but it's also about the AI cost, the agents that you're running. Yeah, we are figuring that out, how we are going to make that shift as a company. Great. Thank you. If I can just add one more thing about that. What I'm really looking forward to, the DORA metrics have come up and that was something I worked on about a decade ago, and I'm really looking forward to the day when we have a set of metrics that can actually share what you just talked about. You can conjure up software from scratch in an hour, right? Yeah. Right now, just talking about merge requests and pull requests, and lead times. It's such an incomplete expression of the magic that's happening right now. We're not there yet, but I look forward to that happening soon. Me too. Thank you, Matteo. From the AWS side, how are enterprise customers defining business success with agentic software engineering? Maybe use cases and how these drive investment. What we see with our enterprise customers is a shift from individual productivity to group productivity. Over the last couple of years, every developer started using these tools and recorded some productivity when coding, and everyone around them, other roles, tech and non-tech roles, started doing the same. We saw in the enterprise that in order to achieve greater outcomes, sometimes it is necessary to align tools and technology with people and processes, and in general, rituals and ceremonies. With people and processes mostly I mean the evolution of roles. We see in many enterprises some of the responsibilities that used to define clear boundaries of roles responsibilities becoming a little bit blurrier. For example, engineers taking over some of the product management duties because this is bringing more efficiency when interacting with AI, just as an example. When we think about rituals and ceremonies, we see sometimes smaller teams working on shorter sprints in order to actually work a little bit more efficiently. When I think about use cases, given this is common for all the enterprises, I would say one prominent one is brownfield modernization. In the enterprises, there is a lot of legacy, sometimes spanning multiple repos that have implicit dependencies and a lot of tribal knowledge. When AI can help us understand this code, we can actually immediately see some of the ROI because we can see maybe moving from releasing a change that used to take months now maybe going live in weeks or days. I would say the second use case that is quite emerging, we spoke about it today, is infusing every step of the SDLC with AI. Not only coding, but also code reviews. Everything related to operations and orchestration of going live into production. I would say the third use case is also something that goes beyond developer and maybe it goes and touches in other phases of the SDLC that maybe relate to, for example, product managers or design or user experience experts. For example, using AI to do data analysis, analyzing customer signals, maybe doing user testing with synthetic personas rather than real personas, or complementary to real user testing. Prototyping, even. This is probably something that, again, is quite prominent in the enterprise because we see it is easier to connect that to some kind of ROI. Fantastic. Thank you. All right, Gene, I'm conscious we have three minutes left, I'm going to ask you another question, then I'm going to do a quick whip-around at the end. In your research, what separates organizations that make AI work at a systems level from the ones that are stuck in just deploying tools? Oh, my goodness. I'll actually quote someone who was part of the dev productivity teams at Amazon for the software builder experience. There was a cross-population study as they tried to implement Andy Jassy's edict of "Everyone has to use AI." He said when he studied 15 teams about who really excelled, he said it was really three things. It was AI fluency, how good are they at AI? How much have they practiced? Two is, do they understand where the bottlenecks are? The third one I thought was really intriguing was the quality of the leader. Is the leader focusing on improvement, making time to get better at our craft? That just really resonated with me. I think that definitely distinguishes and resonates with me. Awesome. Thank you. All right. We're going to do a quick whip-around with this final question, 30 seconds each. Software engineering is changing fast. To close out, what is a wild prediction from each of you on what comes next? Ryan, I'll start with you. This might be unpopular. I think natural language is the only programming language you're going to need to know. We see this internally. Teams who are spending their time refining specifications are significantly more productive than those who don't. Okay. Thank you, Matteo. Everything is changing super fast. I would say my prediction is that this year, next year, every AI engineer will be raising agents and nurturing them. My prediction is that this will require more technical skills rather than not. Awesome. Mans, what about for you? Yes, I think every company will have full agentic teams, but also agents which are managing their own budgets, hiring other agents when needed, scaling up when needed, scaling down if needed. From there on, humans are only setting direction, giving the goal and the right context to get to that goal. Great. Thank you. Gene, take us home. I guess I agree with everyone. I think we're starting to see glimpses of a world where everybody codes. It's not just developers. Where marketing people code, UX, design, CFOs, CEOs. I think it means we're going to 100X the number of developers we have on the planet. You do the math, there's $20 million now, times 100 is about $2.8 billion developers. That's about a third of the world population, and that feels right to me. All those graphs Manav showed this morning of the growth rates, like, "Get ready. More is coming. It's just the beginning. Yeah. Streaming platform and I was in charge of engineering when they went through a very rapid growth. This inspired some of the research that we're doing. I got connected to the people at Stanford University. One of them is Yegor. You see him here. Unfortunately, he couldn't be here today, but together we kicked this off. Just to give you some context, we've been doing this quite a while. Our research has been shared by Elon Musk. It was notably the piece on ghost engineers. Marc Andreessen, we do various events, and also the mainstream media picked up on our publishings. Obviously, we also submit to the major AI and software engineering conferences. We publish a bunch of papers each year and all the research is ongoing. Before we dive into that, I will have to explain you a little bit on the methodology that we're using. How do you even measure software engineering productivity? We didn't know that when we started. There were things like counting commits, counting PRs, counting lines of code. None of that seemed something that is really a good way to measure software engineering productivity. We were kind of trying many different ways to figure out what could work. What seemed to work is actually an expert panel that looks at code written by the engineers. They give their feedback. We ask them questions on implementation time, quality, maintainability, complexity. Then, that was the first surprise here in the study. The experts were in very high agreement. If you guys have been in engineering meetings, it's really hard to get engineers to agree on anything. For us, that was a big surprise. It was exceptional. We used something called the intraclass correlation coefficient to calculate the agreement. Okay, we found a way to actually measure it. Can we do that at scale? In order to do that, we tried to train a model that would replicate the expert panel so that we could look at it at thousands commits in very little time. Currently we have hundreds of companies enrolled in the studies. I haven't updated that number in a bit. I think we're north of 200,000 engineers that were analyzed. The beauty is we can go back and, I think roughly this represents, depending on how you calculate, maybe even 1% of the software engineering population. You have to assume that this is not a perfect way to look into productivity, but better than other metrics. All the following slides are based in that methodology. If you assume that now we have a way to measure productivity consistently, let's take a look at what AI is actually doing. We'll have a couple of sections. We'll be looking into how AI benefits are unevenly distributed, how structured practices are important. We'll look at a real company case study, and we'll give you some benchmarks on AI spending because it seems a lot of people have question marks on what's the appropriate budget allocation and some organizational implications. The first finding here really is that AI is not really delivering benefits to everyone at the same time. What I can show you here is that we picked 46 teams that used AI, and when we started this, we had also a control group of 46 teams that didn't use AI. While we kept doing this, our control group fell apart at one point in 2025 because there were no more teams not using AI. We kind of had to extrapolate from the control group in the beginning. Initially, we had a 4.8% difference and now the recent update really is a 59% difference in output. The gap it's gotten wider. Maybe some of you saw that Fable is out, the Mythos-class model. Let's see if we see another spike in the gap. If we look at more recent data, this becomes even more dramatic, right? The bottom quartile teams get almost no benefit still. The top quartile teams often double productivity. The same technology, very different outcomes. We see a very strong power law effect. The key takeaway here is access to AI is not the differentiator. If that's what you're using to measure, you've got to start looking into who is successful at using it. If we take that from the team level and also take it down to the individual level, you see the same effects. Heavy users outperform the light users, but team effects matter even more than individual effects, right? Even if you have a top performer in the team, they will likely be slowed down by the laggards in the team, and they won't be able to deliver to their full value and possibility. That means AI productivity is a team phenomenon and, in the same time, this also changes who succeeds inside of organizations. Throughout our measurement, when we looked at who is advancing through the performance quartile. What you see here is Q4, that's the top-performing quartile. Q1 is the bottom-performing quartile. We have not seen a lot of movement in the past. The P value was pretty stable, 0.70. Now the rank stability, it fell to 0.45. Since AI, and we haven't really seen that in any period, no matter what the change was, whether it was going remote, any kind of transitions we had in the past, the rank stability was never that low throughout the study. Surprisingly, we see actually a lot of movement upwards from the bottom quartile to the top quartile. It's interesting in a way because we think, and we're hypothesizing here, and based on interviews that we conducted, is that we see that maybe very senior engineers that were doing supporting functions, supporting tasks, code reviews and whatnot, now they get time to delegate these to a model, to an agent, and they can contribute, and they're outperforming everyone. Yeah. The takeaway is AI is changing the skill set matter, and that's why we see those changes. People ask us, "Okay, if we give people AI, then what happens? Does more AI usage deliver better results?" The answer is not necessarily. We see that the successful teams, they work in a clean engineering environment, and clean environments achieve larger AI productivity gains, which is not really surprising. For me, as an engineer, I'm somewhat moderately offended by that because what organizations haven't done for the human engineers, they're now doing for their LLMs. Yeah. The key point is the environment quality matters, and the reason is that clean environments allow AI to operate more autonomously. You can see that also when you look at the task composition and the environment cleanliness. When you look how things come together, there is a dipping point where if you fall below that threshold, then it's just the agents can't really deliver good results. AI just amplifies the environment in which it operates. Yeah. Let's keep moving. How should companies then measure whether AI is actually delivering value? I think what we could do is, and ideally we would look at business outcomes. Ultimately, I think that would be what would be best. It's a very noisy signal because there's just too many confounders. In absence of an ability to properly measure that and correlate that, we recommend looking at the engineering outcomes. That's a relatively clean signal. Then that gives you a pretty clean framework. You can start measuring the AI usage, you start measuring the AI outcomes, you connect the two, and you avoid jumping directly to revenue conclusions, essentially. There are several practical ways to measure the AI adoption. I was hitting on that a little bit earlier. There was an access-based way of doing it. It is essentially companies that are in the rollout phase. They want to make sure everybody can access it, everybody can use it. In the end, it is not ideal, right? What we saw is that access and usage telemetry is the gold standard. While people having access is good, is better than not having access, but because of the discrepancy showed on the earlier slides, you need to double down on the people that really are killing it. We can actually look at that retroactively in our study because we have the Git history, and that is that. The next thing is, okay, how do we measure the engineering outcomes? Here is how we think about that. Essentially, the primary metric that we're using in the study is the engineering output as per the expert panel and per the machine learning algorithm that we've trained. What we use as guardrails is rework, refactoring, quality, tech debt risk. Also the DORA metrics are super important in terms of measuring flow efficiency. Happiness metrics are useful to check in on your team, but not necessarily as a productivity metric. The takeaway here is you can try to maximize the output while keeping the guardrails healthy. With that being said, I think we're ready to move to the next finding here. What we did, and we submitted this paper to ASE, the conference is in October. We have peer reviews and they're very favorable. We have defined four levels. Ad hoc prompting, rules and project context, task-specific agents, and orchestrated multi-agent workflows, which where we saw is the best results are delivered. The point here is that the AMI maturity leaves artifacts that we can analyze. We built a classifier to identify these artifacts and do an analysis of what is going on there and how your engineers and teams are talking to their agents. We analyzed hundreds of repositories. We used embeddings, paths, and content. We use that to detect the actual AI maturity signals and got really strong validation results in the sense that we saw strong clustering effects that tie into a higher performance, higher output. That maturity has a measurable impact on quality as well. When you look at it, repositories with no structure, they suffer a lot more degradation in terms of quality as you keep using your agent, and cognitive complexity for the engineers keeps increasing. Static warnings go up. No matter how you slice and dice it's not a good idea without proper harnessing, proper instrumentation and tooling. Essentially, structure protects your quality. When we look at individual developers, we see the exact same thing. We see the PR throughput is dramatically improved, duplication is decreased, and revert rates are improved. There is only benefits that we see here. There is simple documentation, context practices. Those really create outsized return. There is no measurable net negative effect on that. Everything you do in that direction, you will have some gains. In order to validate that, we can look at an actual case study of a company that did that. It's a real enterprise example. We tracked output quality and churn together, the company really started at a below average output. At one point the CTO said, "We want to 2X everything." We had the CTO mandate. Nothing much changed in the metrics initially. What really started to change things was, first of all, the rollout, the adoption. Access actually didn't matter. They adopted AI in May 2025. They started in the 25th percentile. Productivity doubled. They moved towards the 60th percentile now. The large gains are achievable, even at 600 engineers. Productivity gains alone, they don't really matter if the quality collapses. When you look here at the quality analysis of the company, before AI, they were more or less stable. Sometimes it was going down, then it was going up again. When they focus on, you probably know how it is. You have to ship something quickly, it degrades, then you spend a little time in optimizing, improving it. It goes up again. In the beginning of their AI adoption journey, you see a steep cliff when quality went down. With proper tooling and instrumentation, they essentially were able to reverse the trend. Productivity remained high while quality stabilized. Recently in the last few months, they have even achieved an improvement in quality in their agentic workflows. Quality degradation is in fact manageable at this point. The next metric, it's also giving you an important story. The churn rate, essentially we cluster rework and refactoring in that churn rate metric. That one is also down. AI improved also execution quality, not just speed. Yeah. Understanding what causes these changes is now really the next challenge. What we are trying to do now is we're trying to put a pilot program together where we tie all the engineering metrics, they show correlations. The correlations are not causes. The real drivers here may be meetings, calendar load, we don't really know yet or don't understand yet how these things tie together. This is for us the next step in our study to really find the causes and not just the patterns. If anyone here in the room is interested in joining that effort, feel free to reach out. Let's switch gears from productivity to investment. Everybody's wondering also what is an appropriate amount of money for me to spend? At Stanford, we launched the AI Spend Index, where we gotten consent to publish summaries to spend from companies in our study. We track AI spend per developer. There is benchmark organizations. What we see, the high-performing companies, they spend significantly more relative to the ones that are in the lower performing quartiles. Under-investment can become a competitive disadvantage. However, when you look at that, the benchmark, it becomes more valuable when you can look at your peer and industry data. It's similar concept to levels.fyi. You can compare against your peers, contribute data, essentially, you can unlock more visibility. The value is really where you sit in your cohort and how you do relative to them. You can compare based on a few things. You can open up the website and see where you sit and hopefully contribute also some data back if you find it useful. Yeah. I think this is now the science part that is pretty well understood. We're entering a bit conjecture territory here because we're really trying to tie everything back together. What we see often is that AI speeds up the individuals, but a lot of organizations that we have in our study, they actually fail to capture the gains, and we want to get a better understanding on why that is. We see that lower success correlates very highly with size. The hypothesis here is that enterprises spend a lot of time on internal alignment. It's in principle, people know what they could do or should do, but then you need a lot of time to get everybody who needs to be bought in, bought in, and get that done. It's a lot of meetings, stacks, approvals, politics, and AI doesn't automatically eliminate these costs or improve it. A lot of productivity gains get absorbed by coordination and network complexity explains why. When you look at it, a very simple thing, we've published on that a few years back already. We see as the number of nodes go up, the communication overhead just becomes crazy. That's why single-threaded ownership tends to become so important. Every time now, it's amplified by AI. While you could move faster, you kind of lose the wins or the benefits in alignment. This is also the reason why startups tend to benefit more than enterprises. Then you really have a challenge with regards to organizational design. What we see is AI native companies, they're really operating fundamentally differently relative to the classical traditional organizations that we have in the study. They're really built to capture 24/7 engineering. It's non-stop. You got your agent running. It's not just another tool. I think there's scaling through compute. There is persistent organization knowledge and compounding capability improvement. What we see is they tend to have found a good way in order to remove any blockers where a human decision would be the slowdown. Yeah. With that being said, if you're interested in enrolling your company in our research programs, there is a number of ways to reach out. You can participate in the research. You can contribute to the AI spend. You can help us explore more causal discovery, discuss company-specific opportunities. We're also open to that. Yeah. I think hopefully this was relevant to you. In our view, the organizations that learn fastest how to measure and operational AI, they will capture the majority of the gains. Thank you so much. That's it from my side. Thank you, Simon. Thank you for sharing your research with us. This brings us to the end of today. Whether you're tuning in online, thank you for joining us. To the developers in the room and online as well, don't forget the developer show and also the hands-on lab.
Loading workspace