Good afternoon, and thank you for attending Rosenblatt Securities' sixth Annual Age of AI Technology Summit. My name is Kevin Cassidy. I'm a Semiconductor Analyst at Rosenblatt, and I'm very pleased to introduce, from Rambus, Steve Woo and Matt Jones. Steve is Rambus' Fellow and Distinguished Inventor, and a longtime favorite presenter at this conference. Thanks for coming back again, Steve. Thank you for having me. Yeah. Matt Jones is a Senior VP of Corporate Strategy. I think between Steve and Matt, we've got some very knowledgeable people around the DRAM market, technology and markets. I'm going to start off with a few questions for Steve to help investors understand what's happening in the DRAM market. The hot topics are AI, of course, and processing and memory. If you have questions from the audience, please go to the chat section of your Zoom screen and type that in, and I'll see it and ask the question to Steve and Matt. Thanks again for joining us. Maybe Steve, I'll start it off with agentic AI has caused this fresh demand for CPUs, and I think it's top on investors' minds these days of where did this come from? CPUs are now cool again. What does that mean from a DRAM point of view? Yeah, it's a great question. We're seeing this big sea change that's happened in the last few months. I've actually got a couple slides that I'd like to show here, and I can walk through a little bit about how we got to where we are today. Let me go into slide mode here, and then hopefully you guys can see this. Yeah, I see it. Okay. Yeah. I think it helps to understand where we were in the recent past, and then that helps us understand how we got here. If you think about what happened in the beginning of AI when we were talking about, "How do I accelerate this?" What you do is you take the best components you have at the time, and you figure out how to organize them in a way that makes sense for your application. For AI, what that meant was take a bunch of existing GPUs, and then we'll build these chassis where you have a lot of PCI Express slots, and you can put a lot of these cards into one big server. You can see here in the front of this server, there's a whole bunch of cards here. I think there's, like, 10 cards. That became an AI server, and that helped us really advance AI. Well, if you think about over the last three, four years, what started to happen was people thought, "Well, it's really nice to have a server box that can help me accelerate AI, but I need something at bigger scale. I need something that looks more like a rack." People started to make rack scale systems. What they started to do was each of these units that you see here, there's eight of them in this rack, they're actually composed of two halves. There's a half with a bunch of NVIDIA GPUs, and there's a half with a couple of Intel CPUs. What you started to see is people started to come up with very unique form factors, and they really started specializing the silicon for AI. Once the market took off, you started to see this specialized silicon start to be used. What you also saw was this two kinds of memory that were meeting in the same platform. There was the high bandwidth memory, HBM, in this case, HBM4, for example, was being used for the GPUs, then lots of DDR memory was being used for the CPUs. If you think about where we are today, what's really exciting about it now is we've got this interesting use case, large language models. They are so interesting and so important to the industry now. In fact, there's specialized silicon being developed not for the whole LLM process, but for subportions of the LLM process. I'll talk about this more in a minute, but there's two very important phases. One's called prefill and one's called decode. They have very different properties, and that's caused people to make specialized silicon for each of those phases. What's interesting about LLMs is you're trying to figure out, well, what is it that the user's asking? I'll show you an example in a minute. A lot of what's going on is you're trying to parse the language, trying to understand what the subject is, then trying to form a good response. More and more these days, those responses include things like looking around the web for some links and some information that can supplement your answer. It could also, in the case of people that use it to develop software, it could involve tool calls. You could be calling compilers, for example, or you could be calling APIs as you generate the code. In fact, you could be doing other browser operations as well. What the CPUs are doing is they're orchestrating everything that's going on, and they're doing a lot of the things that they've traditionally been doing to augment what the LLM is providing as an answer. The big upshot of where we are today is we've got specialized silicon, very high bandwidth platform, so very high bandwidth memories, high capacity memories, lots of parallelism and specialized silicon that are all coming together to really cause this interesting collaboration now between GPUs and CPUs, and that's part of what's making CPUs cool again. I wanted to show two more slides, sorry. One is, what is this LLM thing and how it works, and you'll begin to understand why it is that CPUs are in demand again. One of the fun things that people talk about is this thing called KV cache. What is a KV cache, right? Well, when you ask a question, like you have a user there saying, "Hey, are dogs mammals?" Right? This very instrumental structure called a KV cache is put to use. It's created and put to use. So during the prefill phase, what the LLM is doing is it's trying to understand, well, what are you even asking about? Something about dogs, something about mammals. It starts going to this, building this KV cache, where it puts relevant terms, relevant items to dogs and mammals together, then after it figures out what it is you're really asking, it starts to produce an answer. When you get up to the point where the KV cache is ready to go, and the LLM understands what it is you're asking, then it starts to generate a response, and that's shown in blue. That's called the decode phase. What's interesting is the prefill phase is very compute bound, but the decode phase, where we start to generate an answer, that's much more memory bandwidth bound. What's interesting is as each word or as each token in the response is generated, it's also added to the KV cache. The KV cache keeps growing, and that forms the context for your conversation. This is why LLMs can refer back to things that have been said earlier. It's because the KV cache is always growing during your session. The big emphasis has been, how do I support longer, more meaningful conversations? That means I need more capacity. The bandwidth, of course, is important to help me get the answer out, but the capacity's important because the contexts of our conversation are always growing. It turns out that this HBM memory, which is attached to GPUs, it's great. It's super high bandwidth. The issue is you just can't get very much capacity. Now what's happening is the KV cache is starting to spill over into another memory system. As I showed before, people are pairing GPUs with CPUs. Part of the function of the CPU, in addition to doing these tool calls and orchestration, is to hold a bigger portion of the KV cache. What we do is we move the parts we're currently working on into the HBM memory, and we take parts that are maybe a little bit older, and we move that out into the DDR memory of the CPU. Now you can actually see the two of them are working together, to try and support the larger KV cache that's needed in longer conversations. One last thing I wanted to show is kind of how to think about this from a memory hierarchy standpoint. There's all different kinds of memory that's available to processors. Right on the processor, there's usually some kind of caches that are made out of SRAM. Very, very small capacity, but wicked fast. Of course, we would love to put everything in cache. It's just in practice for most problems, you really can't do it. On the DRAM side, what we have available to us really are two kinds of DRAM. There's the HBM memory, which is sitting on package right next to the GPU. We see a lot of the evolution of this in our memory controller business. We sell a lot of HBM memory controllers. There's incredibly aggressive HBM device and memory system bandwidth roadmaps. You almost just can't supply enough. If we can get to infinite bandwidth, that'd be great, but until you get there, you got to keep working, right? There also is a demand for more capacity, but if you had to trade them off, you really can't live without the bandwidth. Bandwidth is always kind of biased if you had to pick one or the other. If you think about the DDR memory, which is part of the CPU, similar kinds of things, right? Again, it's used for all kinds of tasks, like orchestration, things like that, but also as what we call an offload, where the larger portions of the KV cache that aren't currently in use, that's where they're stored. We like to see higher capacity because we like to have longer conversations, bandwidth is really critical here as well. The key message here is that, we're seeing now GPUs and CPUs work together, which means the HBM and the DDR do have to work together, and we're seeing a demand for higher bandwidths all the way around. Right. As you're saying, as the HBM density grows, so does the DDR density. They have to go in tandem? That's right. Yeah. It's just that users are liking longer and longer conversations and, really they both have to grow in order for the whole system to move forward, like you mentioned. Yeah. It was interesting, another company we cover, Penguin Computing, introduced a KV cache system, an appliance that is 11 PB of a combination of, it's DRAM, but it's a combination of HBM and DDR, and I think it goes for $500,000 or so based on today's price of DRAM. Yeah, it goes to show you how vital the KV cache really is to LLMs and, with the specialized silicon being built, there are now solutions like this that can be applicable, maybe not for everybody, I'm not going to have it in my home anytime soon, but obviously it's something that can fill a need if there are people willing to pay for it. How about as we go to the more CPUs compared to GPUs, that's been part of the things that investors are hearing from both AMD and Intel saying that they're using more cores, but the ratio of You showed before, 8: 2 maybe, or 8: 1 for GPU to CPU, but as we go to more CPUs and the price and availability of DRAMs getting difficult, I get a question a lot from investors, what about de-speccing? What's happening to the servers? That plays right into Rambus, since you get a certain amount of $ per module, DRAM module? It's a great question. It is definitely the case that CPUs are becoming more used in modern AI. There's been this talk about maybe it'll move down to 1 to 1 or 2 to 1, something like that. I think school's still out. It's definitely moving in the direction of more CPUs, but there's a lot of questions, I think, on exactly how many CPU cores do I need to support a GPU user. That'll influence the raw number of CPUs a bit. I think we're definitely watching it and it's kind of hard to say exactly what it'll be, but it'll be interesting to see where it settles. This interesting thing about de-speccing, what DIMM modules are great at is they're modular. On the left, I'm showing a motherboard, and these black and blue striped things that are going vertically, those are the DIMM slots, and that's where you plug these memory modules. What makes DIMMs great is they're very serviceable, so if something goes wrong in the field, no problem. I just pop the old one out, pop another one in. Very upgradable, can add capacity, can fix errors, those kinds of things. The other thing that's really neat about them is, as a manufacturer, you can make a very late stage decision on how much memory to put into your box. If I had to solder it down to the motherboard, I have to make a much earlier decision. In times like these, where the prices are fluctuating and it's not exactly clear what the availability's going to be like, it's nice to be able to plug something in at the last second. What you see at the upper right here are two different modules. What's interesting is that one of them is 64 gigabytes of capacity, the other's 32 gigabytes capacity. They both run at the same speed, meaning they both provide the same amount of bandwidth. What that allows a systems person to do is to say, "Hey, I can optimize my cost a little bit by going with a lower capacity module, but I don't have to give up at all on the bandwidth. The performance of my system will stay roughly the same." It's a nice trade-off to be able to do that. There are other memories that do this as well. HBM does the same kind of thing. When people talk about de-speccing, really what it means is, I'm going to reduce the amount of total capacity to try and manage the dollars and the availability of the devices, but I'm really not going to give up on the bandwidth, so I'm still going to have very high-performing systems. Let me make sure we get that clear. There's two things about DRAM, the speed of how fast they spit the bits out to the CPU, and then how dense they are, how high up they are. That's right. I guess rather than putting, say, 64 GB modules in, only putting in, say, 4 GB instead of 8 GB, does that make sense, or is it that you want to get each one of those channels talks to the core CPUs inside the chip? Maybe. Yeah. That's right. What's interesting about it is, these CPUs have many, what are called channels. Think of them as like lanes on a freeway. If you were to take the CPU and hook up only half the number of modules, in some cases, it's like closing down half the lanes on a freeway, right? That's actually not as helpful to you. When you populate all the module slots but with 32 GB DIMMs, then all the lanes are open. It's just you're putting less capacity on each lane, but not changing the actual speed that cars are able to go down those lanes. From a bandwidth perspective, the ability to move data back and forth to the processor, it's much better to fill up every lane or basically have every channel populated with devices that can still achieve the top bandwidth. You're getting a full utilization out of the CPU. That's right. It's staying busy. You paid $10,000 for it. You want it to be busy. That's right. Yeah, the worst thing that could happen is for your CPU cores to sit idle because they're waiting for data. If you don't have the bandwidth, if you can't move that data in and out quickly, then yeah, you run that risk of low utilization and low performance. You did also ask this question about when you look at these modules, you can actually see they're constructed largely the same. The only difference really is that there's fewer DRAM devices between the two modules. They still have RCD devices, they still have SPD hubs and PMICs and things like that. From Rambus' perspective as a component supplier, yeah, per module, we still are selling the same kind of number of devices and things like that onto them. Right. Yeah, I think that's a very key point for Rambus' point of view, is that the number of modules is what matters to you, and as long as you're fully populating the modules in every server, you're going to see the benefits from that. That's right. What about, I get questions from investors often now that Arm and even Arm themselves have come out with their Arm AGI CPU, there's lots of different ASICs or ASIC CPUs that are based on Arm Graviton and others, and even Qualcomm. How does that change? Does Rambus care whether it's an Arm CPU or a X86 CPU? Well, really what tends to matter more is the use case. On this next slide here, you can actually see lots of different kinds of systems that use DIMMs. On the upper left is a PC gaming platform, so high-performance PC, and you can see on the bottom here, there's some DIMM slots. On the upper right, there's actually an Arm-based Ampere CPU server system that Gigabyte offers. Again, you can see, it's a little small, but you can see these horizontal areas on either side of the CPU socket, that's again DIMM slots. On the lower left, you can see a workstation. It's a little hard to see, but they're black. Basically, there's black DIMM slots here, then on the bottom right is an example of a supercomputer. This is the Cray Shasta supercomputer, which is part of the fastest U.S. supercomputer right now. It's a little hard to tell, but you can kind of see these vertical segments here that are on either side of these copper-colored areas. Those are DIMMs as well. In all cases, pretty much it's the same general idea of a DIMM. The real difference is exactly how many devices are on there and how fast they're running. In some cases, in, say, the consumer market, sometimes you sell DIMM modules that don't actually have an RCD, but they still have PMICs, and they still have other things as well. Largely, the DIMM solution, it's used very widely across lots of different market segments. You will either include some components or not, depending on the market. In the case of the PC market, there's a lot of unbuffered solutions, but they still need other things like power management and all that. Sometimes what'll happen is, you can actually see this module in the upper left here, it's got twice as many DRAMs as the one on the bottom here. That means it's going to consume more power. Sometimes, what you'll do is you'll take a PMIC that's a little bit different design. It's just to optimize, to provide more power versus less power, and then that actually is helpful from a power consumption standpoint. Largely it's the same solution. It's just little variations to optimize them for high volume segments where it makes sense to do that. There was one outlier, it goes back a couple of years ago when NVIDIA designed their, I think it was Grace Hopper. They used LP, low power DDR5 because power, of course, was a major issue. They didn't use any modules at all. They soldered the DRAM directly down on the motherboard. I guess I hadn't seen that in a long time. The DRAM industry has been around for 40+ years, and they moved to modules for a reason. Maybe NVIDIA made that decision, but in the next generation, they did switch to a module. It's a small outline CAM. They didn't have the buffer, the registered, the RCD or any of the massaging. That's changed now. I think it's like NVIDIA's evolving quickly on how to use a DRAM. Can you talk about your exposure now with the, now that's Rubin or even the Vera Rubin CPUs? Let me go back and talk about the decision to use CAM modules instead of soldered down. It really follows, some of the benefits of doing that are very similar to DIMMs. With a module, you can easily replace if something goes wrong. You can also make a late-stage decision on how much capacity to include into your system. If you're soldering it down, it has to happen much earlier in the whole process, and then you're fixed in your configuration. That's one of the big benefits of using something like these SOCAMM modules, right? For Rambus, we have recently introduced a SOCAMM2 set of chips. You can actually see it here. You can see that we're supporting the voltage supplies. We have different voltage regulators that go on the module, and also an SPD hub. It's just part of our chipset family. Again, they're very in line with the kinds of things we're already doing. They're, I guess in a lot of ways, variations of what we're already making. I think as AI starts to, as the number of platforms proliferate, and as you start to see AI move away from maybe even just the data center into other types of form factors, slowly these technologies will start to waterfall out, and it makes sense for us to start to take some of our technologies, go into adjacent markets, and service them that way. I think we've talked about this in past years just because it's the PMIC that you were making. Say there's a lot of companies that make voltage regulators. Why Rambus? Why are you making better voltage regulators for DRAM modules? As an engineer, you can get a set of specs and say, "Well, I've got to produce this voltage, and it's got to be this little noise," or whatever. A lot of the difficulty, you can see how packed in everything is on this module. It's not enough just to be able to produce the component. You have to make it immune to all the things that are going on around you, and you have to make it work and fit into that environment. One of the things about Rambus is we've got more than 30 years of experience on memory modules. We understand the environment, we understand the process of how to qualify memory, and we have great relationships with the memory manufacturers, and we have long existing relationships with processor and systems companies as well. I think if you think about everything that goes onto the module, really, a memory manufacturer, they've got the DRAMs, and they're trying to buy all these other components. It's very helpful if a company can, kind of under one roof, can supply all the technology and really understands the difficulty of how hard it is to get components into and make them reliable in those environments. Great. Just another topic, as long as we're talking about the shortage of DRAMs, too, in the past, we've talked about CXL and the various stages of CXL. What do you see as the adoption? I've had some investors say, "Well, with the shortage of DRAM, that more companies are going to CXL because they can mix and match all different types of DRAM that might be available, just so the system can have some type of storage. It may not be as fast, but it's there." What do you see happening in the CXL market? It's a great question. I think the interest in CXL is picking up a little bit again, simply for the reason you mentioned, which is, oh, maybe I can reuse some of my old DRAM in those systems. Really, its adoption has been complicated on the application side. It is a multi-tiered kind of memory, and I think some of the early studies really showed, look, things are going to have to change in the infrastructure in order to really make this useful. Some of those changes are pretty complicated. I would say that right now, that imbalance in the supply and demand is definitely reviving interest, but I think it's also still pretty early. It's very hard for the application guys to really make use of the architecture as it is. There's a lot of ongoing work right now to try and figure out both how to change applications and then what's really going to be needed as they learn more. Maybe we need to put something into the next spec for CXL. Right now, I'd say it's early days. Some of it is driven by the market dynamics that are going on. Certainly, some of the revival of interest is from the market dynamics. If those market dynamics remain, then it's possible we may see more of an interest in CXL. For right now, I think it's still early days. Well, last week we had COMPUTEX in Taipei, maybe if we were talking about that CPUs are cool, now PCs are cool again. You showed an example of a gaming PC, very high-end. As agentic AI starts moving to our PC platforms, the new processors are going faster and faster, I know a couple of weeks ago, Rambus announced a chipset for the PC module. Can you talk about that market and how you see that developing? Yeah. Let's see. I think I've got a slide back here for some of the PCs like LPCAMM2 and then the CSODIMM or the CUDIMMs. That market is important, I think as we watch AI move out of the data center into more towards the edge and the endpoints, there's this great established market of laptops and PCs in people's homes, right? It makes a lot of sense to think about, okay, well, why don't we move AI there? By the way, that platform that I showed, that actually is the one I have at home. It's a great board, would recommend, so if anybody wants to buy one, it's a really nice board. Part of why I got it was because it has all the right kind of ability to support things like high performance GPUs and all that for AI. We're expecting, again, that things will waterfall out of the data center and move more into PCs, into clients. For us, there are various things that we've introduced. We have LPCAMM2, the memory module chipset. Again, you can see here the chips that we provide. There are SPD hubs and a PMIC as well. Again, kind of variations of some of the things we've already got. The PMIC is more optimized to the power and form factor of the LPCAMM2. Then for the CSODIMM, CUDIMMs, again, we've got SPD hub, a PMIC that's a little bit different, and a client clock driver here as well. They're really, again, I think variations of some of the things we were already doing. It just made sense because we can kind of see as AI's evolving and moving out of the data center, this is a place it's going to go. Okay. Maybe if I can ask Matt from a marketing point of view, when do you see this being significant revenue for Rambus? Yeah. As Steve said, we're seeing the waterfall happen. It's been a big moment for us, if you will, to transition into the client market. The speed grades here that require things like the client clock driver, that signal integrity chip to reach those speeds are at the very highest end of some of the platforms that are coming out. If you look at Intel's Panther Lake, for example, or Arrow Lake, excuse me, it was at the top end of the speed grade. It's Panther Lake, it starts to notch down to the top two speed grades. You see it, as Steve said, continuing to waterfall down to the mainstream. A little bit offsetting today is the price of memory, PCs a little bit more acutely feeling that than maybe the data center. It'll be some time before this becomes a mainstream and a real revenue driver for us. We're here, we're present, we're building this forward, and it's important for us to see this evolution. It's coming, but I wouldn't put in your model, Kevin, for this year just now. Okay. Will do. Also during COMPUTEX, Jensen announced their CPU chip, his presentation was more around the PC isn't going to be just waiting for you to open it up and start working. It's going to be your R2-D2, your personal robot. It just made me think more of like, well, robots seems to be where everyone is talking about for the next five years of fast development. What would be needed in a robot, or maybe as we get out to physical AI? We just did a panel discussion on physical AI, it's a trend. What happens to the DRAM configurations as we move out with the CPUs that are out on the edge like that? Yeah. I think, obviously, the systems are going to become more capable than they are today, that tends to drive more memory capacity and the need for more memory bandwidth. What we see is in things like some historical examples, things like cars or some cell phones became more capable, they needed more capacity and more bandwidth. What they also did, though, was they always maintained a connection back to the data center. Even today with self-driving cars or sorry, with navigation, where we're using maps and things like that to kind of find our path, we're communicating back with data centers. It's keeping track of all the information on, say, how congested the roads are and all that. There is this kind of coordination between the endpoint devices and the data centers. I think initially for robotics and autonomous devices, I think many of the devices are going to maintain that type of thing. You're going to have to have some type of connectivity, that's going to help alleviate, so that you don't have to do everything on the device. I think over time, I'm just as interested in getting Rosie the Robot in my home. I mean, I'd love to see that, right? I also think it's going to take time. The technology's got to catch up. We need a few learning cycles, we're on the path, which is the good thing. I think in the end, it will inevitably cause the demand for more memory and the demand for more bandwidth. I'm just as interested to watch the market develop, I think, as everyone else. Yeah. Maybe just a couple of other questions that I get from investors a lot, and by the way, if anyone wants to ask a question, just type it in. We've talked about it in the past presentations, of the memory wall, of the CPUs go faster than you can feed memory into them. There have been Celestial AI, as an example, was a company that Marvell recently purchased and is saying, "We're going to do optical connections or make a optical memory appliance." What's your view of that, and does that still need your chipsets in an optical memory appliance? Yeah. The idea is if you can have this optical connection, it sort of does a couple of things for you. One is it does have higher bandwidth, right? The second is that it allows you to take a larger collection of devices and then kind of multiplex all their data onto a smaller number of optical fibers to move the data. The two things I'd say, one is, it's something that I think we see this in data centers for long distance communication. There've been a lot of companies working to try and cut down the distance over which that makes sense economically and from a reliability standpoint. At some point, you still have electrical signals that you have to deal with. The optical signals eventually get converted back to electrical, and then, of course, the electrical signals on the memory modules would be converted to optical to go across these links. I don't think the fundamental problems that we solve with our technologies, I don't think they go away. I think there's just some interface that after you get off the modules, maybe at the appliance level, then you put an optical interconnect, and that allows you to go over a longer distance. Many of the fundamental electrical challenges, I don't see them really going away as we go to these kind of optical interconnects. Okay. A pure signal is needed no matter what. Yeah. I mean, the world has invested for decades in digital signals and silicon, copper wires, that's unlikely to go away anytime soon. Then, one last topic that I'd like to cover is, last year we talked about MRDIMMs, and these were when you mentioned multiplex, the DIMMs glued together as you described it. What's the adoption rate? What's happening with that market? Is it very specialized, very high-end, or is it, again, availability of DRAM slowing that market? I don't know, Matt, did you want to talk about it or do you want me to? Yeah, sure. Kevin, I think we're still in the early days there. One of the key enablers of the MRDIMM technology is going to be the platforms that it goes along with, and those are working from code names for AMD. It's the Venice CPU and the platforms associated with that. With Intel, it will be their Diamond Rapids platform. We're still a ways away from seeing those servers based on those CPUs come to the market. That'll be a critical step one. The bandwidth and the capacity that this unlocks, that memory wall you were talking about, it's a step in the right direction in an elegant solution, using all the same infrastructure and memory chips that exist today to unlock, double the bandwidth, double the data rate, double the capacity in this existing technology today. It remains to be seen, things like AI, as we talk about inference, we talk about agentic use at the top of the discussion here. Those are things that are thirsting for both bandwidth, as Steve was talking about, with KV cache sizes growing, as well as places that we see this continuing to adopt. Some of the mitigators here is certainly pricing of memory. There was a discussion about de-speccing here. This is upspeccing memory in a way. I think it's a technology that I think will play out and have a great home here over time. Like new technology, it's going to take a while to take hold, and we're looking forward to the first servers rolling out, again, based on that Venice CPU, hopefully by the end of this year. Okay, great. I'll do one more poll for questions from the audience. I've left an extra five minutes for questions, but if we don't get any, I want to thank Steve and Matt very much for coming again this year. Very informative. We now know what KV caches are used for, and we also understand why CPUs, when they despec, it's not the module, it's the density of the module. Is that correct? Yep. That's exactly right. Okay, great. Well, thank you for your time. Very much appreciate it. Thanks for having us. Yeah. Thanks, Kevin, as always.
Loading workspace