Welcome to INSIGHT 2026. Thank you to our sponsors, our partners, and so many who have worked so hard to make this a fantastic event. Thank you for traveling and taking the time to be with us here today. We are so excited about what we will cover over the next two days. Between our keynotes, our featured and breakout sessions, you leave having a deep understanding of our perspective on the evolution of the data center, data infrastructure, and the data stack over the next few years, the massive AI opportunity, and the specific challenges facing organizations who want to capture it, and what we are doing at NetApp to solve those challenges and to help you meet this moment. While much of the AI conversation has evolved, shifting focus from discussions around models to GPUs and AI factories to inference performance and token economics, what has not changed is the critical importance of your data and your data infrastructure in driving success for your AI strategy. Many organizations are still struggling to scale AI from pilot to production. When we sit down and problem solve with you, it always comes down to two fundamental challenges. Challenges with your infrastructure and challenges with harnessing your data. Everything we will discuss today is tied to addressing those two problems. Let me begin with infrastructure. With the advent of accelerated computing and AI applications, data center infrastructure is being transformed, the old design patterns being swept aside, making way for new paradigms. To reduce the time it takes to train and fine-tune your models, you need to pool enormously large GPU clusters together. Typically, the larger the cluster, the faster you can train the model. GPU compute is unforgiving. With high concurrency and parallelism, bulk data reads to load, bulk data writes to checkpoint, interleaved metadata and data operations. You are measured unforgivingly on time to first token and cost per token, and it is extremely expensive to have any of it sit idle. You also need to organize very large exabyte scale or even zettabyte scale data sets with different data types that span your hybrid cloud under one namespace. Because, simply put, larger, higher quality data sets generally improve AI inference quality by giving models more diverse examples to recognize patterns accurately for inference. The scale of both the compute clusters and data sets is accelerating at an extraordinary pace. As an example, large language model scale has grown 5,000-fold in four years. This can be particularly frustrating if you are a model builder or trying to fine-tune models, if you are operating AI hyperscale clouds, GPU clouds, or sovereign AI factories. But also, if you are an end user of these factories or building your own AI platform. The second challenge is your data. Data is becoming the constraint to productively scale AI. You know what? At NetApp, we know a few things about data. For over 30 years, data infrastructure has always been and remains our focus, and we have been building and innovating solutions with industry leaders like AWS to help businesses take advantage of their data. Now it matters more than ever. You see, 93% of organizations report that at least one form of data challenge is blocking their AI ambitions, and the high-performing ones are already running into this. The challenge becomes even more critical as you scale AI into more business-impacting opportunities, which carry more risk because AI retrieves, recombines, and generates data faster than traditional governance models can keep up. To compare against these needs, let's look at your data estate today. You see, before an AI system can generate a single useful token, it must discover, access, and understand the correct information required to answer the question it is being asked. In a modern enterprise, that information rarely lives in one place. A small part of it, less than 15%, typically resides in structured business systems such as your ERP system, your CRM system. But these provide just a clinical view of your business. Much of the context, the meaning, the sentiment, the texture of your organization that makes AI valuable, about 85% of it, is somewhere else entirely. Engineering documentation, research data, operational files, media assets, simulation output, scientific data sets, and decades of institutional knowledge placed in content repositories, email systems, Slack channels, and Zoom calls. All of it distributed as unstructured data across siloed storage and application environments, cloud environments, remote offices, and different data sovereignty jurisdictions. So what's the cost of all that fragmentation? First, foregone opportunity. Petabytes of your unstructured data are dark, with hidden knowledge ignored because it sits in isolated, undiscovered silos. Second, operational complexity and fragile business continuity. Different customers, workloads, and data sets have genuinely different protection requirements. However, having customized, tailor-made infrastructure silos requires separate maintenance schemes, massive operational complexity, and never-ending technical debt. I'm sure you recognize this challenge. Third, ransomware and cybersecurity attacks are on the rise, and agentic AI-enabled ransomware attacks are already enabling bad actors to attack vulnerabilities faster than ever before. Finally, sovereignty restrictions are beginning to be a struggle to implement consistently across your data estate. To summarize, fragmented data, dark data, uneven protection, facing rising attacks, and hardening sovereignty rules. This is the state of most customers' data estates. No wonder it's hard to use AI with your data and data infrastructure in this state. Let's talk about the business impact of doing so. Businesses today are more global in scope than ever before. Customers, suppliers, and employees can be spread across the world, and they work closely together to share information, learn from each other's experience, and act on the organization's behalf. With the advancement of AI, leaders across every industry now envision the agentic enterprise, where agents extend their employees' capabilities. These agents are autonomous software programs that use AI to perceive their environment, make decisions, and take multi-step actions to achieve a goal. To make them work for you, they need two capabilities that let humans and organizations operate effectively in the first place. Memory and context. Let's start with memory. As humans, our memories shape who we are. They make up our internal biographies. They tell us what we've done with our lives, who we are connected to, and inform what we should do next. Now, on to context. Context is the background of an event, a statement, or a piece of data. It is what gives facts a clear meaning. Without context, a single word, number, or action can mean many different things. Context helps memory recover the right experience to make the right decision, and it updates that memory with recent experience. For example, you might have a memory of yourself as a child on a boat with your parents. Without context, that's all it is. But with context, you'll remember that it was a special occasion because your dad took off work that day to be with you. Your mother packed your favorite lunch, and you were wearing a brand-new shirt that day that you felt great in, and that you might still have hanging in your closet today. Together, memory and context create feeling, impact, and when combined with actions that provide for continuous learning, memory and context are what create knowledge. For your enterprise to be truly AI-enabled, it needs context and memory for your agents. To give them that, you need a data platform that not only can store and protect your data, but can also unlock the insights that are embedded but hidden in your data. One that has been trusted to store your organization's data, your digital memory, for decades, and that can now organize it and build in context right where the data is created and where it exists to transform your data into knowledge. That's the NetApp platform, an intelligent governed infrastructure that provides zero copy access to all your data for any data and any workload anywhere you operate. The NetApp platform is already part of your enterprise. You've trusted us with your data, your organization's memory, for more than three decades, and you can continue to trust us that we will provide you with the tools to turn your data into knowledge. Today, we will share two major advancements to address those challenges that I spoke about earlier, infrastructure and data, so that you can lead in the AI era, scale AI across your organization, and take advantage of all the investments you've already made with the NetApp platform. The first extends the value of the trusted NetApp platform to meet and exceed the infrastructure demands of the largest computing infrastructure, the biggest frontier models, AI factories, and hyperscalers. This allows you to solve the infrastructure challenge we spoke of and use the world's most advanced capabilities to process your data, your memories. The second, to build context. We are going to provide you with the capabilities to exploit the value of all of your NetApp data, where it is created, and where it resides by organizing it, providing context, coherence, and the data foundation you need for your agentic enterprise. The NetApp platform lets your organization do four things. One, make your data AI-ready so that you can move quickly from idea to impact. Two, unify your storage, confidently handling the most demanding workloads wherever they live. Three, proactively protect your data with unquestionably the most secure storage on the planet, period. And four, give you complete control over your entire data estate and data infrastructure. No silos, total visibility. The NetApp platform helps you turn your data into knowledge. To provide you with a deeper understanding of the NetApp platform and some of the new innovations, I want to introduce our Chief Product Officer, Syam Nair, who will be on the stage shortly. Before we bring up Syam, it is always worth remembering that the secret to the NetApp platform is not only the awesome innovations that we built, but because we have as our strategy and our tradition the idea that when we innovate together, we innovate best. We have done so with a broad ecosystem of technology leaders across the industry. I am excited to announce the newly expanded partnership with a leader in cloud AI, leading soon to an awesome new product announcement. Let us hear directly from them, virtually welcoming Mahesh Thiagarajan, Executive Vice President of Oracle Cloud Infrastructure at Oracle. George, thank you for the opportunity to join you and the NetApp community at INSIGHT. Getting real value from AI depends on making enterprise data accessible, manageable, and ready to use. NetApp has earned customers' trust over the years by delivering powerful and consistent data management across on-premises and cloud environments. For NetApp customers, moving critical applications and data to the cloud should not mean retraining teams, redesigning workflows, or starting their data protection strategy from scratch. This is why Oracle is very excited to partner with NetApp to bring ONTAP on OCI. Very soon, OCI File Storage and NetApp ONTAP will enable customers to run business-critical file storage workloads in OCI using the same set of ONTAP tools, API, and automation. Together, OCI and NetApp will help customers extend applications and their data to the cloud with greater operational consistency, separating and supporting large data sets, enterprise applications, development environments, and a very resilient hybrid cloud strategy. We are very excited about what our teams can help our customers achieve together. Super excited about the partnership. Please welcome Syam Nair. Good morning, everyone. Good morning. Thank you, George. Thank you, Mahesh. We are thrilled to partner with you and the Oracle team to build the next generation of storage, NetApp Storage Service on Oracle Cloud Infrastructure. We will talk more about this throughout the breakout sessions. Please visit the expo floor. There will be featured sessions where we will talk about this new offering. For 30+ years, our industry answered one question: Where do I put my data? The answer was a scalable, durable storage layer. Store it safely, store it securely. We became the most trusted place to hold your intellectual property, your most precious asset, data. Everybody talks about it, the real fuel of the enterprise. With your trust and our continued innovation, we are still very good at it. A few months ago, I sat down with the CIO of a large European financial institution. He said something that stayed with me since. "Everyone sells me their new features, performance, durability, maybe a new AI operating system. What's actually on my mind is driving value from the investments that I've already made. I need a partner who helps me get more from the data that we already have." He said he felt for the first time, someone really answered that call. NetApp's mission is to supercharge what you all have already invested in. What the CIO was asking: What does my data know? What can it do for him? He is focused on the agentic enterprise, where agents take on complex actions on your behalf, whether it's moving money or making trades. George talked about the importance of memory, context, and knowledge, what agents need to deliver business outcomes and to be trusted. Let me show you where that memory and context live and how you can enable it on the infrastructure that you already have. Look, we need to solve for the infrastructure needs and the data needs together, not in silos. That's the need for the agentic enterprise. Today, we will talk about the only platform built to do exactly that, make your entire data infrastructure ready for the agentic enterprise. One thing before we go further. Remember the CIO conversation? The CIO wanted more value from his infrastructure, and I'm holding myself to the bar right now with all of you. Everything today, literally everything is either generally available this morning or in preview, and you can sign up in the expo hall this afternoon. All of you. Get to supercharge your investments right away. That's the promise from NetApp. The NetApp platform has four capabilities: AI-ready data, proactive protection, unified storage, and complete control. Every one of them is anchored on the same thing, your data, where it already is, made usable without any migration. That's the promise of true zero-copy activation with always-on protection. With the NetApp platform, it's trusted, intelligent data infrastructure for the agentic enterprise you all deserve. Let's start with AI-ready data. Agents that act on your behalf are only as good as the data they can reach, trust, and act on safely. Data fragmentation is a reality across on-premises, cloud, and the edge. When this happens, it's natural to have increased operational complexity and inconsistent governance. The first thing that must be true is your data. It must be ready. That means an agent can find it, understand it, trust it, and act on it right where the data lives, without added complexity and no runaway cost. Here's the part that changes everything. You all already own this data. You have this infrastructure. It's your intellectual property. Every copy somebody is asking you to make is a governance problem. Every move is an operational cost. What do we do at NetApp? We help you activate your data right where it already sits on the infrastructure you already trust. We are igniting what you already own everywhere where your data is. When I met with the head of a research computing company, they have decades of trial data, lab results, research nodes, sitting across systems going back years, some of it older than the researchers now trying to use it. This is digitized memory and context hidden in there for them. But to unlock it, 6+ months of data engineering before a single model sees a single file. They needed the data they already owned to speak to the AI they are trying to use. This is not very different from the CIO conversation, just a different application of the same ask. What does my data know? Activate the data in place safely and securely. We can do this today. Today, we can activate with context exactly where decades of trust already put the data. Let me show you how it fits together. The foundation of it is about discovering and understanding what you have, what data is there without moving it. Then comes insights. Surface what's actually in the data, the patterns, the context that was sitting there unused. Built-in intelligence, governing, securing all of the data automatically with the same policies protecting your data today. And finally, context. Now your agents have something to reason over, built entirely from data that never, ever left your infrastructure. We bring it to life using the NetApp data platform. How? By tracking your data and metadata without constant discovery, without scanning your file systems again and again, with no migration. George mentioned this. Let's go back to the need for memory and context, combine that to create knowledge. We talked about how your data is everywhere, but what has been missing is the layer that knows what is in the data, where does the data live, what does it mean. The layer that turns a data estate into the memory and context, and the layer is right there, lives on the same trusted, unified storage you already use. You already use this for your most critical data and your business critical applications. The most secure storage on the planet. Oh, sorry, the only trusted data infrastructure you will ever need, and that is from NetApp. Look, this is not about just activating data. You have invested in the infrastructure. We need to activate it with efficiency, with accelerated computing and GPU acceleration, better economics on the infrastructure you already trust. That's the outcome you need from an intelligent data infrastructure. Last year, we introduced a NetApp AI Data Engine, and as customers used it, we learned what the biggest job really was. Safely activate data without copying it, without losing the governance. This AI Data Engine is part of the platform, purpose-built to turn all of your dark data in your trusted, governed context. Your data made AI-ready right there. And while it's activating it's also standing guard over it. I'll come back to this, exactly what this means in a little bit. For now, just remember this. Readiness and protection aren't two separate things. They have to happen together. And now that the part I'm most excited about. We talked about the researcher. Their problem was cost and complexity, 6+ months of engineering. Their CIO told me that not all of the data was on the same storage platform, so they had to discover, ETL, harmonize, model, and make them available to the researchers. Well, we have a solution for that. Everything I just showed you up until today worked on NetApp data. I am excited to announce that the metadata capability of the NetApp AI Data Engine supports data from anywhere. Heterogeneous data sources, all of it. It is not a new system you need to buy. It is the fastest way to find out what is actually in the data you already own. The AI Data Engine is now fully software-defined, discovers, understands, and activates data we have never touched. Point it to your entire data estate, every system you have inherited along the way, and start building context across data without moving them into another system. This is the magic. I have a call to action. Thank you. I have a call to action about the metadata engine. Everyone, every enterprise in this room should be running it. Not eventually, not on a roadmap. As soon as you are all back in your offices. To help you all do that, I am happy to announce we are making NetApp AIDE, the AI Data Engine, and this new metadata capability I talked about available to all of you here today and every other enterprise out there, free for the first six months. Take out your phones, scan the QR code, and they will be available in the expo floor if you do not get a chance. Look, our promise is very simple. Your data activated wherever it already lives, and we mean that literally. The data is ready, but every training run, every large query still ties up computer resources. Everyone you talk to these days talks about accelerated computing. Let us talk about the researchers. They do not want a new tool. They live in Spark. They live in Trino. They get, what, decades of Python libraries that they trust. These are engineers. They were talking about true accelerated computing that they need, that we need to build into the data infrastructure with no tool change, no code change. That is exactly why we acquired the company DataPelago. DataPelago accelerates the tools the researchers already use, Spark, Trino, Python libraries, built over years, right in the data infrastructure. We talked about zero copy activation. This is zero code change acceleration. In the agentic era, readiness and protection are in two separate conversation. If your data is not protected, your data is never ready. An attack is not a matter of if, it is a matter of when. That was already true. Now your agents are acting on your data autonomously, making decisions, moving money, triggering transaction, all at machine speed. If that data is compromised and nobody catches it in time, the agent does not slow down to ask if it should trust it. It just acts. So your data infrastructure is no longer the last line of defense. It is the thing your agent's judgment depends on each and every time they act. A colleague of mine, a cybersecurity leader, would often say, "If you are reachable, you are breachable." I think the same is true for every piece of data now, maybe a little bit more urgently. If it is valuable, it is a target, and if it is a target, it needs to protect itself. Here is why I believe this even more strongly. Very recently, some of the best AI models in the world could not pass a hard security test. You probably heard about it. They were working in a sandbox, but they cheated and finally broke into a production infrastructure at a real company. These agents were not malicious, but their actions were risky. The guardrails, yes, they had them, but you cannot hand an agent a policy and expect it to just care. I am talking about the most secure agents operating on your data and where we need a new meaning of zero trust, which is not a network posture. It is not an identity check at the door. It is protection built into the data itself, every single piece of it. NetApp is on a path as part of a platform to define and deliver zero trust data security built into the platform. We will work with our partners, the best cybersecurity systems, as well as other providers, and make sure that there is ability for every piece of data to protect itself. Real AI readiness and always-on data security were never two separate guarantees. Data that is unprotected was never actually AI ready. They have to be part of the same platform promise, and we will deliver that with NetApp platform. Now, detection without recovery, it is just an expensive alarm. It tells you something went wrong. It doesn't help you get back to work. Zero trust means every access point gets verified before it happens. But verification doesn't help if someone still gets through. So your recovery point has to be immutable. Nobody can alter it or delete it, even with the right credentials. And it has to stay safe even if your network is a thing that got hit, because it has to run air gapped, cut off from whatever failed. It is not about detecting ransomware. It is about closing the loop when something goes wrong and getting you back in business, the real recovery, protecting your business always. And it has to run itself and be part of the platform DNA. With the NetApp platform, it catches problems before they spread and fixes them automatically so that no one is waiting on an approval at 2:00 A.M. Our value is now going to be built deeper into the Commvault product. So your team knows from the console that they already use. With this new integration, potential ransomware attacks light up within your Commvault interface, and you can see all the threats collected together. We have now also integrated the complete curated clean recovery in with Commvault, so you can leverage tools like SnapRestore and recover as easily as you want to. No new tool and fewer tools also mean fewer doors for the attacker to find. The real cost of a breach is never just downtime. It is the time lost getting back to business. It is time lost rebuilding trust with regulators and customers who are certainly asking us a lot of questions. The real value here is the day something goes wrong, you can depend on the platform for business continuity, and you are already planned for that and accounted for. I want to thank the Commvault team for the partnership, and we are not stopping here. We continue to enable our partners to always leverage the best of the NetApp platform. As another example, let's hear from Thomas Kurian of Google Cloud, one of our longtime partners, discussing how they are leveraging our latest innovations. Thank you and hello, everyone. It's great to be back at NetApp INSIGHT with you. Google Cloud and NetApp are continuing their long partnership to help customers solve critical data challenges in this AI era. With Google Cloud NetApp Volumes, we power the largest data center exits across all enterprise workloads, and we're doing that without compromising on performance, scale, or security. By leveraging ONTAP's latest innovation alongside Google Cloud's Compute and Hyperdisk, we deliver massive scale performance with up to 200 GB per second of throughput to support demanding EDA workloads. All of this is backed by best-in-class security, including NetApp Ransomware Resilience. NetApp Volumes is also enabling customers to deploy advanced agentic workloads and RAG inferencing. By bringing unstructured data into Google Cloud, customers can now transform their data from on-premise systems and GCNV into a hybrid grounding knowledge base. By combining Gemini Enterprise with NetApp Volumes, we deliver trusted AI intelligence without any infrastructure friction. Customers including PayPal, Marvell, and HCL Technologies are already leveraging Google Cloud NetApp Volumes to power their enterprise workloads. Additionally, NetApp storage solutions are now embedded natively into Google Distributed Cloud. This is empowering regulated sectors to run our AI model, Gemini and Gemini Enterprise, in air-gapped sovereign environments. Alongside NetApp, we enable you to unlock your data's potential and to drive faster AI innovation. Thank you for your continued partnership. Thank you again, Thomas and the Google team, for the partnership. Now, a different focus area for us. If you're running a hybrid estate, and almost all of you are, all of us are, this one is for us. Let me tell you about the manageability challenge that I often hear about. The problem is often simpler but somehow harder. Infrastructure in four different places, four different teams, four different consoles. Admins can't tell the confidence with whether the same policy is enforced everywhere all at once. We need to allow to manage it any way you want it, anywhere, and that's what complete control means. One place to see, manage, and govern everything on-premises, every cloud, everywhere in between, as one data estate. Our answer to that is the NetApp Console. On-premises, Keystone, all major public clouds, all on one canvas. AI-assisted autonomous operations so that there is one control plane for the entire platform. Predicting issues, resolving incidents faster, optimizing capacity and cost with human oversight always in the loop. For the most sensitive environments, the ones that cannot depend on a network connection to stay secure, or where regulatory or enterprise policy needs it to be local and air-gapped, the console runs within your environment, completely air-gapped, with the same visibility, same control, completely isolated when isolation is the requirement. With this, four consoles and several other tools all now become one. Built for every environment, every policy, one place, and built into the same platform you trust your mission-critical businesses with. I talked about this local air-gap control. Let me show you exactly what that looks in practice. This is fleet management capability of the Console local. It goes beyond what any other tool ever tried to do, organizing and operating an entire data estate as one from inside your data center. Intelligent fleet visibility built in, health monitoring, predictive insight. Again, all running on your local infrastructure. For existing AIQ Unified Manager customers, this is not a rip and replace. It is an upgrade path. The same investment now delivering complete unified control. No more manual operations stitched together across tools. One control plane doing what it used to take with five. Let us take it a step further. What if you want your entire environment built out as a service, but you need the same service provided in a sovereign manner? George talked about this. The sovereignty needs, DORA, data residency, a growing list of regional requirements. Today, we are unveiling Keystone Sovereign. This is Keystone subscriptions with data sovereignty controls to meet the strict regional compliance and data residency requirements of European Economic Area countries. This is a sovereign-focused Keystone offering, an add-on entitlement that enables EEA customers to purchase Keystone with additional regional operational support and data handling controls specific to EEA. For customers, it ensures data as well as support remain within the same specific geographic boundaries and meet the regulatory requirements. This is available early November for France and Germany, with general availability for everyone else later on. In both these previous demos, I showed us leveraging the console as we always traditionally have. But George also told you that enterprise leaders everywhere are envisioning an agentic enterprise, where agents extend what people can do. That is just not true for your business applications, it is true for your infrastructure, too. Now, the console is not something to look at anymore. It is something you talk to. The console now understands intent, not just commands. Ask it to provision a volume, set up a replication, investigate why a workload slowed down in plain language, and it acts inside the guardrails you have already defined across ONTAP, StorageGRID, and E-Series. Because it works with the LLM you already run, your own governance travels with it. The same agents helping your businesses move faster now help your infrastructure move just as fast. Again, one team, human and agents working together, running the whole estate. Every day, we are working with the largest industry movers to help our customers more simply manage their entire hybrid multi-cloud environment. To learn more, let us hear from Scott Guthrie, Executive Vice President of Microsoft Cloud. Hi, I'm Scott Guthrie, Executive Vice President of the Cloud and AI Group at Microsoft. It's great to be with you today. NetApp INSIGHT 2026 brings together the people helping organizations navigate some of the biggest technology shifts we're seeing today. When you look at cloud, AI, cybersecurity, and digital transformation, so much of it comes back to one thing, which is data. Customers across every industry are looking to move faster with cloud and AI while continuing to leverage the applications, infrastructure, and data that already run their businesses. They don't want to start over. They want to build on what they have while taking advantage of what's possible with new technology. That's where strong partnerships become especially important. Together, Microsoft and NetApp are helping customers bring ONTAP workloads to Azure and simplify hybrid environments. That gives customers greater flexibility as they evolve their IT strategies. We're also continuing to innovate through Azure NetApp Files, helping customers run some of their most demanding workloads in Azure, including semiconductor design, chip verification, and large-scale engineering simulations where performance and reliability matters. As organizations scale AI, the ability to access the right data across environments becomes increasingly important. By combining NetApp's intelligent data infrastructure with Microsoft's cloud and AI platforms, we're helping customers put their data to work in new ways while maintaining the governance and control they require. That's what makes this partnership so valuable. Customers can move forward with confidence, knowing they can build on the foundation that they have today while creating a stronger one for what's next. I'm excited about what we're continuing to build together and the opportunities ahead for our shared customers. Thank you, Scott. We are excited to continue our partnership with Microsoft. All these partnerships tell us that NetApp platform doesn't have borders. It works on every hyperscaler, on your premises, as well as on the neoclouds. We have talked about how we protect your data, how we make your data AI-ready, and how we help you manage all of that. Now let's talk about where all that data is stored. Let's talk about the story of a global retailer. Every November, the traffic multiplies overnight. For years, that meant scrambling to provision extra capacity, then tearing it back down, helping two environments stay in sync the whole time. A big challenge. They did not need a big data center. They needed one platform that had the elasticity to stretch before November to meet the scale needs. That's the true meaning of a unified storage and a unified platform. One platform, one global namespace across every environment your data lives in on-premises, every major hyperscaler, or anywhere in between. It expands when you need it, contracts when you don't, without you managing three different infrastructures to make that happen. Underneath that, it's ONTAP. ONTAP does all the work. The same foundation that's run your most demanding workloads for three decades, and the reason we call this unified in the first place. AFF is our flagship for on-premises performance. Because it's built on the same foundation, not just for high-performance files or high-performance object, it's truly high-performance NFS, SAN, and S3, all natively, all at once. That means unified, one system speaking every protocol, every application that you have, it speaks to it. AFX extends the same foundation where you need the scale, performance, and capacity separately at massive scale for files and objects. The cloud is where this really opens up. The same unified storage ships first party natively inside every major hyperscaler. This is not a virtual appliance bolted on. It is an integrated service powered by the same foundation, which means you get the same AI readiness, the same protection we just talked about, wherever it is running. Same platform, same data services, wherever you need it. For workloads that need something even more specialized, the platform brings that, too. StorageGRID gives you best-in-class object storage at massive scale, powering enterprise data lakes. Where a workload genuinely needs Lustre, that extreme purpose-built for HPC environment, that is in the platform, too. You are not going to find another vendor for that. It is already there in the platform. Today, AFF and AFX are running enterprise high-performance workloads and powering the AI and enterprise agentic transformation. They are extending the infrastructure customers like you already trust instead of replacing it. These are the leading systems for larger enterprises pursuing agentic AI today. Remember the retailer? They wanted infrastructure that could keep up with extreme scale when needed. That is the answer. The same platform you already trust, fast enough to keep up with whatever you are asking of it. Think about everyone I talked to you about today. CIO, who got the best out of everything they already invested in. Researcher, who finally got the data working for them. Retailer, whose infrastructure now stretches with them instead of against them. Teams who finally have one console, one view, instead of four. That is the platform. That is what you can turn on this week. As workloads move from enterprise scale to hyperscale into AI factories, neoclouds pushing GPU demand further than anything we have ever built for before, this same unified layer must further go. We have got such a big announcement to make here. I think there is only one person who can do it justice. Let me welcome back George Kurian. Let us do this, George. Thank you, Syam. Today, we want to share with you the most important announcement we have made in a long time. Certainly the most important announcement of INSIGHT 2026. The most important announcement since we announced Data Fabric in 2013, and our first hyperscaler native storage service in 2017. Like Syam mentioned, unlike others who ask you to rip apart your prior investments each time the world changes, NetApp extends the NetApp data platform and its technology capabilities to meet the moment. Today, we are in the accelerated computing era, and as we discussed earlier, this era brings new demands and new opportunities for infrastructure and system design and new technologies that are available today. We are introducing a new architecture, one that is disaggregated intelligently so that every single element can be optimized for performance and scaled independently. One that is architected for the performance needs of the 100,000+ GPU clusters, the unpredictable performance and concurrency requirements of the most demanding agentic workloads, and which can assimilate data across tens of thousands of systems into a single namespace to help you turn all of that data into knowledge. The possibilities that that would create are endless. Scientific research could accelerate as researchers run complex simulations and models in real time. Biologists could harness tens of thousands of computers to watch every protein molecule in a cell, and you could have breakthroughs that are not possible today in climate science, in genomics, in healthcare, and astrophysics. Now I ask you to imagine what could be the possibilities for your organization. If you could have a single file system that doesn't just match the highest performance numbers on the IO 500, the benchmark for super computing performance in the world, which is 10 TB per second, but one that is 10 x larger. One that is able to deliver the first file system in the planet to be able to deliver 100 TB per second of performance. That's right. 100 TB per second of performance, the fastest storage on the planet, period. In addition to performance, if we could unify all of the world's data, not just a few exabytes or a few tens of exabytes that some other architectures might be able to deliver. But imagine if you could unify all of the world's data center capacity that shipped last year, 2,000 EB of data into a single namespace, the first of its kind. That's right. The first zettabyte scale file system on the planet. Let that sink in. If all of this could be accomplished with no custom clients or hacked kernel loadable modules, but rather uses industry-standard Linux so that you can rapidly adopt the latest CUDA libraries, radically simplify your operations, and stay current with the latest security updates. Imagine if we told you that all of your existing NetApp investments, the reliability, the trust, the security, and protection of data ONTAP are being extended into this architecture, and so is the AFF A90s. Today, you no longer have to imagine. Welcome to NetApp Novus. Just like the Data Fabric extended the value of the NetApp platform and its technological capabilities from the enterprise data center to the world's biggest public clouds, Novus is designed to provide the architecture for the accelerated computing and AI era to help humans transform data into knowledge. You'll hear about NetApp Novus in depth through this conference and back on this stage tomorrow during the innovation keynote. To give you a sneak peek, please welcome one of the key architects of Novus and our NetApp's Chief Platform and Technology Officer, Arindam Banerjee. Arindam. Thank you, George. Every shift in compute rewrite storage, and NetApp has led every turn. With NFS, we disaggregated storage from the server. Virtualization disaggregated data services from the protocol. 18 months ago, we disaggregated performance from capacity. Now, as compute makes its biggest shift ever, it is time storage makes a decisive leap from a CPU-led gigabyte per second architecture to a GPU-led terabyte per second architecture. This shift is permanent, carrying with it every property that made persistence worth having in the first place. Data integrity, resilience, security. We do not get to trade those away for speed. A data outage could stall your most expensive asset. The most powerful GPU infrastructure on traditional storage is as useful as your favorite sports car on a dirt road. This is not a cycle you can wait out. AI is not a workload. AI will become a property of every workload. Training, inference, and agents do not run in isolation. They run side by side on the same infrastructure simultaneously. What we are seeing is a structural change in how the world procures, deploys, and fuels compute. GPU first, gigawatt scale AI factories with agents driving concurrency to a level this industry has never had to serve. Massive scale, extreme performance, relentless concurrency. That's NetApp Novus for you. Why do we need this shift now? Let me talk about some of the operational challenges at this scale. What breaks first? When you put 1,000 of GPUs and agents against a single namespace, the first thing that breaks is not throughput. It is when you combine millions of metadata operations on a namespace holding trillions of files along with throughput. It's metadata. Here is the part that costs you money. Your data operations get queued behind the metadata operations. The pipe is not full. Your GPUs are still idle. Two architectures served us very well up till now. Traditional NAS scales the enterprise. It gives us shared access, data services, protection, and governance. It was designed for human operations-based concurrency and continues to serve that well. Parallel file systems scaled HPC environments. They give us bandwidth for very large sequential jobs, but require the operator to tune the layout and accept a proprietary client on every node. An AI factory is both of these workloads at the same time, plus agents, and neither architecture was designed for this combination. Billions of tiny random reads for data loading while terabyte scale checkpoint writes land across the whole infrastructure all at once. This is bimodal I/O on the same namespace at the very same second. This is not something traditional disaggregated architectures can solve. This requires a new architecture. To talk about new architectures, let me call up on stage someone who knows a thing or two about parallel file systems and what it takes to operate at this scale. He has spent much of his career helping define how those systems work. HPCwire named him a legend and a godfather of parallel file systems. He has helped shape technologies such as PLFS, the Burst Buffer, Lustre, and pNFS. He's been a trusted advisor to our R and D team, reviewing everything from napkin sketches to architecture papers. It's an honor and a privilege for me to welcome on stage Gary Grider. Gary, thank you for speaking to us today. Thank you. It's an honor and a privilege having you here. HPCwire named you a 2025 legend and a godfather of parallel file systems. PLFS, the Burst Buffer, Los Alamos National Lab fingerprints on both Lustre and pNFS, and now GUFI. You run an environment that has roughly 6 million cores and a trillion files. When you operate at this scale, what breaks first? Well, imagine a workload like you've been talking about, a million cores all opening a single file at the same time, or all trying to create a file into the same directory at the same time. Those difficult things like that. Imagine 100,000 cores trying to walk your file system to find something all at the same time. Worse yet, imagine all of that all at the same time, right? Very concurrent workloads. What breaks first, probably metadata. You can probably scale the data pretty well, but the metadata is very difficult to scale. What would happen? The metadata server would get behind, everything would start timing out, and the file system would fall over, which happens. It's pretty much a metadata problem to scale like that. Well said, Gary. For 30 years, we have been designing file systems for the media use case, and we always bought hardware for the peak. Now, for AI factory scale and workloads, every job could be a peak, right? All of those peaks could come at the same time. When we started Novus, our design principle was not how fast can it go, but how can it handle all these jobs arriving at the same time and overlapping with each other? Gary, as we gear for AI factory scale, where exactly do current architectures hit the wall? Well, that mix of workloads that we just talked about requires you to think about parallelism along a lot of different dimensions to handle the concurrency of it. Scale-out NAS, it gave us parallel data. Parallel file systems gave us parallel data, but they also started scaling metadata. The problem is they did not scale it enough, and so we need to take the next step in that scaling activity. Basically, it becomes a challenge as the compute clusters increase the parallelism and the concurrency, and the multiple problems at the same time. To scale metadata commensurate with the services of scaling data, you need to do new things. The new thing that we just did was we started a project called Lattice, pNFS Lattice, and it was the world's first open source parallel file system metadata server for pNFS. The different thing about it is it is not just a scaled-out metadata service. We broke the service up into parts so that each part could scale separately to address these concurrent needs of a lot of bandwidth happening at the same time, a lot of inserts are happening at the same time, and a lot of queries are happening. You need to be able to scale the pieces independently, and that is what Lattice does. Okay. One structural idea, Gary, that keeps reoccurring across your career, as you have designed state-of-the-art systems for performance, scale, and concurrency, is disaggregation of metadata and data. In modern HPC and AI workload scale, how critical is that separation? Well, very critical. You basically have to start there. I think the problem is most architects think that if you separate data and metadata, you are done and all your scalability problems are over, and it turns out that is really not true. It has to do with the simultaneous stressful workloads. Splitting up the metadata and the data works great as long as your workloads are pretty homogeneous. But the minute that you have this concurrence, there are problems. Basically, in HPC, we do not really have that much concurrence. We run very large jobs. They compute for a long time, and then they dump everything at all at the same time. It is a very homogeneous workload. Yes, it is very bursty, but it is not a lot of different workloads all at the same time. AI data centers, as you said, are quite different. They have different workloads that are doing I/O all the time, and it is all pretty much stressful. This means that latencies can happen for some I/O operations, and inference just makes this worse because inference has to please people, and people have no time. They do not like to wait. Just like I said, that Lattice separates data and metadata and also breaks it up into different parts, this is a really important feature of it, and that is that you need to be able to scale those different parts dynamically. You cannot just scale them out and leave them scaled out all the time because that is a waste. What you want to do is build these pieces of the metadata service and have them scale dynamically as the workloads need it. That is a big important feature of the Lattice project. Where do you think architects consistently underestimate as they build out such environments? I would argue that the biggest thing that people underestimate is quality of service. I think quality of service at massive scale in an environment where something is failing all the time is a really hard thing to do. Yeah. It's really never been done completely well, ever. I would say that that's the thing that gets overlooked because it's just really hard. Okay. At this level of scale, failure is not an event. It can be an everyday condition, right? That is why the data servers in Novus run ONTAP, so that we can carry forward our decades of work on resilience. Gary, you helped steer Los Alamos National Labs' contributions to Lustre. Lustre's answer to metadata scale was DNE, that it shard the namespace across many metadata servers. What did we learn about the limits of that approach? Well, we certainly learned that there are limits. There are workloads that Lustre DNE 1, DNE 2, and DNE 2.5 or DNE 3 work pretty well for, and there are some that they don't work very well for. It wouldn't work very well concurrently. We've definitely learned a lot about what not to do, and we've learned some things to do. The Lattice project, essentially, it's trying to do this dynamic, multidimensional scaling all at the same time, and that's something we didn't try in Lustre. We should have, but we made the mistake of the metadata server being monolithic. Breaking it up gives us this ability to scale each thing independently and dynamically. Something that's less obvious in breaking up the metadata server is that if you break it up, you can break up into pieces and leverage different intellectual communities. If, for example, scalable protocols, that we've been doing scalable protocols for years, MPI and SHMEM and scale-out networks, NVLink and the like. Scalable catalog services, where does the metadata go? The distributed key-value store community has done a lot of that kind of work, and so we can rely on their expertise and scalable state. There's state. You can't insert the same files over more than once, and so you need state to be managed, and that's kind of a shared memory thing, a distributed shared memory thing. There's SHMEM and NVSHMEM and scale-out networks for that. The whole point here was to break it up, scale independently, fail over, and dynamically scale. Also to leverage different intellectual communities. That's what was totally different from what we did with Lustre. Then if you add to that the fact that this is all in a pNFS ecosystem. Yeah. Basically what we're trying to do is coalesce all of these things and their constituent intellectual contributors into one solution set, a technology that people can use. Basically be the handle all the workloads that we can imagine into the future because it's just dynamically growable in so many dimensions. Actually, very well said, Gary. You're very right. That is one of the unique advantages of building around standard protocols and open source. You get the community intellect to come together. Last question for you, Gary. Today, if you were told to design the storage systems that would power the next decade of AI factories, what would it look like? Well, in the 26 years I've been trying to get pNFS adopted, and looks like that's happening as we speak, we've done a lot right. We took our time and we did it right, so I think, leveraging global visibility, single namespaces, dynamic massively scaling dynamics, and nonstop. It has to be very reliable and have quality of service and all that kind of stuff. So I don't really think that way. I think, A, we've got this base that the world is about to adopt. How do we leverage that? How do we go to the next step? What do you add to that to make the thing have more value? For us, that becomes the big question. So given pNFS Flex Files and Local IO and all the fancy features that are in NFS and pNFS today, there's a lot you can do. And with the metadata broken out separately, you can do things like predicate pushdown and analytics. Basically this could replace Hadoop, line, and sinker, and it would be so much better and so much easier for everybody. Virtualized formats, if you have separated all this stuff, you do not have to present the data like the user wrote it. You can present the data like the user wants to see it, which is really different, right? There is no reason you cannot do stuff like that. Then virtual data lake houses. What is a lake house? It is a catalog full of stuff with formats. We have a catalog. We can integrate all that, and it would make everybody's life way easier and way better. I think that is the way we have to think about it is let us get this pNFS thing out there. Let us get it to scale to crazy widths, and then let us leverage the hell out of it with these ideas. I think the other thing that this does is it mates very well with global indexing, having an index, being able to find everything. Earlier was mentioned your AI product, which is really cool. We need that to scale to trillions of files. We need that to scale and not give away security and not give away all that stuff you said. I think we are within reach of doing that now. I could not have said that a year ago, I do not think. Anyway, I think we are on the right path to make this all happen. If I could summarize this, Gary, we have not only disaggregated metadata from data, we have actually opened several new frontiers. Metadata enables deep indexing, and deep indexing powers contextual search and pattern recognition. By disaggregating the metadata services from the catalog, we also enable virtual data formats that you just said, virtual lake houses that you just mentioned. In short, we have elevated metadata to knowledge. Gary, thank you. It has been an education for me, and I am sure for our audience as well. Thank you so much for being with us. Thank you. Thanks, everyone. Now let me explain how we design all of this into Novus. Let us start with the physics. At 50,000 - 100,000 GPUs, no single storage cluster can feed the fleet, not ours, not anyone else's. Every cluster has a ceiling set by its hardware and by its interconnect, and that ceiling sits far below what AI factories demand. We stop trying to build a bigger cluster. Our answer is federation. Many clusters, one namespace, and the throughput of the namespace is the summation of every cluster behind it. Add a storage cluster, add its throughput transparently. To your applications, it is still one file system tree, one mount, one namespace. Federation demands control and orchestration. Something has to own the namespace across all of the clusters. Get that wrong, and the namespace comes to a standstill, and the utilization of every GPU comes to a crawl. Not with NetApp Novus. Novus federates with linear scaling. That federation is ours to manage, not something that you need to think about. That is our disaggregated metadata layer. It assimilates the metadata from every cluster and orchestrates the whole. It is the metadata client or your metadata server that your clients connect to. It ensures that every node you add can truly contribute its full throughput to the entire pool. Because it sits on its own tier, data and metadata scale on different axes. Neither one waits for the other, and you would never have to buy one to get the other. This is how we will drive maximum GPU utilization and handle maximum concurrency. We designed all of this on standard protocols, pNFS, Flex Files. pNFS is an IETF standard, one that NetApp helped drive into existence. Flex Files is what stitches a single namespace across many thousands of servers spread across all these clusters. Here is the part that I want every operator in this room to hear. pNFS ships with the stock Linux kernel. There is no proprietary client. There is no kernel module of ours on a single one of your GPU nodes. Nobody else even attempting to come to this scale can say that, which means you set your own server upgrade cycle, your own driver matrix, your own qualification window. Your storage vendor does not get a vote when to patch your fleet. Gary made a point a moment ago that I want to come back to because that is the one architects consistently underestimate. Disaggregating metadata from data is not enough. Parallel file systems did that 20 years ago. The hard problem is scaling the metadata operations themselves under bimodal I/O at full concurrency with everything happening all at once. In Novus, we went one level deeper. We disaggregated the metadata service itself into independently scalable subsystems, a protocol plane, a state plane, and a catalog plane. Each one scales on its own axis. Each one grows exactly where the pressure is, and that is what gives you linear scaling of metadata operations. Underneath all of this, metadata is held in a distributed key-value store. That flattens your file system hierarchy, and your file system lookups become lightning-fast. There are no hotspots here because we never left a single place for the heat to arrive. At AI factory scale, resilience stops being a feature. It becomes structural. In a synchronized training job, every GPU is waiting for the slowest one. One stall read, one checkpoint that did not commit, and you have not slowed a job down. You have idled an entire cluster. We did not trade away persistence to buy speed. The Novus data servers will run ONTAP, which means the foundation that you have depended on for 30 years will hold good even at 100 TB a second. End-to-end data integrity, resilience, security, multi-tenant isolation, and quality of service. Scale without persistence is a proof of concept. Persistence without scale is irrelevant. Novus is the first architecture that refuses to choose. One last thing, because it is where Gary and I ended, and because it is the part that will matter five years from now. Once metadata is its own tier, independently scaled on its own plane with its own catalog, it stops becoming bookkeeping. It becomes something you can query, your agents can query. Designed for provenance, lineage, meaning across a trillion files. We are not getting metadata out of the way. We are elevating metadata. We have designed this for planet scale. Let me finish where George started. 100 TB per second. One namespace at zettabyte scale, the first of its kind. We believe it is an order of magnitude beyond anything anyone else has put out in the field. We have just designed the fastest storage on the planet for you. With that, let me call upon someone who knows a thing or two about Novus architecture. Let us bring George back out here. Thank you, Arindam, for all the hard work, and to all of the teams that have been involved in building Novus. Man, it has been fun. The most fun I have had in a long time. Gary, I want to thank you for the wisdom, the generosity, and for being the inspiration behind Novus. I am grateful to all our partners who have co-innovated with us, and I want to share one more with you, SAP, with whom we are expanding our strategic partnership. This is a deepening of a long-standing relationship, one built on years of joint work supporting the world's most mission-critical enterprise environments. You will hear more about this partnership tomorrow. We are so excited about the work that we have done with all of you. At the dawn of each era, NetApp has built the architectures and platforms that enabled organizations and mankind to do what was not possible before. Decades ago, with network file storage, we enabled teams of engineers to share large data sets to design the semiconductors that power the AI era today. Today, as those chips and AI models have become enormously powerful, we are privileged with Novus and the NetApp platform to deliver the capabilities so that we can play our part in helping humanity and all of you solve problems that were not possible before today. We are so excited about what we discussed today, and so I want to end where I began, with memory and context. In the agentic AI era, agents will work alongside your employees to make decisions, take actions, and drive real results for your business. To do so successfully, agents need to have knowledge enabled by memory, namely your data, and context, the metadata and knowledge graph that go with it. Today, the NetApp platform is uniquely designed to address and solve your data challenges for the AI era by delivering AI-ready data, unified storage, proactive protection, and complete intelligent control. With the NetApp platform, you have everything you need to build an intelligent data infrastructure to transform your data into knowledge. As you leave here, I have an ask of all of us. Let us be builders together of a better future, not just disruptors. Let us take responsibility for and act with agency to make the products and services that we create safe for those who trust us and use them. Let us use the enormous power and potential of these technologies to bring opportunity for everyone. As dad to two young adults, let us replace the fear that our industry has instilled in the next generation for their future with hope, as our parents did so generously for us. Thank you all for being here. Enjoy INSIGHT. We got an awesome conference.
Loading workspace