
Wednesday, March 18, 5:30 p.m. PT 🔗
20 Years of CUDA: Honoring the Architects of the Accelerated Age
What started in 2006 as a daring parallel computing guess has advanced into the foundational heartbeat of recent science and AI.
At GTC, NVIDIA is marking twenty years of CUDA — representing the efforts of over 6 million builders innovating throughout each layer of the computing stack. In the present day, it serves as a generational bridge between the pioneers who wrote the primary kernels and the following wave of builders deploying trillion-parameter AI fashions.
Led by NVIDIA CUDA Architect Stephen Jones, a panel at GTC Wednesday featured a bunch of researchers and engineers from Leap Buying and selling, Meta Superintelligence Labs and NVIDIA who highlighted the a long time of innovation behind CUDA, the way it helps builders remedy among the world’s most complicated issues — and the way techniques just like the NVIDIA DGX Spark desktop AI supercomputer will allow the following technology of CUDA builders.
The group shared reminiscences of the early days of CUDA — when “no one wished GPUs,” mentioned Paulius Micikevicius, a software program engineer at Meta Superintelligence Labs. “We needed to go and beg them to think about using GPUs.”
Throughout that point, Wen-Mei Hwu, senior distinguished analysis scientist and senior analysis director at NVIDIA, then a professor on the College of Illinois Urbana-Champaign, determined to construct a 200-GPU system in two months with a bunch of grad college students.
“A few weeks later, 200 GPU boards arrived, and energy provide and the whole lot — however there’s no chassis. So we ended up constructing wooden frames for every of those boards … and we ran the Green500 [benchmark] and we received No. 3,” Hwu mentioned. “That was the second I noticed that the power effectivity of GPUs has unimaginable potential.”
As the dimensions of accelerated computing has shifted to rack-scale techniques and AI factories, the panelists see desktop AI techniques like DGX Spark as a brand new approach ahead for prototyping and early improvement.
“So long as you have got that functionality to try this preliminary exploration and one thing that matches in your desk or your lap, that’s the vital factor,” mentioned Kate Clark, distinguished devtech engineer at NVIDIA. “I don’t see that going anyplace anytime quickly. We’ll at all times have CUDA all over the place.”
Monday, March 16, 1:30 p.m. PT 🔗
NVIDIA cuDF and cuVS Adopted by World’s Main Knowledge Platforms, Fueling Trendy Enterprise Knowledge Processing
Enterprises are producing a whole lot of zettabytes annually, and organizations are racing to show that info into insights. NVIDIA cuDF and cuVS — accelerated information libraries constructed on NVIDIA CUDA‑X — are being adopted by information platforms throughout industries to ship as much as 5x sooner efficiency whereas decreasing prices for structured and unstructured information processing.
Built-in with the world’s most generally used open supply information engines — downloaded over 200 million instances month-to-month by builders — these libraries are harnessed throughout enterprise information platforms, databases and information lakes. This helps organizations speed up innovation, develop extra correct fashions and course of extra information whereas managing prices.
For structured information, NVIDIA cuDF accelerates open supply information processing engines corresponding to Apache Spark, Presto, DuckDB, Polars and Velox, delivering as much as 5x sooner processing in contrast with CPU-only deployments.
For unstructured information — which represents 80% of at present’s enterprise information and is rising quickly — NVIDIA cuVS accelerates main engines together with FAISS, Amazon OpenSearch Service and Milvus. This helps brokers and functions extract context, info and suggestions from huge shops of textual content, pictures and video in a fraction of the time.
Powering Enterprise Knowledge Processing Platforms
Google Cloud integrates NVIDIA cuDF to speed up Apache Spark inside Dataproc and cuDF could be simply used inside Google Kubernetes Engine (GKE) to cut back processing instances for large ETL jobs from hours to seconds whereas reducing compute prices.
At Snap, which serves greater than 946 million lively customers, NVIDIA cuDF on GKE minimize each day information processing prices by 76%. This permits 10 petabytes of information to be analyzed inside a three-hour window — saving thousands and thousands of {dollars}.
“Our collaboration with NVIDIA and Google Cloud helps us innovate sooner for greater than a billion Snapchatters worldwide,” mentioned Saral Jain, chief info officer of Snap. “By reducing information processing prices and scaling experiments throughout petabytes of information, we’re delivering AI-powered experiences extra rapidly and effectively.”
IBM watsonx.information is a hybrid, open information platform that features open supply analytics engines corresponding to Apache Spark and Presto engines for structured information, and a vector engine primarily based on OpenSearch. In early experiments with Nestlé’s Order-to-Money mart, watsonx.information with NVIDIA cuDF accelerated workloads ran 5 instances sooner, with 83% decrease price financial savings.
“For an organization that serves billions, information underpins resolution making throughout our world operations,” mentioned Chris Wright, chief info and digital officer of Nestlé. “Working with IBM and NVIDIA, a focused proof of idea has demonstrated the power to refresh world operations information in a couple of minutes and at lowered price. Our focus now could be on turning this functionality into tangible enterprise influence — additional bettering resolution pace in areas corresponding to manufacturing and warehousing, and scaling these capabilities throughout our enterprise.”
The Dell AI Knowledge Platform with NVIDIA contains accelerated information engines that allow enterprises to rapidly and securely activate their Dell AI Manufacturing facility with AI-ready information. It options an Apache Spark-based processing engine accelerated with NVIDIA cuDF, delivering as much as 3x sooner efficiency, and an enterprise-grade vector database accelerated with NVIDIA cuVS, delivering as much as 12x greater throughput for vector indexing in contrast with CPUs.
“Objective-built for agentic AI, the Dell AI Knowledge Platform with NVIDIA makes use of accelerated information processing engines to make multimodal information AI-ready in hours as an alternative of days,” mentioned Michael Dell, chairman and CEO of Dell Applied sciences.
Oracle introduced that Oracle Non-public AI Providers Container can significantly speed up vector index creation in Oracle AI Database utilizing NVIDIA cuVS, serving to organizations pace up AI-enabled choices with the newest info.
“Enterprise AI is transferring from experimentation to manufacturing,” mentioned Clay Magouyrk, CEO of Oracle. “Oracle AI Database with NVIDIA expertise delivers AI-ready information inside minutes, enabling functions that have been beforehand inconceivable.”
NVIDIA cuDF and cuVS are supported by main enterprise information platforms together with EDB Postgres AI, NetApp, Snowflake, Starburst and VAST Knowledge — setting the muse for the AI‑powered future of information processing.
Friday, March 20, 2:00 p.m. PT 🔗
Quantum Computing Reaches an Inflection Level With NVIDIA NVQLink
At GTC, NVIDIA made NVQLink publicly out there via a brand new software programming interface (API) referred to as cudaq-realtime — and shared demonstrations advancing the state-of-the-art in quantum error correction.
First introduced at GTC Washington, D.C. in October, NVQLink’s low-latency, high-throughput connectivity between quantum processors and GPU supercomputing is now accessible to the quantum computing neighborhood, with product choices from Dell signaling fast ecosystem adoption.
The cudaq-realtime API, out there within the NVIDIA CUDA-Q software program platform, supplies open supply, turnkey integration of GPU-accelerated supercomputing with quantum processors. With cudaq-realtime, early adopters within the quantum neighborhood have demonstrated tight, real-time management of quantum {hardware} and deployed hybrid quantum-classical functions.
Main U.S. nationwide labs, together with Pacific Northwest Nationwide Laboratory and Lawrence Berkeley Nationwide Laboratory, together with QPU builders Quantinuum and Infleqtion and quantum software program supplier Q-CTRL, have adopted NVQLink, in some circumstances exhibiting order-of-magnitude reductions in decoding and calibration latencies in contrast with earlier work.
NVQLink has additionally begun deployment in industrial techniques, with QPU builder Anyon Computing and quantum techniques supplier SDT unveiling an NVQLink quantum-GPU system at a Korean industrial information heart, heralding the transfer to constructing accelerated quantum-supercomputing techniques in manufacturing environments.
Laptop-Aided-Engineering 🔗
Monday, March 16, 1:30 p.m. PT 🔗
NVIDIA Launches cuEST for Accelerated Quantum Chemistry in Semiconductor Design
NVIDIA this week launched NVIDIA cuEST, a brand new NVIDIA CUDA-X library that shifts electronic-structure calculations onto GPUs. Utilized Supplies, Samsung, Synopsys and TSMC are among the many preliminary adopters.
A number one-edge chip now accommodates over 50 billion transistors. Engineering them requires answering basic physics questions on the atomic scale: how electrons bond, how they migrate and the way they work together throughout movies just some atoms thick.
“As semiconductor scaling reaches the bodily limits of supplies, the business requires an enormous enhance in computing efficiency to simulate the quantum mechanics of next-generation chip designs,” mentioned Tim Costa, normal supervisor for industrial and computational engineering at NVIDIA. “With NVIDIA cuEST, business leaders can transfer previous the quantum bottleneck and take high-fidelity chemical modeling instantly into manufacturing to speed up semiconductor innovation.”
Trade Influence
- Utilized Supplies: Utilized Supplies makes use of cuEST-accelerated density purposeful idea (DFT) to mannequin difficult constructions, predict materials properties and research response pathways.
- Samsung: Samsung built-in cuEST into its inner pipeline, already accelerated on GPUs, to ship one more as much as 5x end-to-end speedup for key quantum-chemistry workloads.
- Synopsys: Powered by cuEST and QuantumATK, Synopsys expanded its performance to incorporate Gaussian-basis DFT, accelerating simulations as much as 30x for semiconductor workflows.
- TSMC: TSMC makes use of cuEST’s accelerated quantum chemistry to advance processes for next-generation silicon design.
From the Lab to the Fab
The most typical technique for atomistic modeling is density purposeful idea. DFT gives a robust steadiness between accuracy and scalability; nonetheless, its computational price has restricted its widespread use in business, maintaining most functions confined to analysis. With cuEST, NVIDIA makes excessive‑accuracy quantum‑chemistry possible at an industrial scale and in actual manufacturing workflows.
Traditionally, the business has relied on CPU clusters to run these simulations, evaluating candidate supplies, together with gate dielectrics and interconnect metals, one batch at a time over hours or days.
cuEST supplies optimized routines so GPUs can speed up the core matrices of a Gaussian-basis DFT calculation, together with overlap, kinetic power, nuclear attraction, Coulomb and exchange-correlation. It additionally helps purposeful approximations starting from customary generalized gradient approximation to hybrid functionals, permitting engineers to steadiness computational price with accuracy.
NVIDIA’s objective for cuEST: transferring high-fidelity materials modeling from the lab to the fab.
Be taught extra about cuEST by becoming a member of the NVIDIA demo sales space and Synopsys’ sales space at GTC, and dive deeper within the GTC session, “Subsequent-Technology Discovery: Agentic AI for Science, AI-Pushed Simulation and GPU-Accelerated Chemistry.”
