AI Model Co-Design: Hardware-Friendly LLM Design
This article discusses the balance of accuracy, throughput, and interactivity in AI performance, focusing on how model design choices impact throughput and interactivity in large language models.
The auto-updated technology brief
This article discusses the balance of accuracy, throughput, and interactivity in AI performance, focusing on how model design choices impact throughput and interactivity in large language models.
LLM training workloads face GPU memory limits due to competing demands on high-bandwidth memory. This article explains how host offloading can alleviate these bottlenecks in JAX-based training.
Telecom operators have invested over $240B in wireless spectrum. AI-native RAN and NVIDIA AI Aerial aim to maximize spectral efficiency, enhancing network capacity and resilience.
NVIDIA introduces a LangChain deep agents harness profile for the Nemotron 3 Ultra to enhance AI agent performance by balancing accuracy and cost through fine-tuning smaller open models.
NVIDIA demonstrates how autonomous coding AI agents can improve vision reasoning models to over 90% accuracy with minimal manual effort, streamlining adaptation to production video tasks.
Coding AI agents are becoming practical operators for long-running machine learning workflows, enabling tasks like repository inspection, runtime setup, issue resolution, experiment launching, monitoring, and result summarization, which is crucial for reinforcement learning research.
Google removed the AI image generation feature from Google Earth less than a day after its release due to concerns about misleading images.
NVIDIA discusses integrating video analytics AI agents that perceive, reason, and act on video footage into enterprise workflows and applications such as content management and messaging platforms.
NVIDIA enables developers to customize the Nemotron 3 Nano AI model quickly using Prime Intellect Lab, addressing challenges like infrastructure and domain expertise.
OpenUSD is an open, extensible framework enabling teams to integrate CAD data, simulation assets, and telemetry into a shared, physically accurate world view. NVIDIA discusses accelerating lightweight USD runtime development using AI agents.
NVIDIA explores estimating probabilities of rare, high-impact events using guided generative models to improve efficiency over brute-force Monte Carlo sampling.
NVIDIA introduces NVLink as a high-performance scale-up network designed to meet the growing demands of AI factories, enabling faster deployment of AI compute infrastructure for larger, more complex workloads.