NVIDIA announced a range of local artificial intelligence tools, inference software optimizations, and new hardware partnerships at IFA 2026. The company said new compact NVIDIA RTX Spark Windows personal computers will arrive in October, featuring designs from Acer and Lenovo.
Open-source inference backends received updates that accelerate local processing. Kernel optimizations, faster prefill, and enhanced speculative decoding in llama.cpp delivered up to 1.9 times higher throughput on a GeForce RTX 5090. In vLLM, throughput rose 1.2 times on the RTX PRO 6000 Blackwell Workstation Edition and up to 1.4 times across two DGX Spark clusters through FlashInfer attention kernels. Both sets of improvements are available directly and inside the LM Studio and Ollama applications.
Agent Software and Routing
Agent developers are adding simplified setup workflows on Windows systems. Nous Research launched one-click setup for Hermes Agent, which detects NVIDIA GPUs and configures llama.cpp automatically. The OpenClaw project, which counts over 380,000 GitHub stars, released a Windows app configured for RTX GPUs with at least 24 gigabytes of video memory. Perplexity also made its Portable Computer agent available on Linux for systems with 24 gigabytes of video memory, with Windows support planned later. The app allows local workflows to selectively escalate tasks to 15 cloud models with user permission.
The company also released a beta version of the Personal AI Router, a free open-source tool that distributes inference tasks across PCs on a local network. The software operates on Windows, macOS, and Linux through graphical and command-line interfaces, integrating with LM Studio and Ollama. Supported hardware includes GeForce RTX 20 Series GPUs and newer, Turing-based RTX PRO workstation chips, DGX Spark systems, and Apple M4 or newer silicon.
Hardware and Model Support
Computer makers Acer and Lenovo presented new RTX Spark hardware scheduled for October releases. Acer showed a compact desktop concept, while Lenovo introduced its Yoga Pro 9n laptop and Yoga 9n 2-in-1 convertible. Each RTX Spark PC features an RTX Blackwell GPU rated at 1 petaflop, a 20-core Grace CPU, and up to 128 gigabytes of unified memory, running alongside the Windows Agent framework.
Game publishers Electronic Arts, Embark, and Ubisoft committed blockbuster titles to RTX Spark systems, following earlier commitments from Krafton, NetEase, Riot Games, and Xbox. CyberLink announced PhotoDirector AI PC Mode for PhotoDirector 365, using TensorRT-RTX and FP8 precision for on-device image editing when RTX Spark launches.
Several open models arrived alongside the platform updates. Releases included the 30-billion-parameter Nemotron 3.5 Lightning, Z.ai’s GLM-5.3-Flash, Meta’s 30-billion-parameter Muse Glimmer, and DeepSeek v4 Flash, which contains 284 billion total parameters and 13 billion active parameters. Video generation models LTX 2.5 and MiniMax-H3 also added local support, with FastVideo releasing a distilled four-step version named FastH3 that runs seven times faster.
