HomeAIHugging Face Adds Baseten as Hub Infer
AI

Hugging Face Adds Baseten as Hub Inference Provider

The partnership brings serverless inference for models like DeepSeek V4 Flash and Kimi K3 directly to Hugging Face model pages and SDKs.

WHAT YOU NEED TO KNOW
  • Hugging Face added Baseten as a supported inference provider on August 6, 2026.
  • Launch models include Kimi K3, DeepSeek V4 Flash, and GLM-5.2 for conversational and text-generation tasks.
  • Integration is supported in Python SDK huggingface_hub 1.26.1 or later and the JavaScript inference library.
EPOCH CAPABILITIES INDEX (ECI)GPT-5.6 Sol162GPT-5.5 Pro161Claude Fable 5161Claude Opus 5159GPT-5.5158Kimi K3156Source: Epoch AI Benchmarking Hub - CC BY 4.0 - as of 2026-07-31

Baseten has joined Hugging Face as a supported inference provider, bringing serverless model hosting directly to the platform's repository, Hugging Face announced on August 6, 2026.

The initial integration introduces support for conversational and text-generation workloads across open-weight models, including Kimi K3, DeepSeek V4 Flash, and GLM-5.2. Developers can access Baseten's catalog through official Hugging Face software development kits, specifically huggingface_hub version 1.26.1 or later in Python and the @huggingface/inference library for JavaScript. Hugging Face stated that support for additional machine learning tasks will roll out in subsequent updates.

Calls to Baseten operate through two distinct configurations on the platform. Users can supply their own Baseten API key to direct calls to the provider, which bills the external Baseten account. Alternatively, requests can be routed through Hugging Face using a standard platform token, consolidating charges onto the user's Hugging Face account at base provider rates with no added markup.

Third-party agent harnesses, such as Pi, OpenCode, Hermes Agents, and OpenClaw, can connect directly to Baseten-hosted models through the router endpoint without additional adapter code. Account settings on the platform allow users to establish provider preferences, which determine how third-party services appear across model widgets and sample code snippets.

Hugging Face Pro subscribers receive two dollars in monthly inference credits that apply across all supported providers, alongside ZeroGPU access, Spaces developer mode, and twenty-fold limit increases. Free registered accounts retain access through a smaller baseline inference quota, while Hugging Face noted it may explore revenue-sharing arrangements with provider partners in the future.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →