Hugging Face Integration Guide

Technology: hugging-face · Category: ai · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/hugging-face

Insight:

Hugging Face is home for open-source and fine-tuned models. Its serverless Inference Providers router replaced the old Inference API — one HF token routes to Together, Cerebras, Groq, Fireworks and more behind an OpenAI-compatible endpoint (router.huggingface.co/v1). Reach for it when you need a model you can self-host, fine-tune, or run cheaply (e.g. a Swahili-tuned model) rather than a frontier API.

██╗  ██╗██╗   ██╗ ██████╗  ██████╗ ██╗███╗   ██╗ ██████╗     ███████╗ █████╗  ██████╗███████╗
██║  ██║██║   ██║██╔════╝ ██╔════╝ ██║████╗  ██║██╔════╝     ██╔════╝██╔══██╗██╔════╝██╔════╝
███████║██║   ██║██║  ███╗██║  ███╗██║██╔██╗ ██║██║  ███╗    █████╗  ███████║██║     █████╗
██╔══██║██║   ██║██║   ██║██║   ██║██║██║╚██╗██║██║   ██║    ██╔══╝  ██╔══██║██║     ██╔══╝
██║  ██║╚██████╔╝╚██████╔╝╚██████╔╝██║██║ ╚████║╚██████╔╝    ██║     ██║  ██║╚██████╗███████╗
╚═╝  ╚═╝ ╚═════╝  ╚═════╝  ╚═════╝ ╚═╝╚═╝  ╚═══╝ ╚═════╝     ╚═╝     ╚═╝  ╚═╝ ╚═════╝╚══════╝

Hugging Face Integration Guide

Focus: Using Hugging Face models, Spaces, Inference Providers, and the official HF MCP server inside Claude Code sessions and automation pipelines.

Overview

Hugging Face hosts 1,000,000+ open-source AI models, datasets, and Spaces. From inside Claude Code you can run models through Inference Providers (a single HF token routing to Together, Cerebras, Fal, Groq, Fireworks and others behind one OpenAI-compatible endpoint), deploy to Spaces, manage datasets, and run Gradio apps — all via the official HF MCP server or the hf CLI. This makes Claude Code a hub for open-model experimentation alongside proprietary APIs.

Here is the big picture — the Hub sits at the center, and Claude Code reaches it through two friendly paths that feed straight into your app.

flowchart LR
  Hub["Hugging Face Hub<br/>models · datasets · Spaces"]
  CC["Claude Code"]
  MCP["HF MCP server"]
  CLI["hf CLI"]
  Inf["Inference Providers"]
  Local["Local weights<br/>download · offline"]
  App["Your app"]
  CC --> MCP
  CC --> CLI
  MCP --> Hub
  CLI --> Hub
  Hub -->|"serverless"| Inf
  Hub -->|"download"| Local
  Inf --> App
  Local --> App

Official Documentation

Resource URL
HF Hub Docs https://huggingface.co/docs/hub
Inference Providers https://huggingface.co/docs/inference-providers
HF MCP Server https://huggingface.co/docs/hub/en/hf-mcp-server
huggingface_hub Python https://huggingface.co/docs/huggingface_hub
Gradio https://www.gradio.app/docs
Transformers https://huggingface.co/docs/transformers
Gradio Spaces https://huggingface.co/docs/hub/spaces-sdks-gradio
Spaces Config Reference https://huggingface.co/docs/hub/spaces-config-reference
MCP Course https://huggingface.co/learn/mcp-course

MCP Server Setup

Hugging Face hosts an official MCP server at https://huggingface.co/mcp. It exposes Hub tools plus any Gradio Space as an additional tool.

# Add HF MCP server to Claude Code (HTTP transport)
claude mcp add --transport http hf-mcp-server \
  https://huggingface.co/mcp \
  -H "Authorization: Bearer ${HF_TOKEN}"

.mcp.json Configuration

{
  "mcpServers": {
    "huggingface": {
      "type": "http",
      "url": "https://huggingface.co/mcp",
      "headers": {
        "Authorization": "Bearer ${HUGGING_FACE_HUB_TOKEN}"
      }
    }
  }
}

Built-in MCP Tools

Tool Description
search_models Search the Hub for models by task, framework, or name
get_model_info Get metadata, card, and usage info for a model
list_datasets Browse and search datasets
inference Run inference on any Inference Providers-compatible model
list_spaces Browse Gradio Spaces
run_space Call a Gradio Space as a tool
create_repo Create a new model/dataset/space repo
upload_file Upload files to a Hub repo

Community alternative

For running Spaces locally there is a community option, npx -y @llmindset/mcp-hfspace. It is no longer the canonical path — prefer the official HF MCP server above.


CLI Integration (hf)

The unified hf CLI (shipped with huggingface_hub) is the current standard. New docs and examples use hf.

Installation + auth

pip install -U huggingface_hub
hf auth login                 # interactive; stores token at ~/.cache/huggingface/token
hf auth login --token $HF_TOKEN   # non-interactive (CI)
hf auth whoami                # confirm the active account

Key commands

# Search & inspect
hf download meta-llama/Llama-3.3-70B-Instruct      # pull weights/files to the local cache
hf upload <user>/my-model ./model-dir              # push a folder to a repo
hf repos create my-model --repo-type model         # create a model/dataset/space repo

# Run a quick chat inference from the shell (Inference Providers)
python -c "
import os
from huggingface_hub import InferenceClient
c = InferenceClient(token=os.environ['HF_TOKEN'])
print(c.chat.completions.create(
    model='openai/gpt-oss-120b',
    messages=[{'role':'user','content':'Say jambo in one word.'}],
).choices[0].message.content)
"

Prefer hf auth login and hf download going forward.


Inference Providers (serverless)

Since 2025, Hugging Face's serverless inference is Inference Providers: one HF token, a single OpenAI-compatible router (https://router.huggingface.co/v1), and automatic routing to best-in-class providers. There is no markup on provider rates, and PRO accounts get included monthly credits.

sequenceDiagram
  participant App as "Your code"
  participant Router as "router.huggingface.co"
  participant Prov as "Provider (Together / Cerebras / Fal / …)"
  App->>Router: "POST /v1/chat/completions (HF token)"
  Router->>Prov: "route by policy (auto / cheapest / named)"
  Prov-->>Router: "tokens"
  Router-->>App: "OpenAI-shaped response"

Providers today (18, per the docs sidebar): Baseten, Cerebras, Cohere, DeepInfra, Fal AI, Featherless AI, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeed AI, Z.ai. The roster shifts — check hf.co/docs/inference-providers for the live list.

Provider routing — append a policy or provider suffix to the model id:

Suffix Meaning
(none) / :fastest Default. Highest throughput (tokens/sec) available.
:cheapest Lowest price per output token.
:preferred First available in your order at hf.co/settings/inference-providers.
:together, :cerebras, … Force a specific provider.

Python (huggingface_hub)

pip install huggingface_hub
hf auth login   # read token from hf.co/settings/tokens
import os
from huggingface_hub import InferenceClient

client = InferenceClient(token=os.environ["HF_TOKEN"])

# Chat completion — provider defaults to "auto" (fastest)
completion = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Explain async/await in one sentence."}],
)
print(completion.choices[0].message.content)

# Force a provider for cost/latency control
completion = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-R1",
    provider="novita",   # DeepSeek-R1's live provider; check the model page for options
    messages=[{"role": "user", "content": "Habari!"}],
)

JavaScript / TypeScript (@huggingface/inference)

npm install @huggingface/inference
import { InferenceClient } from "@huggingface/inference";

const client = new InferenceClient(process.env.HF_TOKEN);

const chat = await client.chatCompletion({
  model: "openai/gpt-oss-120b",
  messages: [{ role: "user", content: "What is a transformer?" }],
  // provider defaults to "auto"; set provider: "together" to pin one
});
console.log(chat.choices[0].message.content);

Drop-in OpenAI replacement (chat only)

Swap the base URL and keep your existing OpenAI client:

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://router.huggingface.co/v1",
  apiKey: process.env.HF_TOKEN,
});

const completion = await client.chat.completions.create({
  model: "deepseek-ai/DeepSeek-R1:cheapest",
  messages: [{ role: "user", content: "Hello!" }],
});

REST (curl)

curl https://router.huggingface.co/v1/chat/completions \
  -H "Authorization: Bearer $HF_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b:fastest",
    "messages": [{"role": "user", "content": "How many Gs in huggingface?"}],
    "stream": false
  }'

GET https://router.huggingface.co/v1/models lists every available model with per-provider pricing, context length, latency, and throughput — handy for building a model picker.

The OpenAI-compatible endpoint is chat-only. For text-to-image, embeddings, or speech, use the inference clients (client.text_to_image(...), client.feature_extraction(...), etc.).


Pricing breakdown

Verified against https://huggingface.co/pricing on 2026-08-23. Inference Providers add no markup on the underlying provider's per-token rate; GET /v1/models shows live rates.

Account plans

Plan Price What you get
Free $0 Hub access, a generous Inference Providers free tier, CPU-Basic & ZeroGPU Spaces
PRO $9 / mo 20× included inference credits, 8× ZeroGPU quota + highest queue priority, 10× private storage, PRO badge
Team $20 / user / mo Org features, SSO, central billing, included org credits
Enterprise $50 / user / mo Advanced security, audit logs, dedicated support, data residency

Inference Endpoints (dedicated GPUs, scale-to-zero)

Hardware Hourly (per card, AWS; scales with count/region)
CPU (Intel Sapphire Rapids) from $0.03
GPU T4 from $0.50
GPU L4 from $0.80
GPU L40S from $1.80
GPU A100 from $2.50
GPU H100 from $4.50
GPU H200 from $5.00
GPU B200 from $9.25

Spaces hardware

Flavor Hourly
CPU Basic Free
ZeroGPU Free (daily quota: 5 min free / 40 min PRO)
CPU Upgrade $0.03
GPU T4 (small / medium) $0.40 / $0.60
GPU L4 from $0.80
GPU A10G from $1.00
GPU A100 $2.50 – $20.00

ZeroGPU hardware is now the NVIDIA RTX Pro 6000 Blackwell (no longer H200): size large (default) is half a card / 48 GB at 1× quota cost, xlarge is the full card / 96 GB at 2×. Quota is daily — 5 min free, 40 min PRO/Team, 60 min Enterprise — extensible past the cap with pre-paid credits at $1 / 10 min of GPU time. ZeroGPU is Gradio-SDK-only.

Storage

$12 / TB / mo (public) to $18 / TB / mo (private); volume discounts to $8–$10 / TB / mo at 50 TB+, 200 TB+, 500 TB+.

HF-open vs frontier (the codeAmani calculus)

For low-to-mid volume, Inference Providers on a PRO plan is often cheaper than a frontier API and gives you open weights you can later self-host. Pin :cheapest for batch jobs, :fastest for interactive UX. When traffic is steady and high, a dedicated Inference Endpoint (scale-to-zero) beats per-token pricing. For free public demos, a ZeroGPU Space runs a real GPU at no hourly cost.


Deploy a Space

A Space is just a Git repo that Hugging Face builds and runs for you. Pick the Gradio SDK and three files do the work: README.md (a YAML config block tells the runtime what to build), app.py (your Gradio interface), and requirements.txt (deps the runtime installs). Push the repo and your app is live at https://huggingface.co/spaces/<user>/<space>.

Here is the shape of it — Claude Code assembles the three files, pushes once, and the Space build serves your app.

flowchart LR
  CC["Claude Code"]
  Files["README.md · app.py<br/>requirements.txt"]
  Repo["Space repo<br/>git on the Hub"]
  Build["Space build<br/>install deps · launch"]
  Live["Live app<br/>App tab"]
  CC --> Files
  Files -->|"hf upload"| Repo
  Repo --> Build
  Build --> Live

1. The three files

README.md — the YAML block at the top is the Space config (sdk: gradio initializes the latest Gradio; pin sdk_version for reproducible builds):

---
title: Swahili Helper
emoji: 🇰🇪
colorFrom: green
colorTo: blue
sdk: gradio
sdk_version: 6.25.0
app_file: app.py
pinned: false
short_description: Swahili chat demo
---

# Swahili Helper
A small Gradio demo running on Hugging Face Spaces.

app.py — a minimal Gradio interface (the runtime runs app_file automatically):

import gradio as gr

def greet(name: str) -> str:
    return f"Habari, {name}!"

demo = gr.Interface(fn=greet, inputs="text", outputs="text", title="Swahili Helper")

if __name__ == "__main__":
    demo.launch()

requirements.txt — the Spaces runtime installs these on build:

gradio

2. Create + push from the CLI

# Create the Space repo (Gradio SDK) — one time
hf repos create swahili-helper --repo-type space --sdk gradio

# Upload the whole folder (README.md, app.py, requirements.txt) in one commit
hf upload <user>/swahili-helper ./swahili-helper .

…or with the Python SDK

Useful inside a Claude Code automation that generates a Space programmatically:

from huggingface_hub import HfApi

api = HfApi(token=os.environ["HUGGING_FACE_HUB_TOKEN"])

api.create_repo(
    repo_id="<user>/swahili-helper",
    repo_type="space",
    space_sdk="gradio",
    exists_ok=True,
)

api.upload_folder(
    repo_id="<user>/swahili-helper",
    repo_type="space",
    folder_path="./swahili-helper",  # contains README.md, app.py, requirements.txt
)

After the push, the App tab shows the build logs and then the running app.

Gotcha: pin sdk_version

If you leave sdk_version out of the YAML block, Spaces builds against the latest Gradio on every rebuild. A breaking Gradio release can then silently break a Space that worked yesterday — pin sdk_version (e.g. 6.25.0, current major is Gradio 6) so builds stay reproducible. Other notes: Spaces default to free cpu-basic hardware (set GPU flavors via the Settings tab, not the YAML — suggested_hardware is only a hint for users who duplicate your Space); ZeroGPU Spaces are Gradio-SDK-only; and never commit your HF_TOKEN into app.py — add it under Settings → Variables and secrets and read it with os.environ.


Environment Variables

# Required
HUGGING_FACE_HUB_TOKEN=hf_...

# Aliases (all work)
HF_TOKEN=hf_...
HUGGINGFACE_TOKEN=hf_...

# Cache directory (optional)
HF_HOME=~/.cache/huggingface
HF_HUB_CACHE=~/.cache/huggingface/hub

# Use local files only (offline mode)
TRANSFORMERS_OFFLINE=1
HF_HUB_OFFLINE=1

Automation Workflows

Claude Code Hook: Auto-summarize Model Cards

.claude/settings.json:

{
  "hooks": {
    "UserPromptSubmit": [
      {
        "matcher": ".*huggingface\\.co.*",
        "hooks": [
          {
            "type": "command",
            "command": "python scripts/fetch-model-card.py"
          }
        ]
      }
    ]
  }
}

Slash Command: Run HF Inference

.claude/commands/hf-infer.md:

Run Hugging Face inference on model $ARGUMENTS using the InferenceClient.

Use the Bash tool to execute:
```bash
python -c "
from huggingface_hub import InferenceClient
import os, sys
client = InferenceClient(token=os.environ['HF_TOKEN'])
model, *prompt_parts = '$ARGUMENTS'.split(' ', 1)
prompt = prompt_parts[0] if prompt_parts else 'Hello'
print(client.text_generation(prompt, model=model, max_new_tokens=200))
"
```

Usage: /project:hf-infer mistralai/Mistral-7B-v0.1 Explain async/await

CI/CD: Model Evaluation Pipeline

# .github/workflows/evaluate-model.yml
name: Evaluate Fine-tuned Model
on:
  push:
    paths: ['models/**']

jobs:
  evaluate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: '3.11' }
      - run: pip install huggingface_hub evaluate datasets
      - name: Run evaluation
        env:
          HUGGING_FACE_HUB_TOKEN: ${{ secrets.HF_TOKEN }}
        run: python scripts/evaluate.py --model ${{ github.sha }}
      - name: Push results to Hub
        run: python scripts/push-results.py

transformers.js — models in the browser / on the edge

@huggingface/transformers (transformers.js) runs models client-side via WebGPU/WASM — no server, no per-call cost, and it works offline after the first load. For codeAmani this is the low-bandwidth / intermittent-connectivity play: ship a small embedding or classification model to the device and skip the round-trip entirely on 2G/3G.

npm install @huggingface/transformers
import { pipeline } from "@huggingface/transformers";

// Lazy-load a tiny model once; runs entirely in the browser tab thereafter.
const classify = await pipeline("sentiment-analysis");
const out = await classify("M-Pesa payment received, asante!");
// → [{ label: "POSITIVE", score: 0.99… }]

Good fits: on-device sentiment/intent, semantic search over a small local corpus, redaction before a network call. Heavy generation still belongs on Inference Providers or an Endpoint.


Apply HF to the rest of the stack

Hugging Face is not an island — it slots into the other codeAmani tools:

import { InferenceClient } from "@huggingface/inference";
const hf = new InferenceClient(process.env.HF_TOKEN);
const vector = await hf.featureExtraction({
  model: "intfloat/multilingual-e5-large",
  inputs: "Bei ya unga ni shilingi ngapi?",
});
// → number[]  → upsert into pgvector / Pinecone

Common Use Cases

Use Case Approach
Open-source model inference InferenceClient or HF MCP inference tool
Dataset exploration MCP list_datasets → get_dataset_info
Space deployment hf upload + Gradio launch()
Model fine-tuning Upload training data, trigger AutoTrain via API
Embedding generation sentence-transformers via Inference Providers
Image generation SDXL via client.text_to_image()

This workspace has the huggingface-skills plugin installed — reach for these instead of hand-rolling:

Skill Use it for
huggingface-skills:hf-cli Scripting the hf CLI
huggingface-skills:huggingface-spaces / :huggingface-gradio Building & shipping Spaces / Gradio apps
huggingface-skills:huggingface-zerogpu ZeroGPU-backed demos (free GPU on PRO)
huggingface-skills:huggingface-lora-space-builder LoRA fine-tune + Space in one go
huggingface-skills:trl-training / :train-sentence-transformers Fine-tuning LLMs / embedders
huggingface-skills:transformers-js Browser/edge inference
huggingface-skills:huggingface-datasets Dataset loading & curation

Troubleshooting

Issue Fix
401 Unauthorized Re-login: hf auth login or check HF_TOKEN
Model loading slowly Use InferenceClient (serverless); avoid cold starts
Gated model access Accept terms at huggingface.co/model-name first
MCP server not connecting Verify token in headers; try curl https://huggingface.co/mcp
Out of memory locally Use Inference Providers instead of loading weights locally

Official docs: