|
// you commented nv
free API keys to 80+ AI models
As promised: all the links, the cheat sheet, and a guide to what you can start building with this today. NVIDIA hosts 120+ models behind one free API key — here’s the whole thing, tested live before this went out.
| > get_your_key 3 STEPS, ~2 MINUTES | 1. Free account at build.nvidia.com — email in, Developer Program is free. 2. Open any model in the catalog — each model page has a live playground and a code panel. 3. Hit “Get API Key”. Your key starts nvapi- and works on every model in the catalog, not just the one you clicked. THE LIMITS — STRAIGHT 1,000 free inference credits on signup (up to 5,000 on request) 40 requests per minute, per model £0 — no credit card, ever, for the free tier |
Honest framing: that’s plenty to build and test, not to run production. A credit is roughly one request on big models, a fraction on small ones — you’ll prototype for weeks on it. |
| > the_cheat_sheet ONE SNIPPET, EVERY MODEL | The whole trick: NVIDIA’s endpoint speaks the OpenAI SDK. Change two lines — the base URL and the key — and every script, tool and tutorial built for OpenAI now runs on NVIDIA’s catalog for free. from openai import OpenAI
client = OpenAI( base_url="https://integrate.api.nvidia.com/v1", api_key="nvapi-YOUR-KEY", )
r = client.chat.completions.create( model="z-ai/glm-5.2", # any ID from the catalog messages=[{"role": "user", "content": "your prompt"}], ) print(r.choices[0].message.content) |
I ran this exact snippet today against z-ai/glm-5.2 — it works. Model picker — 4 IDs worth your first credits (all pulled live from the API today): z-ai/glm-5.2 the flagship — my default for anything that needs to be smart nvidia/llama-3.3-nemotron-super-49b-v1.5 NVIDIA’s own workhorse — strong reasoning, mid-size nvidia/llama-3.1-nemotron-nano-8b-v1 small and fast — bulk jobs, classification, cheap volume meta/llama-3.2-11b-vision-instruct reads images — screenshots, photos, scanned documents Swap the model= line, nothing else changes. The full list is 120+ deep — browse it at build.nvidia.com/models. |
| > things_to_build 5 STARTER BUILDS FOR A BUSINESS OWNER | lead-qualifier bot Every enquiry from your website form gets scored, summarised and routed — hot ones ping your phone. → small fast model (nemotron-nano) — classification is its whole job review-reply writer Pulls each new Google review and drafts a reply in your tone for you to approve. → small fast model — short outputs, high volume, costs nothing bulk product descriptions CSV of products in, clean unique descriptions out — hundreds in one run. → flagship (glm-5.2) for quality, nano when you’re doing thousands doc-summariser for onboarding New client sends contracts and briefs; you get a one-page summary with dates, numbers and red flags. → flagship — long documents need the big context and the judgment search over your own files (RAG) Point it at your folder of proposals and SOPs, then ask questions in plain English and get answers with sources. → flagship for answers + the vision model for scans and screenshots Every one of these is the snippet above plus a loop. Pick the one that eats the most of your week and build that first. |
| > the_catch WHY IT'S FREE — HONESTLY | NVIDIA isn’t being generous — it’s a hardware standard play. Every developer who prototypes on their endpoint is a future customer for their GPUs and DGX Cloud when the thing works and needs to scale. You’re the pipeline, and that’s fine — take the free compute. So use it for what it’s for: prototyping. Prove the build works, show the client, validate the idea — on £0. When it’s real — credits burning down, 40 req/min pinching — you move: pay per token on a production API, or self-host the same models on your own GPU. The snippet doesn’t change; only the base URL does. That’s the point of building on the OpenAI-compatible standard. Running Claude for the building itself? Read this first: never hit your Claude limit again → |
Grab your key, run the snippet, reply and tell me what you’re building.
Talk soon, Mike
|