automate_&_hustle PAYOFF
REVENUE-FIRST AI
// you commented nv

free API keys to 80+ AI models

As promised: all the links, the cheat sheet, and a guide to what you can start building with this today. NVIDIA hosts 120+ models behind one free API key — here’s the whole thing, tested live before this went out.

> the_links  EVERYTHING FROM THE REEL

build.nvidia.com →
the catalog — every hosted model, one place, this is where your key lives

build.nvidia.com/models →
browse all of them — filter by LLM, vision, speech, embeddings

developer.nvidia.com/developer-program →
the free Developer Program signup that unlocks the credits

docs.api.nvidia.com →
the API docs when you want to go past the snippet below

> get_your_key  3 STEPS, ~2 MINUTES

1. Free account at build.nvidia.com — email in, Developer Program is free.

2. Open any model in the catalog — each model page has a live playground and a code panel.

3. Hit “Get API Key”. Your key starts nvapi- and works on every model in the catalog, not just the one you clicked.

THE LIMITS — STRAIGHT
1,000 free inference credits on signup (up to 5,000 on request)
40 requests per minute, per model
£0 — no credit card, ever, for the free tier

Honest framing: that’s plenty to build and test, not to run production. A credit is roughly one request on big models, a fraction on small ones — you’ll prototype for weeks on it.

> the_cheat_sheet  ONE SNIPPET, EVERY MODEL

The whole trick: NVIDIA’s endpoint speaks the OpenAI SDK. Change two lines — the base URL and the key — and every script, tool and tutorial built for OpenAI now runs on NVIDIA’s catalog for free.

from openai import OpenAI

client = OpenAI(
    base_url="https://integrate.api.nvidia.com/v1",
    api_key="nvapi-YOUR-KEY",
)

r = client.chat.completions.create(
    model="z-ai/glm-5.2",  # any ID from the catalog
    messages=[{"role": "user", "content": "your prompt"}],
)
print(r.choices[0].message.content)

I ran this exact snippet today against z-ai/glm-5.2 — it works.

Model picker — 4 IDs worth your first credits (all pulled live from the API today):

z-ai/glm-5.2
the flagship — my default for anything that needs to be smart

nvidia/llama-3.3-nemotron-super-49b-v1.5
NVIDIA’s own workhorse — strong reasoning, mid-size

nvidia/llama-3.1-nemotron-nano-8b-v1
small and fast — bulk jobs, classification, cheap volume

meta/llama-3.2-11b-vision-instruct
reads images — screenshots, photos, scanned documents

Swap the model= line, nothing else changes. The full list is 120+ deep — browse it at build.nvidia.com/models.

> things_to_build  5 STARTER BUILDS FOR A BUSINESS OWNER

lead-qualifier bot
Every enquiry from your website form gets scored, summarised and routed — hot ones ping your phone.
→ small fast model (nemotron-nano) — classification is its whole job

review-reply writer
Pulls each new Google review and drafts a reply in your tone for you to approve.
→ small fast model — short outputs, high volume, costs nothing

bulk product descriptions
CSV of products in, clean unique descriptions out — hundreds in one run.
→ flagship (glm-5.2) for quality, nano when you’re doing thousands

doc-summariser for onboarding
New client sends contracts and briefs; you get a one-page summary with dates, numbers and red flags.
→ flagship — long documents need the big context and the judgment

search over your own files (RAG)
Point it at your folder of proposals and SOPs, then ask questions in plain English and get answers with sources.
→ flagship for answers + the vision model for scans and screenshots

Every one of these is the snippet above plus a loop. Pick the one that eats the most of your week and build that first.

> the_catch  WHY IT'S FREE — HONESTLY

NVIDIA isn’t being generous — it’s a hardware standard play. Every developer who prototypes on their endpoint is a future customer for their GPUs and DGX Cloud when the thing works and needs to scale. You’re the pipeline, and that’s fine — take the free compute.

So use it for what it’s for: prototyping. Prove the build works, show the client, validate the idea — on £0.

When it’s real — credits burning down, 40 req/min pinching — you move: pay per token on a production API, or self-host the same models on your own GPU. The snippet doesn’t change; only the base URL does. That’s the point of building on the OpenAI-compatible standard.

Running Claude for the building itself? Read this first: never hit your Claude limit again →

get one real AI build a week →

Grab your key, run the snippet, reply and tell me what you’re building.

Talk soon,
Mike

AUTOMATE_HUSTLE · REVENUE-FIRST AI · does it make money?