hostfleet Find my setup
AI infrastructure, made visible

Host AI.
Without guessing.

See what runs where, what hardware it needs, and what it can cost — before you deploy anything.

Sourced14 GPU types · 13 providers · verified 2026-08-13
14GPU types tracked
13providers compared
3hosting paths explained
0fake benchmarks
Start with the workload

What are you trying to run?

Pick the closest shape. We will separate the app server, the agent runtime, and the model so you do not buy a GPU for the wrong job.

Your first decision

Do you need to own model serving?

Use an API for the fastest start. Rent serverless GPU for custom weights or bursty jobs. Keep a GPU warm only after the workload proves it needs one.

Open the model hosting guide
Your app
Model endpoint
GPU belongs here — not under the web app.
The core choice

Three ways to use a model

There is no universal winner. Traffic shape, control, and engineering time change the answer.

API
01

Use a model API

Best first move for most apps. No GPU setup, scaling, or idle hardware.

  • Best for prototypes and variable usage
  • You manage prompts and product code
  • Tradeoff less infrastructure control
Start here unless you know why not
GPUϟ
02

Rent serverless GPU

Bring custom weights and pay for active compute. Cold starts and minimums still matter.

  • Best for bursty jobs and custom models
  • You manage image and inference code
  • Tradeoff startup delay and limits
Good bridge to self-hosting
24/7
03

Keep a GPU running

Predictable capacity and maximum control — with a bill that continues while idle.

  • Best for steady, measured workloads
  • You manage the complete serving stack
  • Tradeoff cost and operations
Earn this complexity with data
EstimatedPut your traffic into the calculatorCompare usage-based API, serverless GPU, and always-on GPU monthly cost.
Calculate my cost →
VRAM, not vibes

What GPU fits the model?

These are conservative first-test targets for 4-bit weights, short context, and one active sequence — capacity estimates, not performance benchmarks.

Smaller modelMore VRAM →
Qwen3 8B8B parameters
16 GB
T4from $0.59/hr
Mistral Small 24B24B parameters
24 GB
L4from $0.39/hr
Qwen3 32B32B parameters
32 GB
RTX 5090from $0.29/hr
Llama 3.3 70B70B parameters
48 GB
A40 / A6000from $0.44/hr
Sourced + estimated

Model sizes come from official model cards. Raw weight math and first GPU targets are estimates with headroom. Context, concurrency, runtime, and quantization can require more.

Read the sizing method →
Useful before you deploy

Tools, not another wall of text

Every tool exposes its inputs, claim type, source date, and limitations.

Deep research

Go deeper only when you need to

The short visual answer comes first. These longer pages keep the source trail, assumptions, and implementation detail available for search, GEO, and serious buyers.

ai hosting Serverless GPU pricing in 2026: August 10 rates and deployment matrix An August 10, 2026 source-backed GPU deployment price matrix that separates Pods, serverless workers, managed inference, and GPU VMs instead of treating unlike rates as one market. 2026-08-10 · Read the evidence → ai hosting What GPU do you need to run Llama 70B? VRAM guide for self-hosting open models A source-backed Llama 70B VRAM guide covering 4-bit, 8-bit, and BF16 memory estimates, quantization caveats, and practical current cloud GPU capacity options. 2026-07-29 · Read the evidence → ai hosting Best hosting for AI agents on a budget (June 2026): choose by workload, not by AI branding A June 26, 2026 HostFleet refresh on budget AI agent hosting, split by scheduled jobs, always-on workers, and small self-hosted stacks. 2026-07-10 · Read the evidence → deploy ai apps Where to deploy your Lovable, Bolt, or v0 app: a decision guide (April 2026) Lovable, Bolt.new, and v0 all generate working apps — but the hosting story is different for each one. A feature-by-feature guide to what each platform supports and where your code actually runs. 2026-04-21 · Read the evidence → deploy ai apps What breaks when AI-generated apps hit production: documented footguns A field guide to the failure modes you'll hit when you ship an app written by Lovable, Bolt, v0, or Cursor. Synthesized from public GitHub issues, Reddit threads, and vendor docs — every claim linked. 2026-04-21 · Read the evidence → ai hosting What it costs to run an AI side project on a VPS for 30 days (June 24, 2026): honest budget ranges A practical June 24, 2026 guide to what an AI side project really costs on a VPS for 30 days, using current Hostinger, DigitalOcean, and Hetzner pricing plus explicit workload assumptions. 2026-06-24 · Read the evidence →