VRAM is the constraint, not the core count
For anything model-shaped, the question is whether the weights fit in memory. If they do not, nothing else about the card matters — the job will not run. Rough working figures for LLMs:
| Model size | Inference, 8-bit | Inference, 16-bit | LoRA fine-tune | Full fine-tune |
|---|---|---|---|---|
| 7B | ~8 GB | ~16 GB | ~24 GB | ~80 GB |
| 13B | ~14 GB | ~28 GB | ~40 GB | Multi-GPU |
| 70B | ~70 GB | Multi-GPU | Multi-GPU | Multi-GPU |
Add headroom for the KV cache, which grows with context length and batch size and is what actually causes the out-of-memory error most people hit first.
Dedicated monthly vs hourly cloud GPUs
Hourly instances are the right tool for an experiment: spin up, run, destroy, pay for ninety minutes. They are the wrong tool for a job that runs continuously, because the hourly rate assumes you will not.
The crossover is usually somewhere around 200–300 hours a month. Past that, a dedicated card is typically a fraction of the equivalent hourly spend, and it does not stop being yours when a region runs out of capacity.
Why host GPUs in India
- Data residency. Training data that cannot leave the country rules out most overseas GPU clouds outright.
- Latency for inference. If you are serving a model to Indian users, 8 ms to Mumbai beats 250 ms to Virginia on every single request.
- Rupees and a GST invoice. No forex markup, no reverse-charge paperwork, input credit claimable as usual.
- Someone to call. A GPU that has fallen off the PCIe bus at 2am is a phone call here, not a ticket in a queue.
What we install before handover
Ubuntu LTS or the distribution you name, the matching NVIDIA driver, the CUDA toolkit version your framework wants, and — if you ask — Docker with the NVIDIA container toolkit. We run nvidia-smi and a short burn-in before handing over the credentials, so the first thing you do is your own work rather than driver archaeology.
Not sure what you need?
Tell us the model, the batch size and whether it is training or serving. We will size it honestly, including saying when a ₹7,299$82.94/mo CPU box would do the job — plenty of workloads people assume need a GPU do not. And if we cannot source the card you need at a price worth paying, we will say that rather than take the order.