Market/vLLM
Language models

vLLM

vLLM inference server. OpenAI-compatible /v1/chat/completions endpoint.

Details

Image
vllm/vllm-openai:latest
Ports
8000/http, 22/tcp
Setup
3-4 minutes
Difficulty
advanced
Features
OpenAI-compatible API, PagedAttention, Continuous batching
Compatible GPUs
RTX 4090
24GB GDDR6X
budget$0.44 / hr
RTX A6000
48GB GDDR6
standard$0.49 / hr
L40
48GB GDDR6
standard$0.79 / hr
A100 40GB
40GB HBM2e
premium$1.29 / hr
A100 80GB
80GB HBM2e
premium$1.89 / hr
H100 80GB
80GB HBM3
enterprise$2.99 / hr
Start
locked

Holders of 10,000 $GPU start against the treasury-funded cluster. Sessions are capped at 1 hour. Everyone else can read specs.

Runtime 1 hour (fixed)
App vLLM
Image vllm/vllm-openai:latest
Ports 8000/http,22/tcp

Connect a wallet on Robinhood Chain.