Language models
vLLM
vLLM inference server. OpenAI-compatible /v1/chat/completions endpoint.
Details
Image
vllm/vllm-openai:latest
Ports
8000/http, 22/tcp
Setup
3-4 minutes
Difficulty
advanced
Features
OpenAI-compatible API, PagedAttention, Continuous batching
Compatible GPUs
Start
locked
Holders of 10,000 $GPU start against the treasury-funded cluster. Sessions are capped at 1 hour. Everyone else can read specs.
Runtime 1 hour (fixed)
App vLLM
Image vllm/vllm-openai:latest
Ports 8000/http,22/tcp
Connect a wallet on Robinhood Chain.