Overview
RAIC Inference Endpoints is the easiest way to deploy machine learning models as scalable API endpoints. It provides you with the flexibility to choose from a variety of pre-trained models, specify your compute requirements, and deploy them in specific locations to minimize latency and optimize performance.
It allows you to select two types of models:
- Pre-trained models — Choose from Llama 3.1, Llama 3.2, Mistral, Qwen 2.5, and many more.
- Custom models — Bring your own model weights, or deploy a model produced by
Fine-Tuning. Custom models are managed in the Model Registry
and, once a version reaches
Available, can be deployed to an endpoint exactly like a pre-trained model. See Deploying a custom model.
Deploying a custom model
Custom models reach the registry in one of two ways:
- Your own weights — create a model in the Model Registry, then upload a HuggingFace-format model directory with the CLI. See the Model Registry guide.
- Fine-tuning output — each checkpoint you register from a fine-tuning job becomes a model version. See the Fine-Tuning guide.
The deployment steps are the same for both:
- Confirm the model version shows status
Availablein the Model Registry. New versions progress throughUploaded→Synchronizing→Availablewhile the weights are cached into the target location. A version cannot be deployed until it reachesAvailable. - Deploy the version, either from the version list in the Model Registry or by creating a new endpoint and selecting your custom model.
- Choose one of the GPU types you selected when creating the model, plus GPU count, location, and replica settings.
Requirements
- Models must be in HuggingFace format. They are served with vLLM, so the model architecture must be supported by vLLM.
- Upload the entire model directory, not just the weights file — vLLM needs
config.jsonand the tokenizer files alongside the.safetensorsweights. - The GPU types you select at model-creation time determine where weights are cached. If no
matching GPU type is active in your registry's location, the version will stay at
Uploadedand never becomeAvailable.
note
Supported file types are .safetensors, .json, .md, .txt, .py, .gitattributes,
.yml, and LICENSE. Models that ship a SentencePiece tokenizer.model file cannot
currently be uploaded, because .model is not an accepted extension.