Skip to main content

Overview

RAIC Inference Endpoints is the easiest way to deploy machine learning models as scalable API endpoints. It provides you with the flexibility to choose from a variety of pre-trained models, specify your compute requirements, and deploy them in specific locations to minimize latency and optimize performance.

It allows you to select two types of models:

  • Pre-trained models — Choose from Llama 3.1, Llama 3.2, Mistral, Qwen 2.5, and many more.
  • Custom models — Bring your own model weights, or deploy a model produced by Fine-Tuning. Custom models are managed in the Model Registry and, once a version reaches Available, can be deployed to an endpoint exactly like a pre-trained model. See Deploying a custom model.

Deploying a custom model​

Custom models reach the registry in one of two ways:

  • Your own weights — create a model in the Model Registry, then upload a HuggingFace-format model directory with the CLI. See the Model Registry guide.
  • Fine-tuning output — each checkpoint you register from a fine-tuning job becomes a model version. See the Fine-Tuning guide.

The deployment steps are the same for both:

  1. Confirm the model version shows status Available in the Model Registry. New versions progress through Uploaded → Synchronizing → Available while the weights are cached into the target location. A version cannot be deployed until it reaches Available.
  2. Deploy the version, either from the version list in the Model Registry or by creating a new endpoint and selecting your custom model.
  3. Choose one of the GPU types you selected when creating the model, plus GPU count, location, and replica settings.

Requirements​

  • Models must be in HuggingFace format. They are served with vLLM, so the model architecture must be supported by vLLM.
  • Upload the entire model directory, not just the weights file — vLLM needs config.json and the tokenizer files alongside the .safetensors weights.
  • The GPU types you select at model-creation time determine where weights are cached. If no matching GPU type is active in your registry's location, the version will stay at Uploaded and never become Available.
note

Supported file types are .safetensors, .json, .md, .txt, .py, .gitattributes, .yml, and LICENSE. Models that ship a SentencePiece tokenizer.model file cannot currently be uploaded, because .model is not an accepted extension.