Get Started
This guide will help you get started with Fine-Tuning models. Our Fine-Tuning feature enables you to train popular LLMs on any publicly available dataset hosted on Hugging Face as well as custom datasets. Follow the steps below to launch and manage a fine-tuning job.
Step 1: Select Base Model
You can choose from a given set of Open Source base models to fine tune.
Step 2: Add Dataset
You can either specify the full Hugging Face dataset path you’d like to use for fine‑tuning, or upload your own dataset under the Custom tab. There are two types of datasets:
- Training: Dataset to train your model against
- Validation (optional): Dataset to validate the model against unseen data that is outside of the training loop.
Requirements
| Type | Description |
|---|---|
| File Type | .jsonl (JSON Lines) or .csv |
| Columns | Exactly two columns: prompt and completion (see schema). Any additional columns will cause training to fail. |
Source
- Hugging Face dataset: Use full Hugging Face path.
- Custom Dataset: You can also choose your uploaded dataset.
For details on the required dataset schema, see the Datasets guide.
Step 3: Output Job Details
Provide the following:
- Suffix: A custom identifier added to your Job ID for easier tracking.
- Registry: Choose the model registry where the fine-tuned model will be saved for future deployment.
- Seed: A random seed for reproducibility. This field is required — you must provide a value (there is no automatic default). The same seed is also used when the dataset is auto-split into training and validation subsets (see Dataset Requirements).
Step 4: Configure Hyperparameters
Each hyperparameter comes with a default value, which you can adjust:
| Hyperparameter | Description |
|---|---|
| Batch size | Number of training samples used in one forward/backward pass. |
| Learning Rate | Controls the step size during optimisation. |
| Number of Epochs | Total number of times the model will iterate over the full training set. |
| Warmup Ratio | Proportion of training steps to gradually increase the learning rate. |
| Weight Decay | Regularisation to prevent overfitting by penalising large weights. |
Step 5: Set LoRA Parameters
| Parameter | Description |
|---|---|
| LoRA Rank | Dimension of the low-rank decomposition |
| LoRA Alpha | Scaling factor for LoRA updates. |
| LoRA Dropout | Dropout rate applied to LoRA layers. |
Step 6: Launch the Job
Once a job is launched, it will automatically start the training run. These runs can vary in time, depending on the configurations set.
At the same time, a placeholder model is created in the registry you selected in Step 3, named after the base model plus your job suffix. This is the equivalent of the Create New Model step you would otherwise perform manually in the Model Registry — the model exists immediately, but it has no versions until you register a checkpoint (see Step 7).
You can track the progress of your job by monitoring its status. A job can transition through the following states:
| State | Description |
|---|---|
Initializing | The training environment is being prepared (image pull, resource allocation). |
Running | Training is actively in progress. |
Completed | Training finished successfully and checkpoints are available. |
Stopped | The job was manually stopped by the user. |
Deleting | The job's resources are being torn down or have been fully removed. |
Step 7: Post‑job Artefacts
Once the job reaches Completed:
- Navigate to the job's details view to see the list of checkpoints produced by the run. Training saves one checkpoint per epoch, so a three‑epoch job leaves you three checkpoints to choose between.
- For each checkpoint, you’ll see:
- Training Loss: The error (or loss) the model incurred while learning from the training dataset. A decreasing training loss generally indicates the model is learning, but if it’s very low compared to validation loss, it might be overfitting.
- Validation Loss: The error computed using the validation dataset (unseen during training) is an indicator of how well the model will perform on real-world or unseen data.
Registering a checkpoint
Checkpoints are not registered for you. Compare the training and validation loss across epochs and register the checkpoint that performed best, so you stay in control of which epoch is promoted to a deployable model version and can leave behind any that overfit. If you want to compare epochs side by side, you can register more than one checkpoint from the same job.
Registering a checkpoint adds it as a new version of the placeholder model created when the job was launched. Inside that model, the version progresses through Uploading → Synchronizing → Available. Once it is Available, you can deploy it to an Endpoint, either from the Model Registry version list or directly from the checkpoint in the job details view.
Registered model weights appear under Model Registry > Model > Versions. Learn more about Model Registry.