Fine-tune
Fine-tune means one thing: how a model runs on this computer - speed, reading room, memory. Everything on this page is optional. Empty fields mean the automatics decide, and the value they would use is shown in the box. Set a number and it applies; put a row back to Auto and the automatics are back.
Fine-tune a model
Every downloaded model has a Fine-tune button on its row on the Offline Models page. The dialog opens on what the model runs at now, and where that came from: Automatic, Measured here, or Set by you. Everything in it is for this computer only.
Measure on this computer
The main event. One button loads the model a few times with different setups and times each one on your own hardware. It takes about three minutes, and chats pause while it runs; you can stop it at any point. Nothing runs without the click.
Afterwards a slider from Faster answers to More room picks between the measured setups. The card under it says what each one means in plain terms: the context in pages in view at once, the speed in words a second, then the detail (reading speed, the expert split, the speed-up file, the cache). The automatic pick is marked on the scale. More room lets the AI hold more of a long chat or document at once; faster answers come from less room. Measure again reruns it.
The measurement teaches the automatics several things, and the dialog says what it decided, one line each. It times the Compact context memory beside Standard, and Auto uses Compact on this machine only when that came back clean. When fewer expert layers in main memory ran at least five percent faster than the automatic pick at that context, Auto uses the measured split at the next load. On a graphics card it also tries bigger batches and keeps them per model where they read faster; a mixture-of-experts model whose expert layers sit in main memory keeps a bigger batch where long prompts read at least 15% faster. On a computer whose only graphics is integrated, it tries the processor against the integrated graphics and keeps the processor when it reads and writes at least 10% faster.
The slider measures the reading room the model runs at, so the status line shows Automatic's speed as soon as the measurement finishes. Sliding back to the Auto stop returns to Automatic, and Save clears the entry rather than pinning a number you did not mean to set.
Save Changes keeps the pick and closes the dialog; the outcome shows in the activity card at the bottom right: reloaded now when the model is the one running, otherwise applied at its next load.
Set it yourself
A drawer at the bottom of the dialog, closed until you open it (or already open if you set values by hand before). Each control is one line, with the explanation in its hover text:
- Context size - tokens of reading room. More than the automatics measure fits may load slowly or fail; the model's trained limit still caps it.
- Expert layers in main memory - mixture-of-experts models only. 0 puts everything on the card; fewer than the automatics pick means the rest must fit on the card.
- Context memory - Standard, Compact or Auto. Compact stores the context at half the memory, so a bigger reading room fits. Auto uses Compact only where a measurement here came back clean.
- Leave the speed-up file out - for a model that has one, to measure the difference yourself.
- Back to automatic puts every row back.
Fine-tune this computer
At the bottom of Settings → Engines:
- Worker threads - for the next model load.
- Graphics cards - on a computer with several: Auto (the biggest card alone when the model fits there, pooled only when it does not), Biggest card only, or Pool all cards (more memory and a larger reading room, slower). Measure on this computer times the pooled case too, so the row shows real numbers.
- Helper models - Automatic, or Keep on the processor; see Engines.
Reply style
The sliders for how replies are written are not a fine-tune of the computer, so they live in their own Settings section, Reply style for every AI: creativity (temperature), word variety (top-p), the rare-word floor (min-p) and the repetition brake (repeat penalty). Every AI inherits them unless it sets its own.
An AI's form has the same four sliders under Reply style for this AI. Empty fields mean the default, shown in the box; a value applies to new replies immediately - two AIs on the same model can answer with different temperaments. Works for offline and online models.
Which value wins
An AI's own reply style, then this computer's from Settings, then the model default.