Skip to main content

Overview

GenerationConfig controls how the model generates text, including sampling strategies, length constraints, and decoding methods. These parameters can be set globally on the model or passed per generation request.

Loading Configuration

Core Parameters

Temperature

float
default:"1.0"
Controls randomness in generation:
  • 0.0 - 0.01: Nearly deterministic (use top_k=1 instead)
  • 0.7 - 0.9: Balanced creativity and coherence
  • 1.0+: More random and creative
Note: Qwen recommends tuning top_p instead of temperature.

Top-p (Nucleus Sampling)

float
default:"0.95"
Nucleus sampling probability threshold. Model considers only tokens whose cumulative probability exceeds top_p.
  • 0.1: Very conservative, deterministic
  • 0.7 - 0.9: Balanced (recommended)
  • 0.95 - 1.0: More diverse outputs

Top-k Sampling

int
default:"None"
Limits sampling to the k most likely tokens:
  • 1: Greedy decoding (deterministic)
  • 10-50: Conservative sampling
  • 50-100: More diverse sampling
Setting top_k=1 is equivalent to greedy decoding.

Length Control

Max Length

int
default:"8192"
Maximum total sequence length (input + output tokens). Generation stops when this limit is reached.

Max New Tokens

int
default:"None"
Maximum number of tokens to generate (excluding input). Takes precedence over max_length.

Min Length

int
default:"0"
Minimum total sequence length. Model will not generate EOS token before reaching this length.

Min New Tokens

int
default:"None"
Minimum number of new tokens to generate

Stopping Criteria

Stop Strings

Stop generation when specific sequences are encountered:

EOS Token

int | list[int]
default:"None"
Token ID(s) that trigger end of generation. For Qwen:
  • 151643: Default EOS token ID
  • <|im_end|>: ChatML format end token

Pad Token

int
default:"None"
Token ID used for padding sequences in batched generation. For Qwen, typically set to tokenizer.eod_id.

Repetition Control

Repetition Penalty

float
default:"1.0"
Penalty for repeating tokens:
  • 1.0: No penalty
  • 1.1 - 1.3: Mild discouragement of repetition
  • > 1.5: Strong penalty (may harm coherence)

No Repeat N-gram Size

int
default:"0"
Prevent repeating n-grams of this size. Set to 0 to disable.

Num Beams

int
default:"1"
Number of beams for beam search:
  • 1: No beam search (faster)
  • 4-10: Beam search (slower but potentially higher quality)

Advanced Parameters

Do Sample

bool
default:"True"
Whether to use sampling (True) or greedy/beam search (False)

Early Stopping

bool
default:"False"
Stop beam search when all beams reach EOS token

Use Cache

bool
default:"True"
Use KV cache for faster generation. Should be True for inference.

Complete Configuration Example

Modifying at Runtime

Factual/Deterministic

Balanced

Creative