For the complete documentation index, see llms.txt. This page is also available as Markdown.

Examples

Short, practical recipes for the most common REE workflows.

Common Workflow Examples

Try some of these ready-to-use TUI configurations for common workflows like test runs, production inference, prompt files, and more.

Minimal: Test Model

Use a small test model to verify REE is working.

Note that test models have random weights and will produce nonsensical output. This is expected.

In the TUI, fill in the following parameters:

  • Model Name: hf-internal-testing/tiny-random-LlamaForCausalLM

  • Prompt Text: Hello world

  • Max New Tokens: 24

  • Press r to run.

Production: Reproducible Inference with a Real Model

  • Model Name: Qwen/Qwen3-0.6B

  • Prompt Text: Explain quantum entanglement in simple terms.

  • Max New Tokens: 256

  • Extra Args: --operation-set reproducible --temperature 0.7 --top-p 0.9

  • Press r to run.

Using a Prompt File

Save your prompt to a .JSONL file:

In the TUI:

  • Model Name: Qwen/Qwen3-0.6B

  • Prompt Text: (leave blank)

  • Prompt File: /path/to/your/prompt.JSONL

  • Max New Tokens: 128

  • Extra Args:

  • Press r to run.

Short-Circuiting (Reasoning Models)

Short-circuiting forces the model to exit a generation phase early by injecting a specific token at a given step. This is useful for reasoning models (e.g., Qwen3) that have thinking/end-thinking phases, where you want to limit the token budget spent on "thinking."

Both --short-circuit-length and --short-circuit-token must be provided together in Extra Args.

  • Model Name: Qwen/Qwen3-14B

  • Prompt Text: Solve this math problem.

  • Max New Tokens: 300

  • Extra Args:

  • Press r to run.

v0.4.0: To turn thinking off entirely (rather than truncating it), use enable_thinking=False on InferenceSession.complete(messages=...) in the SDK.

Short-circuiting remains useful when you want a bounded thinking budget on CLI/TUI runs.

Large Models with Pipeline Parallelism

Split a large model across multiple GPUs using pipeline parallelism. This is the mode to use for models above ~32B parameters.

  • Model Name: Qwen/Qwen2.5-72B-Instruct

  • Prompt Text: Summarize the key ideas behind pipeline parallelism.

  • Max New Tokens: 256

  • Partitions: 4

  • Extra Args: --operation-set reproducible

  • Press r to run.

Set Partitions to match how many GPUs you want to split the model across.

If driving REE from the CLI instead of the TUI, use --n-partitions <N>.

SDK: Tool Definitions with InferenceSession

This example shows how to pass tool definitions into an SDK inference session. REE forwards the tool definitions to the model's chat template when the tokenizer supports one.

First prepare a task directory using the CLI or TUI. Then use that prepared task directory with InferenceSession:

REE does not execute the calculator function or parse the model output into a tool-call object. For reproducibility, record the exact tool output that your application provides back to the model.

SDK: Disable Thinking (Qwen3)

Requires a prepared task directory (see SDK Hello World for the prepare & session pattern).

Validating a Receipt

Validation ensures that a receipt remains internally consistent and that its hashes are untampered and uncorrupted, without requiring re-computation.

After a successful run, switch the TUI to validate mode:

  • Subcommand: validate

  • Receipt Path: Paste the path to your receipt JSON file (e.g., ~/.cache/gensyn/Qwen--Qwen3-0.6B/.../metadata/receipt_20260311_155048.json)

  • Press r to run.

Verifying a Receipt

Verification re-runs the entire inference pipeline and comparing the results with the receipt to ensure reproducibility.

To prove a receipt is reproducible by re-running the full inference pipeline:

  • Subcommand: verify

  • Receipt Path: Paste the path to the receipt JSON file

  • Press r to run.

REE will re-execute the computation and compare the output against the receipt. This is slower than validate since it runs the full pipeline, but it's the strongest proof that the result is reproducible.

Last updated