For the complete documentation index, see llms.txt. This page is also available as Markdown.

Troubleshooting

Common errors, their causes, and how to fix them.

Common Errors

Quick fixes for Docker, CLI, and compilation issues you might hit while using REE.

Error
Cause
Fix

PermissionError: [Errno 13] Permission denied: '/gensyn'

Container runs as non-root gensyn user; can't write to root-owned paths

Use /tmp/ paths for ephemeral runs, or mount a volume with -v

one of the arguments --tasks-root --task-dir is required

Missing required output directory argument

Add --task-dir /tmp/task or --tasks-root /tmp/tasks

one of the arguments --prompt-text --prompt-file is required

Missing prompt input

Add --prompt-text "your prompt" or --prompt-file path.jsonl

argument command: invalid choice: 'bash'

Trying to launch a shell but entrypoint is locked to gensyn-sdk

Use --entrypoint bash to override: docker run -it --entrypoint bash ree

Gibberish / nonsensical output

Using hf-internal-testing/tiny-random-LlamaForCausalLM which has random untrained weights

Expected behavior for test models; use a real model for meaningful output

Shell hangs after pasting command

Trailing \\ on the last line of a command

Remove the backslash from the final line

--n-partitions run fails on a single-GPU machine

Pipeline parallelism requires multiple GPUs; a partition count above 1 can't be satisfied with only one device.

Omit --n-partitions or set it to 1. Single-GPU hosts can only run models that fit in one GPU's memory.

ValueError: 'tools' can only be provided with 'messages'.

Tool definitions were passed with a plain prompt string in InferenceSession.complete().

Use messages=[...] instead of prompt=... when providing tools.

Tool definitions appear to be ignored

The selected tokenizer may not have a chat template, or the model may not be trained to use tools.

Use a chat/instruct model with a tool-aware chat template and verify the model emits the expected tool-call format.

No parsed tool call is returned

REE provides basic tool-definition support, not a full tool executor or agent loop.

Parse the model output in your application, execute tools yourself, and pass follow-up context back into the session.

Tool-Call Workflows

REE does not execute external tools. If a model emits tool-call-shaped output, your application must parse it, run the tool, and continue the conversation.

For reproducible workflows, record the exact tool output passed back to the model. Live APIs, web pages, databases, and search results may change between runs.

Error
Cause
Fix

ValueError: 'tools' can only be provided with 'messages'.

tools passed with prompt=

Use messages=[...]

No parsed tool call in SDK result

REE does not execute or return structured tool calls

Parse result.text, run tools in application code

enable_thinking seems ignored

Plain prompt or unsupported tokenizer

Use messages and a thinking-capable model such as Qwen3

Cannot set thinking on/off from TUI

enable_thinking is SDK-only

Use InferenceSession.complete(..., enable_thinking=False) or short-circuit flags on CLI

Compiler Trace Warnings

When running REE, you may see trace warnings and verbose compiler output in your terminal. These are expected and can be safely ignored. They originate from the ONNX export and MLIR compilation stages.

--tasks-root vs. --task-dir

If you see errors about missing artifacts, make sure you're using consistent location flags. When using --tasks-root, the task directory is automatically derived from the model name. When using --task-dir, you must point to the same directory across all operations.

CUDA Not Available

If you're running on a machine with a GPU but REE doesn't detect it, ensure you're passing the --gpus all flag to Docker:

Use --cpu-only to explicitly force CPU execution when GPU is not available or not desired.

Out-of-Memory (OOM) Issues

If you are using Docker Desktop, you may need to adjust the memory limit. Otherwise, you may attempt to run larger models (models with a higher parameter count) and encounter a failure during the model loading or checkpoint 'sharding' phase.

This typically shows up as a run:failed status with exit code 137, and the logs will show the process dying partway through "Loading checkpoint shards."

NaN Errors & Crashes with Certain Models

Some FP16 models, particularly certain Qwen 2.5 Instruct variants, may produce NaN (Not a Number) errors and crash when run in default or deterministic mode. This is a numerical stability issue: attention score calculations can overflow the FP16 value range during inference.

If you encounter this, try switching to reproducible mode (--operation-set reproducible), which handles these edge cases more gracefully. Note that even in reproducible mode, some affected models may still produce degraded output quality (repetitive text or unexpected tokens).

This is a known limitation related to the ONNX export pipeline's use of FP16 precision and is being actively addressed in future releases.

Last updated