PostHog has released Jeeves, an experimental open-source model for tasks that require choosing among defined options and returning calibrated probabilities. The project adds an explicit reasoning stage to a Jev-style decision model and publishes weights, serving code, an SDK and a reproducible training pipeline.

Jev-like systems are designed to assign probabilities to possible decisions, but their accuracy can lag behind general reasoning models. Developers often compensate by sending difficult cases to a second model. Jeeves instead trains a Qwen3.5-9B model with low-rank adaptation and a separate pointer head so it can reason before scoring the available choices.

The project uses supervised fine-tuning followed by CISPO reinforcement learning. After generating a reasoning chain, Jeeves places a query representation at a dedicated decision token and compares it with representations attached to each option. A softmax converts those scores into probabilities, with a temperature fitted on development data to improve calibration.

PostHog reports that the approach improves results on out-of-domain tasks and outperforms Jev on the public hard tier of JevBench. The published comparison is limited to the benchmark’s 231 public easy, standard and hard items; it does not include a sealed judge tier. The repository also notes that some comparison figures come from a different-sized Kev model because no Kev-9B Jeeves benchmark result has been published. Those qualifications matter because the numbers are project-reported rather than an independent evaluation.

Reasoning produced a measurable gain in the team’s own testing. The same checkpoint scored 0.804 on a 2,962-item test split without thinking and 0.840 when reasoning was enabled. The team selected training step 402 because later reinforcement-learning steps made the output probabilities too confident on the training pool.

Jeeves also includes a speculative diffusion drafter intended to speed generation. Inspired by Orthrus, it is adapted to work with the Gated DeltaNet layers in Qwen3.5 by allowing mask tokens to cross-attend to post-convolution keys and values. The default configuration uses a block size of four to keep the cost manageable when multiple questions are processed together.

Developers can download the released weights and run a local server, or fuse a trained checkpoint into a standalone model. The included Python client is presented as a drop-in replacement for Jev’s typed SDK, connects to a local endpoint by default and requires no API key. Data-preparation scripts fetch public datasets at pinned revisions, leaving each under its original license.

The release gives researchers a concrete implementation to test, but broader claims will depend on independent benchmarks, sealed evaluations and performance across real applications. For now, Jeeves is best understood as a published model and recipe for exploring whether structured reasoning can improve probabilistic decision systems.