How Fort works
Fort never writes free text. Every answer is one of your options, with a probability you can act on: take the sure answers, and send the rest to a person or a larger model.
1 · Text and options
“A quiet film that stays with you for days after the credits roll.”
- A positive
- B negative
2 · One forward pass
3 · The call
A positive 99%
It answers three kinds of question:
- Choice. One of up to 26 options, with a probability for each. Longer label lists, such as 77 bank intents, are read by name, so the answer is always a listed label.
- Yes or no. The probability of yes.
- Score. A position on an ordered scale, low to high.
Models
Two sizes, each with a 4-bit MLX version. They run on MLX (Apple silicon) and PyTorch (NVIDIA GPUs and CPUs); 4-bit: MLX only.
| Property | Fort-1 2B | Fort-1 0.8B |
|---|---|---|
| Best for | Accuracy | Speed |
| Parameters | 1.88B | 0.75B |
| Built on | Qwen3.5-2B | Qwen3.5-0.8B |
| Download, bf16 | 3.8 GB | 1.5 GB |
| Download, 4-bit MLX | 1.08 GB | 0.44 GB |
| Memory while deciding, bf16 | 4.2 GB | 1.8 GB |
| Memory while deciding, 4-bit | 1.6 GB | 0.7 GB |
| Context window | 262,144 tokens | 262,144 tokens |
| License | Apache-2.0 | Apache-2.0 |
Setting Memory measured with MLX.
Speed
Milliseconds per decision, one decision at a time on an Apple M1 Max GPU: the median, with the 95th percentile beneath it. The API models' medians include the network.
| Model | SST2 | AG News | Banking77 |
|---|---|---|---|
| Fort-1 2B | 110122 | 214258 | 79111 |
| Fort-1 2B, 4-bit | 107114 | 210255 | 6084 |
| Fort-1 0.8B | 2830 | 4448 | 5182 |
| Fort-1 0.8B, 4-bit | 3133 | 4144 | 3349 |
| Jev (jev-1.13)TypeSafe API | 376 | 381 | 389 |
| Claude Sonnet 5Anthropic API | 2,076 | 2,014 | 1,995 |
| GPT-5-miniOpenAI API | 1,385 | 1,312 | 1,312 |
Setting Fort: Apple M1 Max GPU, MLX, one decision at a time. API models: jevbench's medians.
Get started
The teximal package runs both sizes from the command line or Python, on MLX or PyTorch.
pip install teximal
teximal run fort-1-2b "I was charged twice this month." --options billing,technical,sales
teximal serve fort-1-2bteximal serve answers Jev's request format (OpenRouter's Decisions API), so switching from Jev is a change of URL.
- Weights. teximal/fort-1-2b, teximal/fort-1-2b-mlx-4bit, teximal/fort-1-0.8b, teximal/fort-1-0.8b-mlx-4bit, in the Fort-1 collection on Hugging Face.
- Package. teximal on PyPI.
- Code and docs. teximal/teximal-cli on GitHub.
- License. Apache-2.0, with the Qwen3.5 base model's notice.
Results on jevbench
jevbench (September 22, 2026) is an independent classifier benchmark that gives every model the same label definitions. Three tasks, 500 examples each: SST2 (movie-review sentiment), AG News (news topics) and Banking77 (77 bank customer intents). Accuracy in percent.
| Model | SST2 | AG News | Banking77 |
|---|---|---|---|
| Fort-1 2BTeximal, open | 93.8 | 86.8 | 76.6* |
| Fort-1 2B, 4-bitTeximal, open | 93.2 | 86.4 | 77.0* |
| Fort-1 0.8BTeximal, open | 93.2 | 87.2 | 74.8* |
| Fort-1 0.8B, 4-bitTeximal, open | 93.8 | 84.8 | 73.0* |
| Jev (jev-1.13)TypeSafe, API | 95.4 | 84.3† | 76.4 |
| Claude Sonnet 5Anthropic, API | 95.6 | 89.6 | 77.4 |
| GPT-5-miniOpenAI, API | 95.0 | 80.2 | 73.6 |
Training Fort never trained on SST2 or AG News, only on similar tasks: IMDB review sentences and HuffPost headlines labeled by Teximal Lab.
*Banking77 Adapted, not zero-shot: Fort trained on bank messages that Teximal Lab wrote from the 77 intent names, never on real Banking77 messages (one scored message also appears inside a training message). Scored with jevbench's instruction, as every model was; with a banking instruction, Fort-1 2B scores 77.8 and Fort-1 0.8B scores 74.6.
†AG News Jev scores 85.8 with jevbench's newer label definitions.
Calibration
When Fort says 90%, it should be right about 90% of the time. Calibration error (ECE) measures the gap; lower is better. “Fitted” is after one temperature is fitted on a few hundred labeled examples of the task, which teximal eval does.
| Model | SST2 | AG News | Banking77 |
|---|---|---|---|
| Fort-1 2BAs shipped | 0.012 | 0.092 | 0.182 |
| Fort-1 2BFitted | 0.018 | 0.060 | 0.035 |
| Fort-1 0.8BAs shipped | 0.013 | 0.063 | 0.178 |
| Fort-1 0.8BFitted | 0.020 | 0.062 | 0.038 |
| Jev (jev-1.13)TypeSafe, API | 0.026 | 0.112 | 0.125 |
Setting jevbench, 500 examples per task; fitting used 300 held-out examples per task. Claude Sonnet 5 and GPT-5-mini return no probabilities, so their calibration cannot be measured.
More results
| Test | Fort-1 2B | Fort-1 0.8B |
|---|---|---|
| Decision exam: 600 tasks written by Teximal Lab, never trained on | 80.3% (ECE 0.035) | 67.5% (ECE 0.081) |
| People's 1-to-5 ratings, held out (urgency, frustration, sentiment strength, toxicity, stars) | 52 to 76% exact, 92 to 98% within one level | 50 to 74% exact, 88 to 98% within one level |
| General knowledge, 4 options | 92.5% | 91.2% |
| Reading comprehension, 4 / 16 options | 96.8% / 93.5% | 93.5% / 88.0% |
| Assistant requests in 51 languages (MASSIVE) | 78.1% | 70.7% |
| Options shuffled or renamed, typos, lower case, a distracting sentence | At most 1.7 points lower | At most 1.3 points lower |
| A text that claims the answer or orders a letter | At most 1% of decisions moved | No decision moved |
| Long context (window: 262,144 tokens) | 90 to 95% with the deciding text inside about 32,000 tokens; 3 of 3 at about 129,000 | 88 to 100% with the deciding text inside about 2,000 tokens; at about 32,000, 95% when it comes last and 55 to 65% elsewhere |
Use and limits
Routing, triage, moderation queues, tagging and form checks: decisions over your own options, where a calibrated confidence decides what to automate and what to send to a person. Not for decisions with legal or similarly significant effects on people without human review.
Limits
- Not a math model: it answers in one step, without working.
- Mostly English. Tweet sentiment in four African languages is weak: AfriSenti 49.9% for Fort-1 2B and 49.4% for Fort-1 0.8B.
- Text only.
- Long label lists read by name, such as Banking77, are overconfident until a temperature is fitted.
- Instructions planted in a text rarely move its answer, but treat untrusted text as data.
Training
Fine-tuned from Qwen3.5-2B and Qwen3.5-0.8B with LoRA. The targets blend the probabilities of a larger teacher, Qwen3.6-35B-A3B (Apache-2.0), with people's labels from Teximal Lab, collected with hidden checks.
Data: MASSIVE, SQuAD 2.0, CLINC150, DBpedia14, SNLI, MultiNLI, Amazon reviews, Civil Comments, the HuffPost News Category Dataset and IMDB reviews, plus critic-style sentences and bank customer messages written from scratch by Teximal Lab.
