Research

Fort-1 model card

Fort is Teximal's family of decision models. It picks one of your options for a text and tells you how sure it is. Fort-1 comes in two sizes, each with a 4-bit MLX version, all open weights under Apache-2.0.

Download on Hugging Face Docs on GitHub

28 ms
Median time per decision for Fort-1 0.8B on SST2.Apple M1 Max GPU, MLX, one decision at a time.
0.44 GB
Download size of Fort-1 0.8B at 4-bit. It uses 0.7 GB of memory while deciding.MLX.
0.012
Calibration error on SST2 for Fort-1 2B, against 0.026 for Jev.ECE, lower is better. jevbench, 500 examples.
Apache-2.0
Open weights for both sizes and their 4-bit versions.Built on Qwen3.5.

How Fort works

Fort never writes free text. Every answer is one of your options, with a probability you can act on: take the sure answers, and send the rest to a person or a larger model.

1 · Text and options

“A quiet film that stays with you for days after the credits roll.”

  • A positive
  • B negative

2 · One forward pass

  • A99%
  • B1%

3 · The call

A positive 99%

Fort-1 0.8B's scores for this review.

It answers three kinds of question:

  • Choice. One of up to 26 options, with a probability for each. Longer label lists, such as 77 bank intents, are read by name, so the answer is always a listed label.
  • Yes or no. The probability of yes.
  • Score. A position on an ordered scale, low to high.

Models

Two sizes, each with a 4-bit MLX version. They run on MLX (Apple silicon) and PyTorch (NVIDIA GPUs and CPUs); 4-bit: MLX only.

PropertyFort-1 2BFort-1 0.8B
Best forAccuracySpeed
Parameters1.88B0.75B
Built onQwen3.5-2BQwen3.5-0.8B
Download, bf163.8 GB1.5 GB
Download, 4-bit MLX1.08 GB0.44 GB
Memory while deciding, bf164.2 GB1.8 GB
Memory while deciding, 4-bit1.6 GB0.7 GB
Context window262,144 tokens262,144 tokens
LicenseApache-2.0Apache-2.0

Setting Memory measured with MLX.

Speed

Milliseconds per decision, one decision at a time on an Apple M1 Max GPU: the median, with the 95th percentile beneath it. The API models' medians include the network.

ModelSST2AG NewsBanking77
Fort-1 2B11012221425879111
Fort-1 2B, 4-bit1071142102556084
Fort-1 0.8B283044485182
Fort-1 0.8B, 4-bit313341443349
Jev (jev-1.13)TypeSafe API376381389
Claude Sonnet 5Anthropic API2,0762,0141,995
GPT-5-miniOpenAI API1,3851,3121,312

Setting Fort: Apple M1 Max GPU, MLX, one decision at a time. API models: jevbench's medians.

Get started

The teximal package runs both sizes from the command line or Python, on MLX or PyTorch.

pip install teximal
teximal run fort-1-2b "I was charged twice this month." --options billing,technical,sales
teximal serve fort-1-2b

teximal serve answers Jev's request format (OpenRouter's Decisions API), so switching from Jev is a change of URL.

Results on jevbench

jevbench (September 22, 2026) is an independent classifier benchmark that gives every model the same label definitions. Three tasks, 500 examples each: SST2 (movie-review sentiment), AG News (news topics) and Banking77 (77 bank customer intents). Accuracy in percent.

ModelSST2AG NewsBanking77
Fort-1 2BTeximal, open93.886.876.6*
Fort-1 2B, 4-bitTeximal, open93.286.477.0*
Fort-1 0.8BTeximal, open93.287.274.8*
Fort-1 0.8B, 4-bitTeximal, open93.884.873.0*
Jev (jev-1.13)TypeSafe, API95.484.3†76.4
Claude Sonnet 5Anthropic, API95.689.677.4
GPT-5-miniOpenAI, API95.080.273.6

Training Fort never trained on SST2 or AG News, only on similar tasks: IMDB review sentences and HuffPost headlines labeled by Teximal Lab.

*Banking77 Adapted, not zero-shot: Fort trained on bank messages that Teximal Lab wrote from the 77 intent names, never on real Banking77 messages (one scored message also appears inside a training message). Scored with jevbench's instruction, as every model was; with a banking instruction, Fort-1 2B scores 77.8 and Fort-1 0.8B scores 74.6.

†AG News Jev scores 85.8 with jevbench's newer label definitions.

Calibration

When Fort says 90%, it should be right about 90% of the time. Calibration error (ECE) measures the gap; lower is better. “Fitted” is after one temperature is fitted on a few hundred labeled examples of the task, which teximal eval does.

ModelSST2AG NewsBanking77
Fort-1 2BAs shipped0.0120.0920.182
Fort-1 2BFitted0.0180.0600.035
Fort-1 0.8BAs shipped0.0130.0630.178
Fort-1 0.8BFitted0.0200.0620.038
Jev (jev-1.13)TypeSafe, API0.0260.1120.125

Setting jevbench, 500 examples per task; fitting used 300 held-out examples per task. Claude Sonnet 5 and GPT-5-mini return no probabilities, so their calibration cannot be measured.

More results

TestFort-1 2BFort-1 0.8B
Decision exam: 600 tasks written by Teximal Lab, never trained on80.3% (ECE 0.035)67.5% (ECE 0.081)
People's 1-to-5 ratings, held out (urgency, frustration, sentiment strength, toxicity, stars)52 to 76% exact, 92 to 98% within one level50 to 74% exact, 88 to 98% within one level
General knowledge, 4 options92.5%91.2%
Reading comprehension, 4 / 16 options96.8% / 93.5%93.5% / 88.0%
Assistant requests in 51 languages (MASSIVE)78.1%70.7%
Options shuffled or renamed, typos, lower case, a distracting sentenceAt most 1.7 points lowerAt most 1.3 points lower
A text that claims the answer or orders a letterAt most 1% of decisions movedNo decision moved
Long context (window: 262,144 tokens)90 to 95% with the deciding text inside about 32,000 tokens; 3 of 3 at about 129,00088 to 100% with the deciding text inside about 2,000 tokens; at about 32,000, 95% when it comes last and 55 to 65% elsewhere

Use and limits

Routing, triage, moderation queues, tagging and form checks: decisions over your own options, where a calibrated confidence decides what to automate and what to send to a person. Not for decisions with legal or similarly significant effects on people without human review.

Limits

  • Not a math model: it answers in one step, without working.
  • Mostly English. Tweet sentiment in four African languages is weak: AfriSenti 49.9% for Fort-1 2B and 49.4% for Fort-1 0.8B.
  • Text only.
  • Long label lists read by name, such as Banking77, are overconfident until a temperature is fitted.
  • Instructions planted in a text rarely move its answer, but treat untrusted text as data.

Training

Fine-tuned from Qwen3.5-2B and Qwen3.5-0.8B with LoRA. The targets blend the probabilities of a larger teacher, Qwen3.6-35B-A3B (Apache-2.0), with people's labels from Teximal Lab, collected with hidden checks.

Data: MASSIVE, SQuAD 2.0, CLINC150, DBpedia14, SNLI, MultiNLI, Amazon reviews, Civil Comments, the HuffPost News Category Dataset and IMDB reviews, plus critic-style sentences and bank customer messages written from scratch by Teximal Lab.