A 58M-parameter French reasoning LLM trained entirely from scratch — on a gaming laptop's RTX 4060. It works through problems in a <think> scratchpad before answering, and beats public French GPT-2s twice its size on language quality and arithmetic. Qwen-style architecture, Muon optimizer, custom digit-split BPE tokenizer, full pretrain → midtrain → SFT → GRPO/RLAIF pipeline.
view source ↗