Loading Open Internet
    I built a custom 2-Bit Ternary Inference Engine from scratch in Rust + native PyTorch QAT. I'm running GPT-2 XL (1.5B) entirely offline on a Surface Pro 7 at 115 tokens/sec.