A 110M-parameter instruction-tuned chat model, built from scratch by Dev Intern AI (Malawi). Engineering: Lance Muyawa.
No KV cache in this architecture, so replies are slower than a typical cached model, especially on CPU.