Preprint reports faster CPU text generation from a hybrid language model
In a one-seed comparison, the 160-million-parameter model reached 1.76 times the throughput of a matched all-attention twin at 2,048 tokens, while downstream scores were essentially tied.