MoE is probably good enough.
I wanted a denser kind of sparsity. Not “this token sees 2 of 64 experts,” but a model that stays fully capable without paying full price on every forward pass. There is a romantic version of this where every parameter is always a little bit awake, just cheaply, and nobody has to gamble on a router.
I did not find a version of that story I would bet a training run on.
Mixture-of-experts has obvious sins. The router is a single point of stupidity. Experts collapse. Load balancing turns into a second loss function you babysit. Some tokens get the B-team, and you only notice in the evals you did not think to run. Stare at it long enough and you can convince yourself that a dense model with the same total parameter count would have been cleaner, more even, more “actually intelligent.”
Maybe. I also ship things.
A production model has to be smart enough, fast enough, and cheap enough to call from a phone, from a batch job, or from a tool loop that might run 40 times before breakfast. MoE is a slightly ugly way to own more capacity than you pay for on any one token. That is a very old systems trick. Caches do it. Indexes do it. Anycast and every CDN on earth do it. You do not keep the whole world hot. You keep a fast way to reach the part you need.
The bill does not disappear. Every expert still has to sit in memory somewhere, and serving gets fiddlier. But memory is a cheaper problem than paying full compute for every parameter on every token.
Dense models feel more honest. They also feel like a luxury belief. If you have the money and the energy budget to keep everything awake, congratulations. Most of the useful work I see looks more like this: get a good enough specialist onto the token, do not blow the latency budget, and leave room for retrieval and tools to do the factual work.
I am not saying the research is done. Better routers would help. Better training, so experts do not become decorative, would help. I would like a world where sparsity is a property of the computation, not a pile of expert islands and a traffic cop. Until that world shows up in a checkpoint I can run, MoE is the available compromise.
There is also a taste issue. People treat architectural purity as a moral stance: dense is “real,” MoE is a hack. Transformers were a hack that worked. Every production system I have loved was a hack that worked, then got cleaned up just enough that you could sleep.
If MoE leaves a little intelligence on the table, fine. I will take the intelligence I can afford to serve and spend the savings on memory, tools, and evals against the actual job. Those have been the bigger leaks in every product I have touched.
“Probably good enough” is not a slogan. It is the sentence you say after staring at a prettier alternative long enough to notice it does not have a ship date.
Justin DaCosta builds training systems for Amazon Nova and ships iOS apps on the side.