Blown Away Yet Again

Walking Astra to the playground when I'm notified that DSv4 Flash Vision now works

The new Anthropic Fable 5.1 release is outrageous. One of the things that I always found perplexing is that various models don't work on various bits of hardware despite the fact that writing software is supposedly obsolete. Well, I believe that software writing as coding is not a job anymore. But one thing I couldn't square was why hardware was so hard. Sporadically someone would write a driver for something but for the most part companies like Tenstorrent or AMD simply don't have inference running on most of their hardware while Nvidia has a moat that's entirely software.

This particular mismatch between a belief I had and a fact I was looking at constantly bothered me. Where was I wrong? It turns out I was wrong in the sense that it was really just a few steps too early. This was going to happen. It just hadn't already. My test for this was always whether I could have a frontier AI with its native harness get me a model running that no one had yet implemented.

And today it happened. There's no way to get TP=2 DeepSeek V4 Flash Vision working on 2x RTX 6000 PRO Blackwell Workstation Edition Max-Q with DSpark. No one had implemented this, and in fact if you tried running the upstream work that ran on B300, it wouldn't work with TP=2. I fully expected that the prompt I fired off while I prepared Astra for the playground would have no effect and I'd come back to be frustrated yet again.

I'd spent some half an hour on the problem already and we had the model running normally. We also had drafting working on the prior non-vision version of this model due to a lot of work upstream including some by my friend Alex Bilichenko. But 15 minutes after my prompt as I was walking out the door of the apartment building I saw that this last bastion of software immune to LLMs had fallen.

And the model has been running in prod for me ever since.