AI Updates
Lots of stuff today in the AI world.
No Need To Complain About Stack
[edit | edit source]Cloudflare released their next CLI to replace wrangler. It's called cf, and it's written in TypeScript. On Hacker News this choice of stack was disliked by enough people to be at the top of the page[1] when I viewed it.
I don't understand why this is written in Typescript. This is a great example of how agents can write code (I'm sure they wrote `cf`), yet having fundamental computer science knowledge is still critical. Do not force your users to manage the dependencies of your cli. Do write your cli in a compiled language. Understand the reason for those decisions and tell your agents to use the correct architecture.
Now, I personally don't much care for TypeScript because it has the annoying issue that people are constantly losing their dev machine environments to build scripts and so on. Rust is just as bad, and perhaps worse because build scripts cannot be avoided. These ecosystems are not well-suited to agentic development on a non-dev-isolated machine.

In any case, a TypeScript to Rust port should be quite feasible, though I really should have done a TypeScript to Go port. Go doesn't seem to have the same problem and now that code is not meant to be read, it seems fine to write in. Well, I decided to do the usual stuff. I armed a prewalk, described that I wanted a compatible Rust port, gave it a read-only key, and fired off one round after which I then armed a goal, said that it should use codegen if that's what upstream was using and went to bed only to return to a fully Rust port that was fully functional as far as reads were concerned.
Now this is not particularly notable today, but it is mind-blowingly impressive. The code is mostly generated from the OpenAPI spec for the API so it's quite easy to do with an agent, but the fact that it all works really well is wild. I had to fix the color use since it always printed the special characters but that made it impossible to pipe to jq and so on, but it otherwise seems quite functional. Someone else's software is now just a starting point for one to make one's own. Any irritants are rapidly fixable, and the stack is irrelevant for the most part since the code is inert text and sufficient for a rewrite.
Hell is Other People's Software
[edit | edit source]In talking to a friend, Chris, I realized that software is now like dreams: I am fascinated by my own nocturnal mental perambulations but I find other others' dreams tedious with the exception of those particularly close to me. Doubtless this is so when reversed. People do things like write general purpose software when special purpose software is what I want.
One thing that struck me is that the Bitwarden chrome extension has a giant WASM bundle it loads and does something with each time you click the button. This slows down start time substantially. Writing a tiny extension that just uses a smaller cut-down WASM bundle is quite easy, and when you do that, it opens in milliseconds. The final software can be so small that a human could look at it, and with a few LLMs peeking one can work on quite a bit of the security. One hopes that the limited form of specific software one has made here helps make it immune to the Cavendish Effect.
This is, of course, the same principle that led to me building a text viewer that is fast. Multi-GiB JSONs or Markdowns open fast and reformat to pretty-print before the original file could even be opened in Zed. This kind of thing is not universal, obviously, since it's not like I'm using my own browser.
The Esoteric is Accessible
[edit | edit source]
One big deal of the past is that CV algorithms were annoyingly hard to work with. And everyone knows that even if you got a flow around the algo development, you still needed labeled data. In Masters of Doom, iD Software aimed to develop level-development tools and so on as soon as possible. Labeling and so on present the extra part of the software that can be made nice.
Now all of this is trivial to do, and consequently, developing an appropriate CV solution also becomes much more accessible.
Frontier Labs Stymie Model-Use Improvement
[edit | edit source]
I have AI integrated into a large number of personal tools. One of the things that has it is the software that I use to author this blog. It has the usual grammar checks and spelling checks, but also has a web-search tool and a mechanism to tell me when I'm not really supporting my points well, or oversizing some text relative to others. It's not perfect and I usually ignore its style suggestions, but I like using it. Once or twice I've abandoned a post that it's successfully found a good argument against.
I've written bits of it with all sorts of models, and though Opus 5.5 is my current favourite, the tool was written by an amalgamation of various agents. Recently, I was concerned about cache breakage because follow-ups were taking just a little too long. The model is hosted on Nvidia RTX 6000 PRO Blackwell Max-Qs and it's DeepSeek V4 Flash Vision so I'm used to getting near instant streaming back.
Trying to use Opus 5.5 or Sonnet 5.5 to identify the problem was absolutely infuriating because it would frequently trigger safeguards. In this case, this isn't even a per-call safety issue; it's Anthropic judging that I'm trying to extract context from their model[3].
⏺ API Error: Opus 5.5's safeguards flagged this message (https://www.anthropic.com/legal/aup). This
sometimes happens with safe, normal conversations. Claude Code can't respond to this message with Opus
5.5.
Double press esc to edit your last message, or try a different model with /model.
Send feedback with /feedback or learn more: https://support.claude.com/en/articles/8106465
Details: `[reasoning_extraction]`
Request ID: req_011CfYhcdxs8acG8EAQAC6Q3
Message ID: msg_011CfYhceMvX2jDZhRR4WGtD
Well, in this case, all I was trying to do was make sure that my prompt has the reasoning sections included appropriately so that DeepSeek's prefix cache works well. In the end, I just used DeepSeek on its own and it 'recursively self-improved' pretty quickly.
Curiously, the frontier lab models aren't even optimal at security tasks. If you send an Opus-fixed piece of code through DeepSeek V4 Flash Vision it can help with advice that the Opus can find worthwhile. Some degree of diversity in advisors helps a lot, so I'm used to using multiple models on the same codebase both simultaneously and at different times.
GPT-6.1 Sol
[edit | edit source]This was a strange and sudden release since GPT-6 Sol is already so good. While I'm currently happily using Opus 5.5, I still keep my Sol subscription because it works as a great /prewalk tool for smaller models. I didn't notice 6.1 Sol being noticeably superior in any respect to 6 Sol.
Assistants
[edit | edit source]
AI Assistants have gotten quite good. OpenClaw was the original and I still call this class of assistants claw-likes in contexts where they can be confused. When Instinct released to much fanfare, I was surprised to find that it was some kind of inferior version of what I've had for a long time[4]. I'd stolen some features from Grokbot, and I had The Everything Store, and browser profiles for Julie and me and this could do many of the things we needed. My favourite part about the whole thing was that the local model powering it could access so much private data of ours.
But when I saw Muse work, things had gotten good. Muse's Stripe Link integrations and the fact that it could voice-call other people and talk to them and then seamlessly hand-off the conversation to the user were incredible. The occasional bloopers on Twitter were fine. A few rookie mistakes, some just a result of fast models hallucinating, and so on, but it seemed really good.
Now, of course we have Gemini Spark, and OpenAI's Dots. It's a bittersweet feeling when you know your software is beaten by some kind of market leader, but hopefully one of these will attract a lot of energy and become the de-facto one and I can swap to using it. I'd hate to switch only to find that they have decided to abandon the market! Like a Japanese holdout I am the single one among my friends to still have his Mac Mini[5].
One fortunate thing is that these tools use a facsimile of my browsing history. My agent integrates with my archiver, and my logged-in browser history helps hide the bot control, so things like Amazon banning Meta's Muse[6] haven't affected me yet. Once again, I am saved by virtue of not being the Cavendish.
Notes
[edit | edit source]- ↑ Admittedly, I have roughly a third of the comments blocked via Overmod.
- ↑ slowin on Permalink: HN • news.ycombinator.com • 49881834
- ↑ Getting reasoning makes distillation better; and previous Claudes were susceptible to this via various techniques - including taking a strong model and swapping to a weak one to prompt inject a reasoning extraction.
- ↑ I don't mean this in a "rsync is enough for Dropbox" way. It's just that Instinct was a particularly mediocre tool.
- ↑ The Mac Mini became somewhat of a craze for a personal assistant after OpenClaw's release. While inference on it is tremendously slow, it is a nice quiet and relatively cheap device that offers MacOS in a corner of the room doing things. I have my assistant run there, though it doesn't run any inference.
- ↑ Nat, Ives (2026-09-22). "Amazon Stops Meta's Muse at the Door". Wall Street Journal (in English). ISSN 0099-9660. Archived from the original on 2026-09-30. Retrieved 2026-09-30.