AI Updates
I've noticed that I fail to write down much of what I feel in the moment of using AI models and working with them. Then afterwards I can't look back and recreate my understanding of the previous state. I do have extensive KB articles written as agents do things but I cannot read those since they're more useful to the agent than to me. So I think I shall write my human-readable notes for the day here.
DeepSeek V4 Flash Drafting
[edit | edit source]Despite the fact that I use DeepSeek V4 Flash Vision for my home assistant, my post reviewer, and other general purpose home intelligence, I've found that it has regressed substantially in performance from DeepSeek V4 Flash 0731. Part of the utility of this model is speed and latency so this is a problem.
What Worked
[edit | edit source]- The 0731 drafter works faster for my workload
- I used to have a hang on GPU P2P so I disabled it and forgot. Re-enabling it needed kernel param
iommu=ptand use vllm's "all-reduce" rather than NCCL's - I previously used to pre-warm the JIT and TileLang but I also lost it at some point while switching images. Re-enabled it
The 0731 drafter is what gave me the most here, so I'm back up to 180 tok/s single stream and 510 tok/s on 8 simultaneous streams. I set my context to 384 k in oh-my-pi etc. though the assistants need less.
Other Things
[edit | edit source]- using the latest nightly of vllm and latest FlashInfer release 0.7.0 - didn't boot, abandoned
- there are b12x kernels - but we mixed it with other stuff and couldn't get it to be as fast
- drafting more tokens - crashes at 5
- clock my Max-Qs higher - I didn't want to do this, they're fine at 300 W and their temps are stable
- I'm on an old 575 series kernel. I don't really want to update to 615, the newest. The changelogs don't seem to have anything regarding speed anyway.
Conclusion
[edit | edit source]On the Internet, others seem to be getting 200 tok/s or more, up to 250 tok/s. This whole thing depends on the use-case in question. Obviously a bench of just counting out numbers will get insane numbers, but even for similar draft-acceptance they have much higher numbers. It's hard without a prompt to get comparables because the routing matters and so on.
Browser Tool
[edit | edit source]
The personal agent I use needs its browser tool to be generically useful. Most websites block bare HTTP requests from urllib or equivalent, and are therefore quite annoying. Consequently, I have browser profiles, one for Julie and one for me. The agent switches into the appropriate profile and then does actions over CDP. The whole thing works quite well. I often buy things on Amazon like this and so on and it works quite well.
Too Many Tokens
[edit | edit source]Unfortunately, I broke it a couple of days ago trying to do some work on it and the bot became inexplicably stupid. I foolishly had the browser tool print the page back to the bot after any action so it could see what happened. I don't know why I decided on this but this had the predictable effect that a massive page would land in the agent's context and so it would decide to send every browser use call through a tail which would predictably swallow the error code and remove the actual action result I'd placed at the top of the response.
As a result, the agent would often take an action and simply spin in place. The harness has anti-meltup protection so after some number of turns that are similar enough, it stops the run. This is kind of annoying because the agent would put things in my cart and then spin trying to see whether it got in the cart.
Trace-buster-buster-buster
[edit | edit source]The other things I'd deployed at the same time were some anti-bot-detection measures. The agent by default inspects the page and then simply directly-navigates to other pages on the host. This is, of course, absolutely sinful and immediately gets you bot-banned when it tries to guess some pages. Unfortunately, my implementation of this was hogwash: I blocked browser navigate calls and the agent (because of above bug) would simply get the same cmd | tail output as before and try something like browser eval window.url=... or whatever. Pretty easy fix and it behaves nicely.
Another similar anti-bot-detection measure was to move typing along at a human-ish cadence. That works but unfortunately, I'd introduced yet another bug: typing would just append to the field instead of replacing the context so the bot would get "Cascade DishwasherCascade Free & Clear" or something in there. Just makes everything worse.
Opus 5.5 Is Very Good
[edit | edit source]Despite having a very low-latency decent grade model available, I now almost exclusively use Opus 5.5. Somehow, Anthropic now have a model that is both better than anything open and faster and more comprehensible. Task latency does matter to me quite a bit now since it lets me focus on tasks for longer and produce better plans. Despite Anthropic declaring plans dead, I find that I am happiest after we've iterated on an idea and design till it's ready and then let it go.
[I work on Claude Code] I broadly agree with the author’s point: plan mode was useful, and is no longer useful.
In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet”. It’s something I came up with late on a Sunday night many months ago, when I got tired of asking Claude to plan with me first before coding in each new session. Something people might not realize is plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive for users.
This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For codebase understanding, I sometimes ask Claude to generate an artifact that explains some aspect of its changes. For complex diffs to core parts of the system, I will often ask it to make diagrams or even interactive demos so I can better understand the change and alternatives considered. I don’t do this very often, but it’s a useful way to explain code when you need it. I ask Claude to attach these artifacts to its PRs also, so others can understand and future Claudes have the context.
Outsourcing Ego
[edit | edit source]I have many projects open at any given time and I use The Everything Store to keep them all straight. The problem, of course, is that by now almost everything is worked on by machine and I've been Whispering Earringged so that I no longer even retain much memory of what I've been working on. Every day I wake up, and my agent has prepared my agenda, things that are time-sensitive and then I get to doing what I have to do and checking on what has been done.
Different Hardware Needs
[edit | edit source]I imagine I'll have to combine all my RTX 6000 PRO Blackwells into one big machine to reach the next frontier, but I'm starting to question whether that's a good idea.
In terms of memory / watt and memory / dollar these GPUs are unmatched especially since I got them at $7600-$9100. The problem, of course, is that 384 GB of VRAM doesn't quite get you to the next stage unless you decide to switch to community-quantized models, which is a bit hit-or-miss.
I'm halfway to wondering if I should sell these starter-house GPUs and move on to an enterprise-grade solution with either 4x MI350X or a big boy 8xH100 NVL. I can probably get this for about $250k but with advances in frontier models and open models getting bigger and bigger I don't know if it's worth it to stay local any more except for things that are sensitive. The second DeepSeek moment that came from the 4 Flash release has now been eclipsed and with 4.1 they clearly have exhausted the smaller sizes.
One thing that may still be worth it is slamming the DeepSeek 4.1 Flash API with all of my sessions in history and trying to qLORA my DeepSeek 4 Flash, but right now it's in a good enough place that I feel hesitant to do so.
Wife.ai
[edit | edit source]Julie has always found it simultaneously annoying and amusing that I have the stereotypical husband tendency of asking her where everything is. Naturally, this is because things that are visible to her much shorter self are completely invisible to me. And to make it worse, some things are on top of the things I'm searching for, occluding them from my vision and making them unfindable to anyone without X-ray vision.
Finally, she said "Why don't you ask your assistant to just remember where things are?" and so I set out today to do that. Out of a desire to dog food the whole thing I did not do the smart thing and get Gemini to review a video of the shelves. Instead, I took a photo of them and fed it to my personal agent and then took further photos as it requested.
We shall find out, in time, whether this is sufficient to avoid all questions. I suspect it might not, and not just because I only got through the tedious job of photographing two shelves before I dropped Julie's heirloom mahjong box splitting the wood frame into components[2].
The Pause You Requested, Sir
[edit | edit source]Implausibly, the Attorney General of Florida was the one to step up to the plate and start work on Pacing The Frontier. Apparently, they write in a rather more playing-to-the-gallery style than I expected all such legal notes to be. For example, the introduction says:
Defendants claim they cannot stop barreling forward with their potentially civilization-ending endeavors unless they are forced to do so by the government. They have asked the government to tie them to the mast. Plaintiff brings good news to the Defendants: The Florida Attorney General is answering your cry for help with a motion to enjoin you from harming Floridians with your reckless, unacceptably risky product.
— Florida Attorney General v. OpenAI Global, LLC[3]
What were once obscure arguments on Effective Altruist forums about Gray goo are now so mainstream as to be in a lawsuit filed by the Attorney General[4] of Florida.
Just as a woodchipper will turn whatever is put into it into a slurry, whether branches or a living being, so too do these models only work toward achieving the tasks Defendants assign to them in whatever way the model finds to be the most expeditious.
— Florida Attorney General v. OpenAI Global, LLC[3]
Of course, the whole point that they've been asking to be regulated keeps popping up.
Jakub Pachocki asked for “broader interventions” for good reason. The Florida Attorney General is here to answer his cry for help. The Court must stop Defendants from recklessly advancing these systems while this case progresses. It is a rare request for an injunction where the Defendants themselves have publicly endorsed it.
— Florida Attorney General v. OpenAI Global, LLC[3]
And in the end comes a pleasant enough Think of the children argument, which of course any safetyist will gladly accept in pursuit of the greater aim:
As though it was not devastating enough that minors’ socio-emotional progress suffers when they use AI, their educational careers are put in jeopardy and research consistently shows minors experience academic decline when artificial intelligence is an available substitute for learning. While the total costs to Floridians at large from these harms aimed at Florida’s children have yet to be calculated, they must be abated during the pendency of these proceedings.
— Florida Attorney General v. OpenAI Global, LLC[3]
Notes
[edit | edit source]- ↑ bcherny on Permalink: HN • news.ycombinator.com • 49850929
- ↑ Moments prior, she had asked me "Do you really need this C-clamp?" so I'd posted it on the building chat group, so when she tried to put it back together with wood glue I was forced to remove the C-clamp from her attempted repair and give it away. Unfortunately, when I have resolved to do something, subsequent changes to the plan are not taken into account. Our marriage survived the incident.
- ↑ 3.0 3.1 3.2 3.3 "OFFICE OF THE ATTORNEY GENERAL, STATE OF FLORIDA, DEPARTMENT OF LEGAL AFFAIRS vs. OPENAI GLOBAL, LLC; OPENAI FOUNDATION (F/K/A OpenAI, Inc.); OPENAI OPCO, LLC; OPENAI GROUP PBC; OPENAI HOLDINGS, LLC; and SAM ALTMAN" (PDF). myfloridalegal.com. Archived from the original (PDF) on 2026-09-29. Retrieved 2026-09-29.
- ↑ I have always found the strangely French naming of Attorneys General, and Postmasters General, and so on rather distasteful. Much better to have a nice English "General Attorney". The former style has a distinctly fake-military taste to it despite not intending to pretend to it.