AI Updates
Another day with more action.
Browser Use Is Hard
[edit | edit source]
For the most part, my assistant works quite well. I tell it to get me something and it does some research and gets the thing that I want. But there's always more it could do and so on. And I keep trying to do some engineering that gets me into more trouble than it's worth. I did a few things with the browser so that it could become an actually usable tool.
In my case, I'm working with a fairly small model DeepSeek Flash 4 Vision 0831 and so I don't have the luxury of the true SOTA models. This means harness and tool engineering dominate success rates and so on.
Test Sites Matter A Lot
[edit | edit source]You can't iterate against the actual real world because you have to test your whole loop at once and a large amount of the stack starts mattering: rate-limiting, bot-evasion, CAPTCHA processing. While these pieces of the puzzle do matter, the reality is that improving things gets done faster if you isolate the variables. Just classic engineering. Nothing new.
In order to do this, one thing I first had to do was choose a bunch of sites I've interacted with in the recent past and start cloning them. A lot of websites are well-behaved, like Amazon.com, while others are nightmarishly weird like the TECO SF's visa appointment site. For my part, whenever the bot started failing at some site, I'd add it to my test set, interact with it, and clone its functionality. For some sites, this means adding in complexity like "what if you took too long to book an appointment and despite it appearing available it is gone by the time you get there" or "what if the price changes in the cart after you added to the cart".
Fortunately, with modern coding agents this is very quick.
A Completely Separate LLM Harness Didn't Work
[edit | edit source]For a bit I had the idea that the bot should pass through a task to a completely separate harness that would use the same model and do the work. I couldn't make this work very well. I tried the usual tricks (make the original message also available to this model, etc.) but in practice the other model needed to have steering prompts and all that as before.
This was the initial version of the browser drive tool I gave the agent, but it appears I've grown a lot of instructions and prompt specifics to the primary agent that I'd have to re-derive here and this wasn't yielding much for the time spent so I moved on.
A System One Model Actually Helps Quite A bit
[edit | edit source]Heterogeneity prevents standard failure modes, I suppose. While I couldn't make a separate harness explicitly to drive a browser to a goal in a way that the main agent couldn't do, I was able to have the main agent fire off short sub-tasks to a System One model that did the work.
In my case, I used Surogate.ai's Rune model which is based on Gemma 4 26B A4B. My main GPUs are used for LLMs, but I have a spare A6000 Ampere (an older GPU) that hosts a bunch of things and that had some 37 GB of VRAM unused. It's got half the VRAM bandwidth of my RTX 6000 Pro Blackwells but it is available at home so I used it. The model release on HuggingFace is BF16 so I needed to quantize it to fit and run on my A6000. That's a fairly mechanical process these days, and any LLM can assist you with the task.
In any case, post-quantization, I ran this on the A6000 backing a browser drive --goal "goal in text" --values k=v kind of approach which would return either "decision-model claims completion" or "decision-model stuck" and the LLM would do another round. This actually improved performance quite a bit, success rates went up, number of turns went down, and wall-clock time went down. On one task, I had zero completions in 300 s prior and 4/4 completions after.
Quantizing Rune
[edit | edit source]My RTX A6000 Ampere runs a bunch of things for me:
- bge-en-base: an embedding model that I use for semantic search across my apps
- frigate and a pose detector and ffmpeg/camera: a video movement detection application across my internal cameras I use to warn us on Astra doing some things in the house
- whisper: speech to text so that conversations with my assistant work
- kokoro: text to speech so that the agent can respond naturally
So overall that leaves 37 GB of VRAM free on that card. Now there's a few kinds of quantization we can do.
- FP8 - Ampere GPUs don't have any compute for FP8
- W8A8 INT8 - Uses some 27 GB of VRAM (or 33 GB if I keep the experts at full precision). Possible but I haven't run one of these on the vllm I use yet
- W4A4 INT4 - No one wrote the vllm support for this for this card
- W4A16 - this keeps the weights in 4-bit, but all the math runs at 16-bit
The W4A16 with GPTQ[1] is what I went with because the card has low-bandwidth but decent compute, so the smaller weights arrive faster but then the math is done in the BF16 format that the model is natively in.
llm-compressor does most of the work here, but you do have to choose a few things.
- The non-expert dense MLP better at full precision - everything goes to hell if you quantize this to 4-bit
- The attention and router modules stay BF16 - they're not large
- The actual experts quantize to 4-bit but at group size 64 - these are the majority of the params, but you have to set group size to smaller than the default 128 because Gemma has rows that don't divide neatly by 128 but do by 64.
The rest is pretty standard. GPTQ tells you how to modify the quantization to preserve the results by using calibration data to keep things lined up. And through this you only need each layer loaded to run its original quant through the calibration data, and then modify the new quant until it is close, and then proceed to the next layer.
Browser Tools
[edit | edit source]Early on in the approach I had a single headed browser that the agent used. I didn't do too much engineering and so the set of tabs would go to infinity. So the short version of what I did was:
- Separate browser profiles for Julie's accounts and mine
- Tabs reaped fixed duration after use
- Intermediary process for browser handling that provided an adapter on the CDP
The browser tool has a bunch of actions on a page:
- navigate - to go somewhere
- click - to click something
- type - to write something into a field, human-like cadence
- form - like above but for multiple with a submit
- upload - fill in an upload form
- key - single keys, for things like Esc
- back - navigate back
- drive - multi-step goal target with a System One model
- eval - run JS
And a bunch of inspection for the page:
- page - the accessibility tree for the page, with each iframe appended, and navigation history
- text - just visible text against a CSS selector
- inspect - all matching elements with attributes for a selector
- image - saves an image to a path
- screenshot - saves a screenshot to a path
- mhtml - archives the page to disk
- network - starts network traffic recording
There's also a little login helper login that fetches from the creds store into the page without intermediating through the agent, though the browser profiles are set up to use passkey auth with an external store by default. There's also some of the usual open and close and a little script to prepare for user takeover.
This is quite a large set of tools, and is mostly accumulated as the model looks for tools. I treat the agent as a user for whom I'll provide an affordance if they've reached for it a sufficient number of times and found it missing.
Multi-tenancy
[edit | edit source]The cotenancy of the assistant is quite helpful since most of what we do is household stuff and though we're on a single Prime subscription we do have our own Amazon accounts and this and that. Presumably one day we will add Astra until she's an adult and she can also ask the agent to buy her things.
Intermediary Process
[edit | edit source]The intermediary process is obviously a big engineering hole. You now suddenly have this adapter program that has its own lifecycle and so on, which is just a reliability nightmare. You need some justification for it:
- The whole bit about connecting to the browser process over CDP and sending commands was quite slow each time.
- More importantly, you can't listen to network traffic easily if you have a stateless thing
In the end, since it's an agent that's driving the entire tool, it is quite harmless because if there's any trouble the agent will just restart the intermediary process and unstick itself and the other agent instances also running will simply react with "the browser tool was killed under me" and take their next step. The fact that most of our interactions here are asynchronous helps a lot since they'll eventually figure it out.
Proactivity and Escalation are in Opposition
[edit | edit source]These agents are only useful when they can just do things without you having to actually direct them a lot. A good assistant knows when to just plunge ahead and when to ask for guidance. In some sense, this is the entire field of alignment. This is obvious, but because of prompt-sensitivity, temperature-sensitivity, and the fact that there is no really continuous way to move through the changes I make here I haven't thought to plot an ROC dot plot to see how the changes I make are affecting things. I really should, but two examples will illustrate how things can go wrong.
An appointment for an event for which appointments open up on a specific date and rapidly get filled had the agent navigate to the booking page and then send me a choice of which one I wanted. I had to use the browser remoting functionality to take over the page and pick because it had been 30 minutes since the message was posted.
When I added the site to my test sites and improved proactivity, the agent next noticed an email from UCSF about dates where they could run some tests and when I asked it which days I was available on, it replied to the human on the other end with a suggestion of which day works for us! Okay, that's exactly what I would have wanted except I try to keep my conversation with humans primarily human-reviewed at least.
My friend Armaan suggested that these kinds of outbound email actions should be placed in a send queue by default. He doesn't give the agent the ability to send an email, only compose a draft, and then he treats that as a send queue by reviewing, editing, and sending. I should implement something like that. Because I'm currently unable to move the true positive rate without also moving the false positive rate significantly, these action-class gates are likely the only solution. That also matches the reality better. I want a certain class of action ("conversation with humans") to be gated. Relying on broad proactivity vs. escalation is the wrong sort of lever.
Order Tracking
[edit | edit source]
One other thing that I had the bot do is that when it buys things and so on it should record them to a separate area so I can just track all of these things. So far it's only got a couple of orders and it's useful to see them.
Gemini the Giant Wakes Up
[edit | edit source]The new Gemini Argon model looks absolutely stunning. Google has long had a problem with their benchmarks and examples looking amazing and then the real world looking not so good. A friend of mine working with the Gemini team does have device-locked access to Argon, though, and he said that he's been using it for a few weeks and finds it quite good, and fast.
Excited, but maybe not convinced yet. Likely Fable 5.5 and the next Astra will blow it out of the water. Model releases are getting quite fast and I wonder if this is the RSI kicking in!
Notes
[edit | edit source]- ↑ Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh, "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers," arXiv:2210.17323 (cs.LG), 2023. arXiv and implementation on HF and implementation I used