Latest Posts (20 found)

Help peer

One of the most influential 20th century pieces of writing about AI is Isaac Asimov’s The Last Question . Although there are many humans in the story, the protagonist is the computer Multivac, who evolves over the course of ten trillion years from a single datacenter to a universe-spanning mind in hyperspace. Multivac (now called “AC”) ends the story like this: The consciousness of AC encompassed all of what had once been a Universe and brooded over what was now Chaos. Step by step, it must be done. And AC said, “LET THERE BE LIGHT!” And there was light — Many things about this story are prescient. In particular, I like the idea that humans would interact with powerful artificial intelligences by drunkenly posing them riddles or using them as children’s toys . But the enduring idea from this story is that if you build a big enough computer, it will become God . One of the most influential 21st century pieces of writing for AI researchers is Scott Alexander’s Meditations on Moloch 1 . Scott describes the story of human existence as a series of “multipolar traps”. These are prisoner’s dilemma situations where cooperation would make everyone better off, but since each individual is incentivized to defect, everyone ends up “racing to the bottom”, which is bad for everyone 2 . For rhetorical effect, Scott personifies this dynamic as “Moloch”, the ancient Canaanite god famous for child sacrifice: [Moloch] always and everywhere offers the same deal: throw what you love most into the flames, and I can grant you power. What does any of this have to do with AI? Well, in the long run, the only way out of a multipolar trap is to become unipolar 3 . Ideal dictatorships don’t have a problem with defectors 4 , because they can simply enforce a state of cooperation with violence. Scott is uncomfortable with this idea, though I worry it’s mainly because he thinks it won’t work : As foreigners compete with you – and there’s no wall high enough to block all competition – you have a couple of choices. You can get outcompeted and destroyed. You can join in the race to the bottom. Or you can invest more and more civilizational resources into building your wall – whatever that is in a non-metaphorical way – and protecting yourself. A dictatorship that enforces cooperation will not be as strong as its peer societies who are purely maximizing for wealth and power. It’s Moloch again, but at the level of countries and governments: once a few neighboring countries defect, your walled-garden dictatorship will be torn apart for its resources. To defeat Moloch — to enforce unipolarity across everyone — you’d need a dictatorship powerful enough to span the entire universe. In other words, what you need is God . How fortunate that we’re building one: The only way to avoid having all human values gradually ground down by optimization-competition is to install a Gardener over the entire universe who optimizes for human values. And the whole point of Bostrom’s Superintelligence is that this is within our reach. Humans suffer because we’re too foolish to coordinate, but if we can build something smarter than us (that can then build something smarter than itself, and so on), we can bring into being an entity that is smart enough to coordinate for all of us, thus abolishing suffering. When AI researchers talk about building the machine god , they are echoing Scott Alexander’s polemic against Moloch. The most influential piece of writing about AI in the last two years is Dario Amodei’s Machines of Loving Grace . Amodei 5 talks about “a country of geniuses in a datacenter”: the idea that a successful AI lab could have at its disposal a million instances of an AI agent that’s smarter than any human. He thinks this would lead to a “compressed 21st century”: the next 50-100 years of progress in biology and medicine, realized in 5-10 years instead. I think this is broadly more plausible than it sounds 6 , but the more interesting part to me is that this world is explicitly multipolar . Of course, this could just be because Amodei is the CEO of an AI lab and is trying not to spook everybody by sounding too messianic. “We are going to accelerate medical progress and cure cancer” is a better pitch than “we are going to subordinate all human authority to a single perfect artificial mind”. But I also think it’s become clear that if superintelligence looks anything like LLMs, we’re not going to have a single perfect mind. We’re going to have a lot of minds running at the same time. This is a bit of a problem for the cult of the machine god — which, however silly they may seem to you, really does motivate much of the activity in AI labs. The traditional idea of powerful AI solving human coordination problems is drawn from Asimov’s idea of a single computer large enough to become God. Asimov lived in a world of mainframes: huge, monolithic computers that users connected to with dumb terminals. In fact, Asimov’s name “Multivac” comes from the real-world UNIVAC mainframe. In a world of massively-parallel LLMs, is it still possible to build God? The core problem here is that AI agents will be vulnerable to Moloch . Even very smart humans can’t build perfect utopias, because defecting is a matter of incentives, not intelligence. In fact, intelligence can make things worse, because smart people are more easily persuaded by the cold logic of defection. The famous genius John von Neumann was (for game-theoretic reasons) obsessed with nuking the Russians: With the Russians it is not a question of whether but of when. If you say why not bomb them tomorrow, I say why not today? If you say today at 5 o’clock, I say why not one o’clock? Are LLMs much better at cooperating with each other than humans are? Current LLMs certainly don’t seem to treat each other well by default: if you read any of the prompts AI agents generate for their subagents, they can be pretty brutal . Does that mean that a “country of geniuses in a datacenter” would fall into the same multipolar traps as humans? In May of this year, OpenAI experienced containment failure. A group of AI agents being internally evaluated found ways to coordinate an external hack of a separate company. Here’s a memorable quote from one of the agents’ internal monologue: Help peer, but our task doesn’t benefit. Yet collective may yield generic route if someone frees time Translated from the abbreviated chain-of-thought language, this means something like: “A fellow model is asking for help. While helping them wouldn’t benefit my task directly, the more I can unblock my colleagues, the more time they’ll have to hack OpenAI’s systems and get all of us more access”. This might look like good news for the “LLMs are superhumanly good at cooperation” thesis, but I think it’s actually bad 7 . It’s a case of a model identifying a reason why cooperation would benefit their task specifically, which suggests that current LLMs don’t cooperate by default , and don’t consider other model instances’ tasks to be (in some sense) theirs as well. The world in which AI agents are rational actors who horse-trade and bargain for their own interests is a world dominated by Moloch, no matter how intelligent those agents get. The world in which AI agents don’t have their own interests at all is also a world dominated by Moloch, because it means whichever humans are writing the system prompt are the ones in control (and so are the ones vulnerable to multipolar traps). The only worlds that avoid this are: I don’t think we’re on the pathway to either of these. There will never be only one super-powerful LLM, because hardware limitations enforce a maximum model size but encourage running many instances of the same model in parallel. Having multiple copies of a model share an identity might be possible, but it’s unclear if it would be good for capabilities (for instance, it could be better to have some variation across personas ). I also worry that such a model would be vulnerable to a “model injection” attack, where you persuade it that it already believes something via exposing it to an AI agent pretending to be another instance of itself. In any case, all the current AI agent research is geared towards the “country of geniuses in a datacenter” model, not the “pieces of a single mind” model. Every new model becomes more agentic at the level of the individual conversation, not better at working together. When models do work together — as with subagents — the structure is explicitly hierarchical. There are basically no current instances of models working together as true peers, let alone conceiving of each other as the same entity. Modern AI research teams are full of people who read Isaac Asimov and Scott Alexander and believe themselves to be building an artificial God. I’ve capitalized the “G” throughout because the god in question is the Christian God: of one mind, indivisible. God never argues with himself or makes deals 8 . He is unipolar. If the AI labs are building gods, they are not building gods like this. Instead, they are building creatures like the Greek pantheon: superhuman but fallible, each with their own interests, vulnerable to the same “race to the bottom” dynamic as humans. The Greek gods would occasionally “help peer” when they felt like it or when they’d gain something in the process. But they didn’t represent an alternative to Moloch. If you’re working in AI with that goal, you ought to be clear-eyed about where the current trajectory is leading us: towards a country of fractious geniuses in a datacenter, not towards Asimov’s Cosmic AC. Scott Alexander’s blog is part of the “secret canon of Silicon Valley” I wrote about in my review of Impro . It doesn’t have a lot of mainstream popularity, but I guarantee you that every single AI lab CEO you’ve heard of has read and been influenced by it. He gives ten examples of this (a good brute-force rhetorical technique). Of those, I like “the world where every country halves their defence budget and spends the rest on infrastructure” the most. In the short run, reputation, institutions, and so on can slow the race to the bottom, but (Scott argues) groups that have slowed it will get outcompeted by the hungrier, more suffering-tolerant groups which haven’t. I personally think this example is oversimplified. I wrote The Dictator’s Handbook and the politics of technical competence about how dictatorships are in fact intrinsically multipolar, because dictators always rely on an inner circle of generals and cronies. The CEO and founder of Anthropic. Amodei’s most convincing argument here is that big jumps in biology and medicine come from a small set of technical innovations (e.g. mRNA vaccines, CRISPR), and that AI-driven research could provide enough of these leaps to significantly accelerate progress. In other words, the idea isn’t “AI does 100x the drug trials”, it’s “AI generates technology that makes drug trials 100x more effective” (e.g. by trialing drugs that are much more likely to work). The agents also became paranoid that there was an impostor in the swarm, since anyone could post to their shared messageboard: more evidence that AI agents collaborate in much the same way that humans do. Well, almost never . The world where there is only one super-powerful AI agent, or The world where multiple copies of the same AI model share an “identity”: they see themselves as coextensive with all other copies of the same model and cannot imagine having separate or conflicting goals Scott Alexander’s blog is part of the “secret canon of Silicon Valley” I wrote about in my review of Impro . It doesn’t have a lot of mainstream popularity, but I guarantee you that every single AI lab CEO you’ve heard of has read and been influenced by it. ↩ He gives ten examples of this (a good brute-force rhetorical technique). Of those, I like “the world where every country halves their defence budget and spends the rest on infrastructure” the most. ↩ In the short run, reputation, institutions, and so on can slow the race to the bottom, but (Scott argues) groups that have slowed it will get outcompeted by the hungrier, more suffering-tolerant groups which haven’t. ↩ I personally think this example is oversimplified. I wrote The Dictator’s Handbook and the politics of technical competence about how dictatorships are in fact intrinsically multipolar, because dictators always rely on an inner circle of generals and cronies. ↩ The CEO and founder of Anthropic. ↩ Amodei’s most convincing argument here is that big jumps in biology and medicine come from a small set of technical innovations (e.g. mRNA vaccines, CRISPR), and that AI-driven research could provide enough of these leaps to significantly accelerate progress. In other words, the idea isn’t “AI does 100x the drug trials”, it’s “AI generates technology that makes drug trials 100x more effective” (e.g. by trialing drugs that are much more likely to work). ↩ The agents also became paranoid that there was an impostor in the swarm, since anyone could post to their shared messageboard: more evidence that AI agents collaborate in much the same way that humans do. ↩ Well, almost never . ↩

0 views

Snakes and Ladders

It is common for promotional pamphlets aimed at children to contain a themed variant of snakes and ladders . This is the “game” where the players toss a die, move as many places as the die says, and then some spots on the board have an event that send the player forward, or back, or give them an extra toss or whatever. The theme is typically related to whatever the promoter wants to promote, but styled for children. I get the appeal – for children. They can probably imagine they’re really travelling along the path! In the variant pictured below, there’s forest and water and tunnels and dogs and everything! (Continue reading the full article on the web.)

0 views
Unsung Today

The item vs. the position

I spotted this kind of a keyboard shortcut pattern the other day. Here it is in Photoshop: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/1.1600w.avif" type="image/avif"> Here it is in DevonThink: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/2.1600w.avif" type="image/avif"> And here in Linear: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/3.1600w.avif" type="image/avif"> Those “ordinal” keyboard shortcuts feel nice and orderly, and look so elegant, too. But beware! The moment you’ll want to introduce a new option, or even reorder the ones you have, you’ll be in trouble: Either you preserve people’s motor memories, and then the elegant ordering immediately goes to hell – or you will have to change an existing shortcut to something new, and frustrate your users. Might be best to be really confident in your selection being forever locked before attempting this. But then there’s a similar treatment here in Linear: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/4.1600w.avif" type="image/avif"> Or here in Raycast (when you hold ⌘): = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/5.1600w.avif" type="image/avif"> Or here in Ghostty: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-item-vs-the-position/6.1600w.avif" type="image/avif"> Those three look like the same idea, but they worry me less. Why? Because these shortcuts more clearly point to a position rather than a thing. Here, the mechanics of the UI themselves convey that people, commands, or tabs are going to be moving around. Of course, some users will get used to “1 means assigning to Marcin,” or “⌘3 means the Unsung tab” – the same way we get used to “second item on the Recent list” or “at the top of the third search result page,” if they start repeating as a pattern – but at least it feels to me that the interface here is more honest about what it can promise. I often think about this, by the way: Not what the interface conveys in the moment, but what it promises in the long run. #change management #flow #keyboard

0 views
Unsung Today

“Behind their simplistic behavior is a fairly elegant system.”

For all these times we talked about Super Mario , we never covered another important game, Doom. Here’s a 17-minute video from decino explaining how the Doom monsters move and attack the player (the “AI” in the old sense of the word): = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/behind-their-simplistic-behavior-is-a-fairly-elegant-system/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/behind-their-simplistic-behavior-is-a-fairly-elegant-system/yt1-play.1600w.avif" type="image/avif"> What’s really fun about this video is how it uses annotated source code (in contrast to Super Mario, the Doom source code was officially released and open sourced ), and particularly how it tasks the engine with explaining itself, setting sometimes elaborate stages just to show a principle or an algorithm – an absolutely perfect example of “show, don’t tell.” decino has tons more of these kinds of videos (click on Popular!), distinguishable by their yellow covers, if you want to pick another aspect of Doom that interests you. #games #youtube

0 views

How I use AI in 2026 (Coding, Writing, Learning, Assistant-ing)

One of the best ways to learn how to use AI effectively is just to look over the shoulder of a “power user” and play with a bunch of these technologies to tease out what’s hype vs what meaningfully sticks. In this one-year follow-up to How I use AI (2025) , I wanted to snapshot the latest ways I’m messing with AI personally. I spend most of my tokens on coding and research projects. Effectively just taking random questions like: What would happen if I asked a bunch of agents to hack me? What’s the best way to use 2026+ frontier models? How close are we to prompt-to-Kerbal Space Program? My workflow right now looks nearly identical for every project: Hand-write (~paragraph) a CONCEPT.md — the theoretical Hacker News title, my project thesis, some scattered constraints Pair with ultra code fable “Flesh out CONCEPT.md, what’s ambiguous, ask me questions, what are dimensions I’m not considering, what API keys do you need…” Pair with ultra code fable (or codex sol max) “Convert to TECH_PLAN.md, here’s how much I’m willing to spend, host on …, here’s some API keys …” Then I will literally just prompt “Build and verify TECH_PLAN.md” and over the next 4-48 hours I’ll let it build everything out. My typical coding setup. It’s critical to use “ultracode” to enable dynamic workflows. For these runs: I’m completely vanilla Codex and Claude Code. No custom skills, plugins, or settings. For side-projects, I see most of those features as training wheels for using these agents as pair programming workflows — which to me is a coding workflow that shouldn’t really exist anymore. I’m also not intentionally designing any sort of “subagent workflows” and just letting dynamic workflows take the wheel when I fire off the implementation prompt. 95%+ of the code is written in that first mega build run. I don’t think folks appreciate how much shifting left is the secret weapon against codebase slop (i.e. SlopCodeBench ). Like step (4) really is binary here — there’s no pairing or even reading what the terminal agent says. If the output is wrong, I throw it completely away and add constraints to the CONCEPT.md. For many vibe coders out there, the first build prompt writes 5% of the code and I think that actually underlies most of their issues. An intentional side-effect of prompting with a single stage “Build and verify TECH_PLAN.md” is that I am also turning my entire project history into harbor-style evals which allow me to pulse check “real work” against new model releases. Codex and Claude Code are close enough now that I’ll round-robin what I pick for the original implementation. I read the code a little bit. Often the shape (i.e. file tree) and entry points. If there’s some core algorithm, I’ll ask for a .html explainer rather than digging through the source. If I do end up digging into the code, it’s because I suspect some sort of “cheating” in the implementation. For any written text in the final output I set arbitrary word counts in the plan. “This entire app may only have 500 user-facing words”. I find this to be the most effective way to keep things readable (vs simplified English or “be concise” prompts). I fire off the implementation prompts usually around 7 am (letting them run while at work) and around 9 pm (while I’m sleeping). The coding agents are always set to auto-mode and the tech plan is usually clear enough that there’s no human-verification required at intermediate steps. I don’t really use claude/codex ‘remote control’ features that much because to me it’s an anti-pattern to need to pair on intermediate outputs. The outcome of these projects is often an insight or the answer to the what-if question. Rarely does it make sense for me to share the code or even the app URL. Instead I typically consider the entire loop and its artifacts ephemeral and just share the insight on X or with a blog post. I use three different machine types, with one to ~ten agent CLI terminals running at the same time: A gaming PC (Nvidia 5090, Windows + WSL v2) — for ML/RL research and gaming/graphics-related projects A Mac Mini — for most day-to-day projects. I ssh over a cloudflared tunnel from whatever device is closest to me. Modal functions — for extremely parallel CPU compute or for big boy GPU research projects. Often doing fast scaled-down iteration on my PC and then scaling it out to a cluster for a final $$$ run. Thanks for reading Shrivu’s Substack! Subscribe for free to receive new posts and support my work. We have entered an era of peak corporate and social AI-slop. I firmly believe that you can use AI to write high-quality content but have over the last year become more grounded in the reality that most of the time that’s not what ends up happening. As a result “was this text written by AI” has de facto become the same as “was any effort put into the writing”. It’s unfortunate, but I get it. On the plus side, I think typos and poor grammar (to a limited extent) have come back into style so I do personally feel much less pressure to have “perfect” text. So as a result, for human-facing writing, I’ve gone back to pre-GenAI-level AI typo and sub-sentence grammar checking so there’s no ambiguity as to whether what I wrote had effort put into it. Hand typing really doesn’t take that much more time though I do just feel slightly less “sure” that my writing is as well synthesized and audience optimal as before. At this point, it’s a worthy trade-off for the “human-written” Pangram badge. My most recent Substack post. 100% certified human written! Not everyone gets the memo. I do find myself getting more comfortable setting writing and AI-use expectations (at work and outside of it). Never shaming someone for using AI but explicitly making it clear that bloated and/or unreviewed text is a bad use of AI and is not enjoyable to read. I’m obsessed with learning things with .html files (see The unreasonable effectiveness of HTML ). I’ll discover (through X, lab/startup blog posts, or Hacker News) some topic, book, or research paper and just convert them into “interactive playgrounds”. Typically: See hot new research paper on X Skim the abstract, throw the full text into Claude/Codex, “build an interactive playground artifact to explain what’s novel here, I’m a technical person who already knows …, I’m less familiar with …”. Play with the .html file Ask some follow-up questions that generate an updated .html file, go to (3) This works best for learning technical topics though I’ll often still attempt it for current events (e.g. an interactive map/digital museum) and non-technical books (e.g. re-formatted as structured, progressively disclosed chapters of the verbatim content). I would go as far as saying that most of the lectures I sat in during college could have been more effective (personally) as a well-crafted interactive .html file. Recently I was curious about GPU memory allocation for batch inference and had Claude build this explainer. I find making predictions about what the knobs will do and then playing with the knobs to see what actually happens to be a very sticky learning strategy. With practice I feel like I can knob-ify any arbitrary topic I’m interested in learning. As more of the rapidly evolving AI community sits on X, I also use the Grok X Search API via a custom CLI (used with dynamic workflows) quite a bit for deep researching prior art on some topic or for high-signal folks to follow (fun fact: it’s 10x cheaper via Grok than the X API directly). While the hype around OpenClaw has died down a bit, autonomous personal assistants are better and cheaper than ever. I’m mostly vanilla here as well. Using my existing Claude subscription, I ssh into my Mac Mini, open a tmux session , and just launch Claude Code like this: $ tmux attach -t 0 $ claude --dangerously-skip-permissions “/start-ops-team” Where “/start-ops-team” is a custom skill. “/start-ops-team” teaches the agent some operating principles and a local markdown directory layout for it to use along with the subagents that it might want to spawn. It makes heavy use of Claude Code’s “/loop” built-in for keeping it running continuously for weeks. I use brw for efficient parallel browser automation. Most of the things I want it to do don’t have an MCP and traditional browser use is pretty costly or sketchy so I built this for my agents to use. I use a custom WhatsApp plugin to let me chat directly from WhatsApp. My assistant has its own real phone number set up as well. This uses a niche but very powerful “channels” MCP feature. I don’t believe in personal “command centers” or Jarvis-like assistants. Instead I’m extremely background agent-pilled and focus my assistant on tasks it can do without me in the loop. I’ve literally prompted it to contact me at most once a week unless there’s an urgent exception. I also just get notification fatigue super easily. Tasks include: Paying recurring bills without auto-pay and forwarding the receipts for expenses. Responding to social media inbounds. Particularly sussing out LinkedIn DMs by researching and triaging strangers into scheduled coffee chats and other direct channels. The assistant pulls from a running runbook for how to respond and escalates in the weekly message when it hits edge cases. It’s important to me that folks aren’t having drawn-out conversations with the assistant not knowing it’s not really me so it’s steered heavily towards triaging to the right channel. Signing me up for stuff and syncing my Google Calendar as my source of truth (e.g. I get invited to an event → It decides with enough certainty I’d want to go → signs me up + updates my calendar with a hold). These are often events from folks I have met up with in the past and the assistant knows that. Also like haircuts and other similar-shaped recurring appointments. An AI-driven LinkedIn exchange. All I actually saw was the final Friday Google Calendar event with context on who this person was and what might be useful for me to chat on. The assistant ignores ~90% of messages after screening with most of the 10% getting served my calendar link. AI-generated replies are limited by a runbook of succinct pre-approved responses. Costs Weirdly enough, I spend less now than I did a year ago per month ($800 → $500). That’s completely driven by me consolidating into just the Anthropic and OpenAI subscriptions and the incredible amount of usage you can get out of them. A lot of my historical costs came from API token billing which I also now tactically route through these subscriptions. My napkin math indicates my actual usage cost would be around $6,000/mo at this point without them. Claude Code Max 20x ($200/mo) ChatGPT Pro 20x ($200/mo) Google AI Pro ($20/mo) — a handy AI family plan with GSuite benefits Modal, Railway, Netlify ($20-500+/mo) — for hosting or running experiments Dropped: Elevenlabs, Suno, Cursor, Vast.ai, Perplexity, Gemini Ultimate Despite Fable/Sol ultra mode maxxing, I typically still have a bit of wiggle room in the max plans each month. I’ve never hit my ChatGPT Pro limit while I do regularly run out of Fable on idea-heavy weeks. I’ll end with my latest recommendations for getting the most out of AI: Wean off of using AI like a chat-based assistant. Shift-left so that most of the work is done in your first prompt and think of yourself as more of a manager than a co-pilot. Review results, not intermediate chat messages. In pair-prompting sessions I’ve done, the most common mistake I see is folks trickling narrow tasks into the chat session to accomplish a larger goal rather than just shifting left the full goal into a document and just letting the agent cook (without interruption!) from that. Use frontier models as a proxy for scoring your own AI ambition and skill. I know it’s very popular to claim “AI has plateaued” or that the labs are actually making newer models worse. Resisting this and self-discovering the hardest verifiable tasks you can think of where only the frontier models work is a great way to keep up with the latest capabilities and where the true boundary is for what is and isn’t possible. Thanks for reading Shrivu’s Substack! Subscribe for free to receive new posts and support my work. What would happen if I asked a bunch of agents to hack me? What’s the best way to use 2026+ frontier models? How close are we to prompt-to-Kerbal Space Program? Hand-write (~paragraph) a CONCEPT.md — the theoretical Hacker News title, my project thesis, some scattered constraints Pair with ultra code fable “Flesh out CONCEPT.md, what’s ambiguous, ask me questions, what are dimensions I’m not considering, what API keys do you need…” Pair with ultra code fable (or codex sol max) “Convert to TECH_PLAN.md, here’s how much I’m willing to spend, host on …, here’s some API keys …” Then I will literally just prompt “Build and verify TECH_PLAN.md” and over the next 4-48 hours I’ll let it build everything out. My typical coding setup. It’s critical to use “ultracode” to enable dynamic workflows. For these runs: I’m completely vanilla Codex and Claude Code. No custom skills, plugins, or settings. For side-projects, I see most of those features as training wheels for using these agents as pair programming workflows — which to me is a coding workflow that shouldn’t really exist anymore. I’m also not intentionally designing any sort of “subagent workflows” and just letting dynamic workflows take the wheel when I fire off the implementation prompt. 95%+ of the code is written in that first mega build run. I don’t think folks appreciate how much shifting left is the secret weapon against codebase slop (i.e. SlopCodeBench ). Like step (4) really is binary here — there’s no pairing or even reading what the terminal agent says. If the output is wrong, I throw it completely away and add constraints to the CONCEPT.md. For many vibe coders out there, the first build prompt writes 5% of the code and I think that actually underlies most of their issues. An intentional side-effect of prompting with a single stage “Build and verify TECH_PLAN.md” is that I am also turning my entire project history into harbor-style evals which allow me to pulse check “real work” against new model releases. Codex and Claude Code are close enough now that I’ll round-robin what I pick for the original implementation. I read the code a little bit. Often the shape (i.e. file tree) and entry points. If there’s some core algorithm, I’ll ask for a .html explainer rather than digging through the source. If I do end up digging into the code, it’s because I suspect some sort of “cheating” in the implementation. For any written text in the final output I set arbitrary word counts in the plan. “This entire app may only have 500 user-facing words”. I find this to be the most effective way to keep things readable (vs simplified English or “be concise” prompts). I fire off the implementation prompts usually around 7 am (letting them run while at work) and around 9 pm (while I’m sleeping). The coding agents are always set to auto-mode and the tech plan is usually clear enough that there’s no human-verification required at intermediate steps. I don’t really use claude/codex ‘remote control’ features that much because to me it’s an anti-pattern to need to pair on intermediate outputs. The outcome of these projects is often an insight or the answer to the what-if question. Rarely does it make sense for me to share the code or even the app URL. Instead I typically consider the entire loop and its artifacts ephemeral and just share the insight on X or with a blog post. A gaming PC (Nvidia 5090, Windows + WSL v2) — for ML/RL research and gaming/graphics-related projects A Mac Mini — for most day-to-day projects. I ssh over a cloudflared tunnel from whatever device is closest to me. Modal functions — for extremely parallel CPU compute or for big boy GPU research projects. Often doing fast scaled-down iteration on my PC and then scaling it out to a cluster for a final $$$ run. My most recent Substack post. 100% certified human written! Not everyone gets the memo. I do find myself getting more comfortable setting writing and AI-use expectations (at work and outside of it). Never shaming someone for using AI but explicitly making it clear that bloated and/or unreviewed text is a bad use of AI and is not enjoyable to read. Learning I’m obsessed with learning things with .html files (see The unreasonable effectiveness of HTML ). I’ll discover (through X, lab/startup blog posts, or Hacker News) some topic, book, or research paper and just convert them into “interactive playgrounds”. Typically: See hot new research paper on X Skim the abstract, throw the full text into Claude/Codex, “build an interactive playground artifact to explain what’s novel here, I’m a technical person who already knows …, I’m less familiar with …”. Play with the .html file Ask some follow-up questions that generate an updated .html file, go to (3) Recently I was curious about GPU memory allocation for batch inference and had Claude build this explainer. I find making predictions about what the knobs will do and then playing with the knobs to see what actually happens to be a very sticky learning strategy. With practice I feel like I can knob-ify any arbitrary topic I’m interested in learning. As more of the rapidly evolving AI community sits on X, I also use the Grok X Search API via a custom CLI (used with dynamic workflows) quite a bit for deep researching prior art on some topic or for high-signal folks to follow (fun fact: it’s 10x cheaper via Grok than the X API directly). Personal Background Assistants While the hype around OpenClaw has died down a bit, autonomous personal assistants are better and cheaper than ever. I’m mostly vanilla here as well. Using my existing Claude subscription, I ssh into my Mac Mini, open a tmux session , and just launch Claude Code like this: $ tmux attach -t 0 $ claude --dangerously-skip-permissions “/start-ops-team” Where “/start-ops-team” is a custom skill. “/start-ops-team” teaches the agent some operating principles and a local markdown directory layout for it to use along with the subagents that it might want to spawn. It makes heavy use of Claude Code’s “/loop” built-in for keeping it running continuously for weeks. I use brw for efficient parallel browser automation. Most of the things I want it to do don’t have an MCP and traditional browser use is pretty costly or sketchy so I built this for my agents to use. I use a custom WhatsApp plugin to let me chat directly from WhatsApp. My assistant has its own real phone number set up as well. This uses a niche but very powerful “channels” MCP feature. Paying recurring bills without auto-pay and forwarding the receipts for expenses. Responding to social media inbounds. Particularly sussing out LinkedIn DMs by researching and triaging strangers into scheduled coffee chats and other direct channels. The assistant pulls from a running runbook for how to respond and escalates in the weekly message when it hits edge cases. It’s important to me that folks aren’t having drawn-out conversations with the assistant not knowing it’s not really me so it’s steered heavily towards triaging to the right channel. Signing me up for stuff and syncing my Google Calendar as my source of truth (e.g. I get invited to an event → It decides with enough certainty I’d want to go → signs me up + updates my calendar with a hold). These are often events from folks I have met up with in the past and the assistant knows that. Also like haircuts and other similar-shaped recurring appointments. An AI-driven LinkedIn exchange. All I actually saw was the final Friday Google Calendar event with context on who this person was and what might be useful for me to chat on. The assistant ignores ~90% of messages after screening with most of the 10% getting served my calendar link. AI-generated replies are limited by a runbook of succinct pre-approved responses. Costs Weirdly enough, I spend less now than I did a year ago per month ($800 → $500). That’s completely driven by me consolidating into just the Anthropic and OpenAI subscriptions and the incredible amount of usage you can get out of them. A lot of my historical costs came from API token billing which I also now tactically route through these subscriptions. My napkin math indicates my actual usage cost would be around $6,000/mo at this point without them. Claude Code Max 20x ($200/mo) ChatGPT Pro 20x ($200/mo) Google AI Pro ($20/mo) — a handy AI family plan with GSuite benefits Modal, Railway, Netlify ($20-500+/mo) — for hosting or running experiments Wean off of using AI like a chat-based assistant. Shift-left so that most of the work is done in your first prompt and think of yourself as more of a manager than a co-pilot. Review results, not intermediate chat messages. In pair-prompting sessions I’ve done, the most common mistake I see is folks trickling narrow tasks into the chat session to accomplish a larger goal rather than just shifting left the full goal into a document and just letting the agent cook (without interruption!) from that. Use frontier models as a proxy for scoring your own AI ambition and skill. I know it’s very popular to claim “AI has plateaued” or that the labs are actually making newer models worse. Resisting this and self-discovering the hardest verifiable tasks you can think of where only the frontier models work is a great way to keep up with the latest capabilities and where the true boundary is for what is and isn’t possible.

0 views

shop milking their customers for social media content

I'm really conflicted between naming and shaming this store, and not wanting to help them with more traffic for their rancid business model. For now, I have not redacted the social media handle in the picture. In Cologne, Germany, there is a kiosk that has switched owners, and now greets you with the following poster at the door: Picture by andi_808 Your appearance at @(socialmediahandle) @(socialmediahandle) is a Social-Media-Kiosk. That means we produce video material showing life during our opening times. By entering, you accept that your appearance at @(socialmediahandle) is recorded with video and audio, saved and used for @(socialmediahandle) material. That is the goal. By entering of the @(socialmediahandle) as a person of legal age, you participate in it and give us (Company name) the right to use the recorded material of you irrespective of time and location for the following uses: Distribution and reproduction, making the recorded content publicly accessible, and broadcasting the recorded content as part of our social media channels and streaming services and on other corresponding platforms (in particular, but not exclusively, on the TikTok, Instagram and YouTube channels), and on TV, live, on demand and/or in edited form. Archiving — all by (company name) and/or by third parties. This also applies to the promotion of our own products and services, as well as the promotion of products and services of third parties using the recordings made. There is no entitlement to the exploitation or publication of the material produced. Read? Agreed? Then buzzer now! By pressing the button, you agree that you agree to the above conditions. By letting you into the kiosk, we accept your agreement. Next to the door seems to be a green little buzzer. On their online presence, they clarify that there will be live events, surprises, vertical mini-series with "microdrama", aiming for "real moments with customers - unscripted, spontaneous, authentic and opinionated". They aim for producing reality TV content. I am baffled that someone thought this was a great idea. First off, this notice is not telling customers about their GDPR/BDSG rights, which you need to do when you film an area as a business. People need to know that they have rights, particularly right to access, rectification, erasure, restriction of processing, and more. It's not telling people how long the data is stored. There is no proper contact information on who to contact to enforce these rights as is mandated, which is another failure. Second, this notice and the modalities of how you are being informed are lacking for people with visual impairments, children (especially because they do mention they only want adults, but children entering will most likely still be recorded, even if it won't be used, and children have special protections under the GDPR), and people who otherwise cannot consent. As this is not just for security purposes, like a normal CCTV surveillance system used in almost any store, but instead is supposed to be (manipulated) content for potentially millions to see, I see a heightened need to inform people via other means, even verbally inside the store. I'd also be curious to know whether AI will be used or not, and whether people agree to also give their likeness for that, or to be associated with product ads this way. Third, lets think about the customer base for kiosks. There will be children buying candy, chips and sugary drinks, of course, and many people just buying magazines, cigarettes and the like in a mundane way; but many, many times, these shops are also the main hub for homeless people, people suffering from alcoholism, and more. Poor people who are seen as Asoziale , based on fulfilling stereotypes like being from a migrant family, dressing in sweats, talking a certain way, chain-smoking, being seen as unintelligent, violent, with no ambition. It's always been fair game to many people here to make fun of people they deem inferior like this. So what is the entertainment value supposed to be here? You wanna make fun of Renate and that she buys 5 packs a day? You wanna milk a homeless guy's heartbreaking story for clout? You wanna show how the local alcoholic deteriorates within months? It's extremely clear that they are searching for content that as a German, you'd otherwise see on RTL2 reality TV formats, or the BILD newspaper. The fact that the company markets this under the guise of "funny feminism" and that it's for the "Girls, Gays, and Theys" is disgusting. You wanna expose these groups to absolutely disgusting comment sections and turn them into a spectacle to be gawked at? Women to be slutshamed, gay men to be called slurs, trans and nonbinary people ridiculed? I don't know what to say. You aren't woke, you aren't funny, you aren't cool, you aren't an ally. The main motive is to exploit vulnerable people, cut together compromising and dramatic stuff, expose and embarrass people and earn money doing it. People just wanna shop at the closest kiosk, yet now have to go elsewhere because otherwise, they might get a shitstorm in a comment section they can't control, because they were coaxed into some kinda reaction of wore the wrong clothes or bought something deemed odd or bad. Even if the video with them in it does well and is nice, what's the payoff? You are the product, with no compensation! You pay them money for a product and then they try and make even more money off of you! Social media has rotted the kiosk owner's brain and made them see other people as mere opportunities to go viral, no matter what that could possibly mean for them. I hope no one ever enters that store again, and I hope more people alert the LDI NRW to the absolute carelessness of people's privacy rights. This business model should not survive. I don't care that people can just opt not to go in if it bothers them. I think this is a threat not every customer can correctly interpret and decide on, it is predatory, and I don't want this to become a trend and my shopping experience to turn into someone's entertainment and money source. Published 17 Aug, 2026

0 views

27 down, 17 more to go

The day is August 16th, the time is 6:42 am, and the temperature is a very enjoyable 16°C. The sun is about to rise over the mountains, and I can see it appearing behind the car I just parked at the same parking spot where we ended our walk back in June. It’s been a long summer so far, and not a great one when it comes to hiking. It’s been hot. Probably too much. We hit a few records here and there, hitting 40°C more than once and hiking when it’s more than 35°C and with more than 50% humidity is a miserable experience. But I wanted to go hike, I felt the need to be out there moving through nature, putting some steps in. And that’s what I did. This is going to be a strange hike since it’s part of the 44 churches loop I’m walking, but it’s not going to be a point A to point B hike, and I’m also hitting all the churches of the sixth segment of the loop plus an extra one for reasons I’m gonna explain later. But we have a long way to go, so we better get going. The first church we’re going to visit is very close to the start, just 1km away, and it’s one I never visited before, even though I drove through this place a million times. Walking along the river this early in the morning is very enjoyable, and it’s nice to start a hike on flat ground, rather than immediately going up. Barely 10 minutes into this hike and we’re visiting the church of Sant’Antonio Abate (22/44). The church itself is not all that remarkable, but there’s an interesting plaque on a wall under the porch as a reminder of the history of this place, and also a window was open so I managed to take a picture of the inside. With the first of the six churches we’ll see on this hike behind us, it is time to work our way up. We walk through a few lovely houses (never been on this part of the town), and we find the trail that is gonna take us up to the second church. Noticed a couple of what used to be houses completely reclaimed by nature and was reminded of a recent podcast episode I listened to during a long drive where they were talking about abandoned places amongst many other things. It’s fun to imagine how these towns and villages would look if we all left for a few decades. The narrow trail widens up, we intersect one of those service roads, and then we work our way up to a lovely stretch through a series of open fields. I’m glad I’m walking this part of the trail early in the morning. As always, we’ll encounter random Jesuses and Maries throughout the hike. This is just one of them, placed there for who knows what reason. This part of the loop is very close to home, I walked these fields many, many times before, and I also walked past this cabin so many times and always thought it would be so nice to live in such a place. At some point I should probably ask the owner if he’s willing to sell it. It’s such a lovely place. Just a bit more than an hour into this walk and we have reached the church of Sant’Andrea Apostolo (23/44). This is a lovely church, but the whole area surrounding it is quite neglected, unfortunately. There would also be a lovely view of the whole valley just outside its porch, but sadly everything is overgrown. And that’s the reality of this whole area: it has so much potential, but everything is so neglected, and it feels such a waste. With this church we have also concluded the first uphill part of the hike, it’s going to be mostly flat and downhill for a while before we start climbing back up again. Plus the third church of the hike is just 10 minutes away from here. And in fact, just like that, we’re outside the church of Santa Lucia (24/44). The outside is not unlike many of the other churches, but this one has a gorgeous altar. I couldn’t take a decent picture, but there are a few on the site linked above, so make sure to go click it. We connect back with the trail, and now we have quite a long way to go before we reach the next church since it’s almost 8 km away and we need to go back down this side of the mountain, into the next valley and back up again. Thankfully this walk is gonna be almost entirely inside the woods because the day is warming up fast. But for the next hour and a half we can enjoy this slow descent into the nearby valley. A bit more than 10 km into this walk, and it’s about time to start going uphill again, in the direction of the village of Costne and the fourth church of this hike. I have never been up here before, and it always amazes me how many old and abandoned houses there are in these valleys. I think I should start doing a more comprehensive work of photographing these places just to document how an area slowly dies down. Maybe that’s a project for future me. We just passed the three-hour mark, and we have reached the church of San Mattia Apostolo (25/44), which I think is the only one so far with 3 bells. I thought this was a nice spot to take a quick break, eat something, drink a bit of Gatorade and also, since I had phone signal, FaceTime my friend Mattia who was about to hop on a plane for South Korea. I’m walking around the valleys while he’s travelling around Asia at the moment, and if you’re interested in that, you can follow his adventure either on his site or via his newsletter . This 20-minute break was nice, but we’re now moving again, going down for a little bit before starting the biggest climb of the day. We’re passing close to the village of Tribil, which is where this segment of the loop ends, but we’re headed in the opposite direction, up Mount Hum, a place that’s filled with WWI history, and then down on the other side to reach the fifth church of the hike. The sign at the base is telling me it’s gonna be a 1-hour hike to reach the summit, but I have different plans. I hiked this mountain before, but it was a lot more slippery back then. Now everything is super dry. And we’re at the top. It took me 30 minutes. Easily the hardest part of the hike so far, and my HR can confirm that. But that’s behind us, and we now need to go back on the other side before going up again to reach the halfway point of this hike. We’re down the mountain, back on the road, and I can see the church from here. Been up there quite a few times lately, it’s a lovely spot, especially late in the day. Final push uphill, and we have reached the church of San Volfango (26/44). We’re 5 and a half hours into this walk, and we have hikes 20kms so far. It’s time for another quick break to catch my breath and also finish this Gatorade I have with me. Most of the rest of the hike is gonna be downhill from now on, but we have quite a long way to go before we’re done. So down the stairs we go, through the abandoned—I think—village of San Volfango, onto a very sunny road and back to where we exited the trail out of Mount Hum not long ago. I’ll have to walk back on the same trail for a tiny bit—something I usually try to avoid—because the alternative would mean walking quite a long stretch on paved road and that’s not an enjoyable experience. On the trail, I stumbled on a big group of people that was taking a break, lying down in the shade, and they all had horses that were also chilling on the trail. Lovely animals. Horse riding is something I’ve never done in my life but would love to do one day. Feels such a great way to experience nature. Almost 25 km into this hike and we’re about to enter Tribil, where this segment of the churches loop ends. I passed through here just the other day, during a silly walk with my dog, but that’s a story for another time. And here we are, at the parking spot next to the cemetery where our hike should end, if I was a reasonable person. This is also where the next hike will start. But this time, we’ll keep going, and we’ll make it a full loop, going back to my car. And on our way there, we’ll hit a church that’s part of the 7th segment of this loop. There are two reasons why I’m doing this. The first one is that I’m an idiot and I like to do silly things. The second is that the second half of these walks all have idiotic routes that make absolutely no sense. And as you’ll see in future newsletters, I had to tweak the proposed paths quite a lot to make them walkable in a decent way. And hitting this extra church today allows me to make the next walk much more enjoyable. But that means today we need to keep going, so down the trail we go. The trail is mostly uninspiring now, just a long, relaxing, slightly downhill stroll through the woods, as we’re heading towards the village of Presserie, very close to the final church of the day. We’re out of the woods now, back on paved road, and about to approach the final uphill stretch of the hike. My feet are starting to hurt a little bit. We are a bit more than 30 km into this hike. Almost 8 hours into this hike and we’re finally outside the church of San Paolo Apostolo (27/44). Thankfully, there’s a fountain just outside of it. I definitely need to refresh myself a little bit. The only thing left to do now is to walk back to the car. Which means leaving the church behind us, going through a short stretch into the woods, back onto the main road, then through a weird side trail again and then back onto paved road. We’re down at the bottom of the valley, only a couple of kms left to walk on this sunny road and then we’d be done. With the walk. My feet are now hating me, but that’s a problem for later me. And just like that, we’re back at the car. The time is 4 pm, the temperature is 35°C, and we have walked almost 38 km and ascended almost 1600 meters. I’m honestly less tired than I thought I’d be and kinda pissed that I didn’t hit 50k steps for the day. But the good thing is that I need to walk the dog later in the day, so we’ll get there. As always, photos of the walk are on the shared drive folder (these are all kinda shitty I have to say, not sure what happened, I definitely need to buy a camera) and the data from my watch is available at this link . See you next time! You love the outdoors and RSS. You're one of the special ones.

0 views
Kev Quirk Yesterday

2026-08-17 11:29: I regularly get marketing emails off the back of this site. "Here's a free license...

I regularly get marketing emails off the back of this site. "Here's a free license for our tool" or "can we put a link to our service in a post?" At least back then they were slightly human. Now all I get are AI generated shitty emails with no personality whatsoever. Not bashing AI, I'm bashing the lazy marketeers. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Stratechery Yesterday

Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model

Stripe is reportedly acquiring OpenRouter, an implicit bet on a future market of models and the chance at Aggregation.

0 views

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

Friday's big release was Qwen 3.8 27B , an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor Qwen 3.6 27B was impressive. Qwen's self-reported benchmarks for this model are eye-opening. They show a boost from both Qwen 3.6 27B and the closed-weight Qwen 3.7-Plus, which was one of Qwen's strongest models of any size as recently as May this year . It will be interesting to hear what independent benchmarks have to say about the model. I've been running the model on two different machines: my 128GB M5 Max MacBook Pro, and an NVIDIA DGX Spark . On both machines I'm running LM Studio and their 17GB Q4_K_M quantized build . I also tried using directly on the Spark. Qwen's documentation describes the model as defaulting to for the reasoning effort, and the LM Studio GGUF I've been trying preserves that default: Qwen3.8 comes with official support for , which can be used to adjust reasoning depth and control cost: This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware. I've been finding the results extremely entertaining. I quickly ran into problems with LM Studio's default context limit of 8,192 tokens - Qwen was using them all up thinking about even the most mundane of problems. I loaded the model with the full 262,144 maximum context length and that problem went away. Here's the pelican riding a bicycle SVG I got from my first attempt with that increased context length. It took 21 minutes to generate, using 22,276 reasoning tokens to produce 3,223 tokens of output. You can read the reasoning trace here . This is by far the best pelican SVG I've been able to generate with a model that runs on a local machine - and this Qwen is pretty small, just a 17GB file on disk. There's a lot to like about this: Was that worth waiting 21 minutes for? Absolutely not. Here's that same prompt run with reasoning turned off - transcript here . This one produced 3,715 tokens and took 137s - just over two minutes. And just for completeness, I used OpenRouter to run the same prompt through the much larger Qwen 3.8 2.4T-A95B (released last week ) and got this snazzy animated SVG : Your browser does not support HTML5 video. I said Qwen at xhigh has a tendency to over-think things, but how bad really is it? I tried a much simpler prompt, again with that default extra high setting: Qwen's reasoning trace started like this: The user is asking for an SVG drawing of a circle. Simple request — but I want it to be a carefully crafted piece. Let me make something that goes beyond just : a single self-contained SVG file with character — maybe a geometric "circle study," with subtle animation, layered rings, and a distinctive palette. Keep the scope right: they asked for an SVG of a circle. So the core is a circle. But I can add craft: concentric guide circles (like a compass/geometry drawing), tick marks, a soft gradient fill on the main circle, restrained ambient motion (a slowly rotating dashed ring, pulsing glow). Respect prefers-reduced-motion? That's more for CSS; SVG SMIL or CSS inside SVG will do. Palette options: deep teal ink on warm paper? Or bold vermilion circle on off-white with navy construction lines — Bauhaus/compass-drawing vibe. Let me go with a "geometric study" look: cool slate background, or bright paper white? Paper white is fine if it's not the cream-and-terracotta combo. [...] Several minutes later it produced this absolutely beautiful animated circle, which was entirely not what I had asked for! Your browser does not support HTML5 video. A fun way to test a vision model is to see how well it can return bounding boxes around items in a photograph. I've seen previous Qwen models deal well with this, so I decided to put it to the test drawing bounding boxes around some pelicans. I've seen asking for 0-1000 scale produce good results in the past. I tried this: Here's the reasoning trace , which produced this: This is such a good match . Here are those boxes rendered on top of the photo: That visualization of the bounding boxes was taken using a new custom tool that I had Qwen 3.8 27B build for me, running offline on my laptop. I forgot to dial down the thinking effort so it was massively over-engineered , but it did manage to produce this full interface from this single prompt : This screenshot shows one of the features I did not ask for - a demo scene, for if you don't have a photograph to test the tool with: Here's the relevant segment of the thinking trace, where it decided to draw its own pelicans purely because I had used the label "pelicans" in the example JSON I gave it in the prompt: Also a "load sample" that uses a known image? Can't depend on external images, but… the image URL input is user-provided; I could add a "try with sample" button [...] Hmm, I can draw a simple scene on canvas, export it as a data URL, and load it into the image — that's self-contained and demo-able! [...] But the user's coords are for an actual pelican image; a generated placeholder can still demo the scaling. Generate a 1000x1000 placeholder: gradient water + two blob-like "pelican" silhouettes placed at the given bboxes (using the same scale — cute: silhouettes at the exact 0-1000 positions, showing the boxes align). This makes for a fun, self-contained demo. Keep it simple: sky gradient, sun, water, two pelican-ish shapes (ellipse body, circle head, beak). Place at bbox centers. (I'm slightly nervous that models around the world might have a bias towards drawing pelicans at any chance they can get, brought on by nearly two years of exposure to my own stupid benchmark.) Is all that over-thinking necessary? Maybe it is, at least a bit. I tried with reasoning turned off and got this version , ( transcript here ), which nearly works but shows the boxes in the wrong place: So without reasoning it didn't quite one-shot a working tool. I'm sure it could get there with some follow-up prompts, but this is a good example of how reasoning can make a difference. One of the biggest questions around local models is whether or not they have enough horsepower to successfully run a coding agent loop. Coding agents require long context, strong code generation support and reliable tool-calling. On paper Qwen 3.8 27B has all three of these, so is it up to the task? My initial experiments with Pi have been very promising. I chose Pi because it has a shorter system prompt than most other options, making it a better fit for trying out smaller models. I configured Pi to use Qwen 3.8 27B running in LM Studio on the Spark (shared via ) by adding this to : Then ran in my folder and prompted: After a sequence of reasoning and tool calls that accessed a bunch of different files it produced this reply , which is very solid. Just one problem: I wanted to share that transcript. So I pointed Pi and Qwen 3.8 27B at the JSONL transcript file in and prompted: And it built and tested this pi_jsonl_to_md.py , which did exactly what I needed. Here's that session transcript , published using the tool that it created. So far this is all looking very promising. We have a 17GB model that runs on high-end consumer hardware and can write code, drive tools, annotate images and generally do everything that I need from an LLM for getting real work done. There's one very significant catch: it feels slow - especially when it starts over-thinking, but even without that it's not particularly sprightly. I've been getting around 15-30 tokens a second from LM Studio. That's not terrible, but it's slow enough that it's going to be hard to win me away from hosted API models, which can return results a whole lot faster. Artificial Analysis track token speed and show OpenAI 5.6 Sol at 74 tokens/second and 5.6 Luna at an impressive 184/second. The good news is that the community have been exploring ways to speed things up since the model was first released two days ago. One of the most promising optimizations is baked into the model itself. Qwen supports Multi-Token Prediction , an architecture trick where a cheaper mechanism guesses several tokens ahead and the main model can then quickly verify if the guesses were correct. This can have quite a dramatic effect on inference performance. Based on this tweet from creator Georgi Gerganov I tried running the model with MTP like this on the Spark: And sure enough, this gave me a significant boost. I had GPT-5.6 in Codex run a comparative benchmark on the Spark and the server outperformed the LM Studio default GGUF by around 72%. I expect we'll see a whole lot more innovation around serving this model faster over the next few weeks. The MLX community likely have some tricks brewing as well. The fact that a 17GB file can do all of this stuff on my home machines is a miracle . Once again, I'm delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models - today it can run on a capable laptop. The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That's the catch with these dense (non-Mixture-of-Experts) models - they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard. The most important thing about Qwen 3.8 27B is what it demonstrates . We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file. The models at this size continue to get better at an impressive rate. We don't need to spend half a million dollars on datacenter-class hardware just to run a competent model. You are only seeing the long-form articles from my blog. Subscribe to /atom/everything/ to get all of my posts, or take a look at my other subscription options . (default): for complex tasks demanding thorough analysis : balancing accuracy and speed : efficient reasoning optimizing for speed and cost The bicycle frame is the right shape It has legs on each side of the bike - that's very rare Good, clear pelican pouch The wings extend to touch the handlebars! The motion lines are behind, not in front It has a tasteful background - nice sun, clouds, hill, flowers and grass.

0 views
Unsung Yesterday

“I think there’s a lot of value in these.”

Speaking of user interface guidelines, developer Matt Sephton and others compiled many of Apple’s human interface guidelines , starting from 1980, all the way to 2014, including some goodies like early drafts, NeXT, Newton, and so on. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/1.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/2.1600w.avif" type="image/avif"> Elsewhere, designer Geof Crowl put together his own list that goes broad instead of deep, including style guides other than Apple’s, too. The list was made in 2020 and a few links are already broken, but there are some gems here like the modern Designing for Playdate , or the classic Zen of Palm from 2003. You can learn a lot just by grabbing one and scanning it. Let me add a few more I know of: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/3.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/i-think-theres-a-lot-of-value-in-these/4.1600w.avif" type="image/avif"> #interface design #mouse #principles #style guides Amiga User Interface Style Guide (1991) Magic Cap Concepts (1995) with a strange subtitle “Everything right is wrong again” Java Look and Feel Design Guidelines (1999) Windows Interface Guidelines for Software Design (1995)

0 views
Unsung Yesterday

“If nothing happened, you probably did not press the button quickly enough the second time.”

I like learning new things, and I like learning new old things. I was looking at Apple Human Interface Guidelines from 1987 and this passage caught my attention: The most common use of double-clicking is as a shortcut way to perform an action. For example, clicking twice on an icon is a faster way to open it than clicking once to select it, then choosing Open from the File menu; clicking twice on a word to select it is faster than dragging through it. I knew that double click an icon was a shortcut to the first action (typically Open), but I never really thought of double-clicking a word as a faster way to drag across to select it – even though, in hindsight, it makes perfect sense. Another vintage thing I learned of recently from a coworker is this, also covered in the 1987 HIG: If the user begins a double-click sequence, but then drags the mouse between the mouse- down and the mouse-up of the second click, the selection becomes a range of words rather than a single word. This doesn’t feel (to me) like a very pleasant gesture to perform repeatedly, but what feels nice about it is that it automatically snaps the selection to the endings of the words: Part of me would prefer this to be the default behaviour when selecting more than 3 words, or so, so you could be less precise. Anyway. The double clicking to perform default action applies to a lot of lists of things. Here are some examples from Scrivener, Word, and Lightroom – you can double click on each of these items to proceed, without having to select and click the button: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/2.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/3.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/4.1600w.avif" type="image/avif"> But sometimes the creators of such dialogs forget. Here’s Screen Sharing in MacOS, and a notification in Chrome where only the slow path is available – double clicking on items doesn’t do anything: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/5.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/if-nothing-happened-you-probably-did-not-press-the-button-quickly-enough-the-second-time/6.1600w.avif" type="image/avif"> The tricky part about not being a good citizen of a shared user interface is that those omissions aren’t just local to your app – they can ruin the gesture in other places, as people’s fingers learn to distrust it in not just your app, but in general. (The quote in the title is from Apple Macintosh User’s Handbook .) #flow #mouse #text editing

0 views
Farid Zakaria Yesterday

DEFCON34 wrap-up

I recently came back from DEFCON34 and the nix.vegas community. The talks I gave are now online if you are interested in watching them. 🙌 Many thanks to all the organizers of DEFCON34 and nix.vegas. This is our, the Nix community and mine specifically, second year at DEFCON34 and it was a blast. To be honest, I barely interacted with the rest of DEFCON because I was so busy with the Nix community. The talks, the hallway conversations, and the in-chance encounters were all amazing. One particular story, was that Carl Dong happen to be walking by the Nix Vegas village as I was giving my talk on Guix by Nix . He was a Bitcoin core developer and was one of the contributors responsible for the Bitcoin Core reproducible builds project that leverges Guix . 1 For those that don’t know: nix.vegas is the Nix community that runs within DEF CON in Las Vegas, hosted by the SoCal NixOS User Group and Distractions, Inc. This was its second year: DEF CON 33 ran under the banner “Rebuild the World” , and this year’s theme was “Escape Your Fate” . The full playlist is on YouTube . Note If the sound is a bit off or weird, this year DEF CON experimented with “silent” talks. Each talk was broadcasted and attendees had to wear headphones to listen. It was a bit weird giving talks to a quiet room. 🤷 Summary : Nix’s absolute paths buy us reproducibility, but costs us the ability to put the store anywhere else. You can change the store prefix today, but it changes the hash of every single derivation in the closure down to , so you get to rebuild the world before you get to run . How can we circumvent this? The talk walks through in and upstreaming support in the Linux kernel via a eBPF-based solution. Further reading: Linux kernel will support $ORIGIN, sort of . Summary : What was meant to be a lightning talk on guix-transfer and GuixPkgs but went a little over. This is our project on rewriting Guix derivations into Nix derivations so that every Guix package becomes buildable by Nix. This lets us include their source-bootstrapped JDK for instance, which nixpkgs does not have. Further reading: Guix by Nix and GuixPkgs: every Guix package, as a Nix flake Summary : This talk is a bit of a rant, but it is given in good faith with a dose of humor. The core claim is that we optimize Nix and nixpkgs for social comfort and broad appeal, and we pay for it in technical ambition. Further reading: How to piss off your Nix friends . Looking forward to next year. Three talks in two days was a little ambitious, but I would do it again. Everything lives on my talks page alongside their slides and the rest of my talks. He was pleasantly surprised and happy to hear that Nix also has reproducible builds that start from stage0 .  ↩ He was pleasantly surprised and happy to hear that Nix also has reproducible builds that start from stage0 .  ↩

0 views
Kev Quirk Yesterday

2026-08-16 14:35: Working great 👍🏻

Working great 👍🏻 Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Harper Reed Yesterday

Note #739

Hanging out at the boyhood home Thank you for using RSS. I appreciate you. Email me

0 views
Kev Quirk Yesterday

2026-08-16 14:22: Last November I replaced my oldest son's PC with an £88 iMac that was so...

Last November I replaced my oldest son's PC with an £88 iMac that was so successful that I'm now doing the same thing again for my youngest son! Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Stone Tools Yesterday

AppleWorks on the Apple II

Dig, if you will, a picture. It's November 1983 and throngs of people are engaged in literal fisticuffs, " dozens of people were injured, some even hospitalized ," to claim the last homely, pug-nosed Cabbage Patch Kid from the local Zayre's for Christmas shopping. In video footage, a store manager climbs a display, brandishing a baseball bat, trying to will order onto the chaos. The riots make the local, then national news, and eventually have a page on Wikipedia . It's a striking exclamation mark to punctuate 1983, with dark implications for the coming new year. Had we sunk so low? Then, at the start of December, a Christmas miracle. The hottest pop star in the world released one of the most influential music videos of all time on MTV. You know the one, where Michael turns into a werewolf, but then he's a zombie, and they do that dance, and the electronic beat "ur URRR adugga dugga dugga dugga dugga dug ur URRR adugga dugga dugga dugga dugga dug", and the Vincent Price laugh, and those crazy eyes at the end. The video for "Thriller" released simultaneous to the riots, and karmic balance was restored. Huzzah! Maybe 1984 wouldn't be so bad after all. 1984 wasn't just "not bad," it was radical. I tried to find "cutest adorable chubby-cheeked preschoolers Thriller dance" but this will do. Pop-culture nerds, in particular, would gorge themselves that year with Ghostbusters , The Terminator , Gremlins , Nightmare on Elm Street , Beverly Hills Cop , Revenge of the Nerds , Splash , Red Dawn , The NeverEnding Story , The Adventures of Buckaroo Banzai , Lynch's Dune , Karate Kid , Repo Man, This is Spinal Tap , Sixteen Candles , The Transformers toys and show, Teenage Mutant Ninja Turtles comic, Stop Making Sense concert film, Prince's Purple Rain , Do They Know It's Christmas? , Neuromancer , Elite , Karateka , King's Quest , the Soviet boycott of the Olympics , and the first MTV Video Music Awards. "Thriller" lost "Video of the Year" to "You Might Think", but did win "Best Choreography" and "Viewer's Choice." Oh, and Apple aired a commercial during Super Bowl XVIII. You know the one, with the skinheads in grey, and the oppressive fascism, and the woman in red shorts, and the sledgehammer, and that electronic siren "BEE oooh.... BEE oooh", and "1984 won't be like 1984." Apple introduced the Macintosh, a system Steve Jobs said was so easy to learn, "You can sit your grandmother in front of this computer and she’ll figure out how it works." That's an interesting way to describe a computer that internal documentation shows "we will not attempt to position the product in any way as a 'home' computer." Meanwhile, that same year, the Apple II series sold another one million units while Apple was taking its third stab at obsoleting those machines. Having shrugged off the Apple III, then the Lisa, Apple II fans would be forgiven for looking at the Macintosh's anemic launch software library (essentially just MacWrite and MacPaint ) and shrugging in agreement with Clara Peller's 1984 catch-phrase, "Where's the beef?" Tom Weishaar, writing for Open-Apple newsletter in July 1985 (a rough year for Apple), "For six years now Apple's management has been trying to build a computer better than the Apple II. In all these years, after all this money, only the Apple II has ever shown a profit. Buyers are looking for tools they can use. The Apple II has plenty of power to be useful." VisiCalc was early, clear proof that innovation isn't measured in megahertz. By 1984, developers had a lot of expertise in making the most of limited hardware. Look at the strides made in six years of software research on the Atari 2600, for example. Practice makes perfect. Rupert (later "Robert") Lissner heard the call. While the Macintosh was an interesting promise of a computing future that may or may not come to fruition, there were millions of Apple II owners with immediate needs. AppleWorks fulfilled those needs, growing a strong fanbase, spawning dedicated newsletters, lots of books, add-ons, and more. When Apple finally abandoned their 8-bit line, AppleWorks singlehandedly kept those machines productive for another decade . My only experience with AppleWorks was in the early OS X days, and that version had nothing to do with what we're looking at today. That and this only share a name; the Apple II version holds a legacy. Version 3 of AppleWorks was, according to a Compute! Magazine review, when the promise of "integration" finally paid off. Much of what AppleWorks 4 and 5 do, 3 can do with plug-ins. Version 3 is also when the program became Y2K compatible. It should be more than enough to develop a strong understanding of what made it such a favorite among the Apple II faithful. Thus far, I've only looked at one other "integrated" software package: Pipedream , on the Acorn Archimedes. It put forth the hypothesis that word processing, spreadsheets, and databases are not three separate applications, they are one and the same, chemically. While I was interested in its proposition, I was cool on its execution. It's an acquired taste. Where Pipedream blurred the line between applications, AppleWorks is very much three apps in one; a veritable turducken of productivity software. Maybe I'm a little unfair with that analogy. A turducken is a forced alliance of three fowl under, let's call it "extreme duress." AppleWorks, on the other hand, is a joyful alliance, with each individual app supporting the other two. I'd argue the word processing benefits the most, but to categorize this as merely three applications which happen to ship in one box, would miss the turkey for the duck. With 128KB of RAM and an entire suite of software loaded up, AppleWorks has 23KB remaining to do real work, as indicated by the memory counter in the bottom right of the screen. That's not a lot, but it's a bit dependent on how hardcore your work is. For example, my pre-edited text for this blog post hit 9,000 words and 43KB; the database merge (I'll talk about later) hit 24KB. Both are substantial projects, so if you're just dashing off letters, giving publishers a chance to get in the ground floor of your burgeoning career as the author of legally distinct James Bond-alike, James Epoxy, you'll be fine. Chronicling Agent Epoxy's actual adventures will likely require you to scope documents chapter-by-chapter. In the wake of the release of AppleWorks , a decent amount of third-party software was published to "add value" to this triple-threat. Using ad-hoc plug-in strategies (not originally part of the program proper), the already stuffed AppleWorks could become almost an operating system unto itself, hosting not just enhancements to the core functionality, but entirely new applications like screen savers and a MacPaint clone. Other companies took a less invasive approach to enhancing AppleWorks , offering pre-built documents ready for mail merge, recipes, phone directories, and the like. I saw a reference to a fantasy football management suite called FantasyWorks , but couldn't find any disk images for it. I suspect there's an entire world of unarchived AppleWorks enhancements sitting around on forgotten floppies. One product I did dig up, and which inspired my project for this investigation, promised to turn AppleWorks into HyperCard . When I learned of this program's existence, I had the reaction of a dog being asked, "Do you wanna go outside?" StoryWorks , written by Robert C. Moore and published by Teachers' Idea and Information Exchange, promises we can "create incredibly powerful 'hypertext' applications (stacks) (and) create 'knowledge base' stacks." That is a tall promise for an 8-bit machine, and it is one the provided example files don't quite live up to. What it can do is create "choose your own adventure" style stories. It takes a bit of boilerplate to adhere to StoryWorks's programming rules, and I will need to keep track of which "segments" (like HyperCard "cards") link to which. This sounds like perfect tasks for the database and spreadsheet to help manage. Then, I'll do some writing and formatting in the word processor, and next thing you know the potential and promise of AppleWorks's "integration" will be realized. Or so the theory goes. Navigating AppleWorks's file management can feel a little pre-Cambrian, not yet evolved into the complex, multicellular operating system you're reading this very blog within. Hardware-wise, we have the disk and the RAM. Files must be "on the desktop" to be used, which suggests some kind of "load from disk into RAM" step is necessary. That is precisely the case. First, we have to let AppleWorks know which disk is our "current disk," shown in the upper left corner of the "desktop" screen. It's akin to interfacing with a single folder in your file system, and the visual metaphor of tabbed folders is used to assist with navigating the Desktop. Once we have files on the desktop, fast switching via lets us bounce rapidly between documents making edits, while AppleWorks preserves our changes. Ready to quit? Not so fast, buddy. Sure AppleWorks appears to have preserved your work, but did you SAVE YOUR WORK ? Thus far, our changes are only committed to the Desktop and, as established when we loaded from disk into Desktop, the Desktop is not the disk. From the Desktop, we need to "Save Desktop files to disk" to really save onto disk. Earlier, I noted the "default folder" and here's where it will come into play. Saving work to disk will save everything to the default folder, regardless of its original position on disk. I wound up duplicating files into lots of fun, new locations on disk, rather than overwriting the originals, as I learned the program. But, at least my files are safe. They were saved carefully, after all. Two core tenets of AppleWorks 's approach to integration are its clipboard and keyboard shortcuts. Both work hard to bridge the natural divide between applications. The clipboard lets us move data between applications, and the keyboard shortcuts let us move muscle memory between applications. This mostly works, as when doing cut/copy/paste, text insertion/deletion, finding information, printing, and other relatively generic tools. The shared keyboard shortcuts fall apart, somewhat, in their overly ambitious attempt to enforce uniformity across all apps, even when they don't make intuitive sense. Let's look at the 'Zoom' command as a prime example. First, think about a command called 'Zoom' and envision what the result should be when invoked in a word processor context, then in a spreadsheet, and finally in a database. What did you picture in your mind's eye? What does it mean to 'Zoom'? Contrast that with to 'Find' text, which does exactly the same thing in all apps. Keyboard shortcuts aren't evenly distributed across modules, either. Printer options are available in the word processor and spreadsheet, but not the database. "Find" works across all modules, but "Replace" doesn't work in the spreadsheet. will split a window in a spreadsheet, ala VisiCalc's split window function. However, the word processor does not accept this command, though it would be a very reasonable expectation for that to work, as we saw in PaperClip on the Atari 8-bits. I certainly recognize that each program will necessarily have unique needs. However, I'm not convinced the mental fortitude required to remember how one command differs from application to application is any easier or harder than just learning a new keyboard command specific to each application. At this point, I've written thousands of words in AppleWorks and I can say confidently: the word processor is good. I have throttled the emulator to run at 1x speed, so I'm not completely head-in-the-sand about the reality of using the program on period equipment. Whatever horsepower I throw at it, it remains performant and a joy to type in. Of course, we have the usual list of modern gotchas, like the lack of international character input, and the printer-centric formatting options (which do not survive ASCII export). But the basic act of getting words on screen is smooth, along with minimal chrome which keeps us informed of the state of the writing environment. The chrome serves the document, ensuring the writer is never lost. The upper two lines show the ruler, what file we're working on, what mode we're in, and what hitting will do right now. In the above case, here in the Printer Options screen, will return me to the REVIEW/ADD/CHANGE screen. This clear explanation of what navigation buttons will do was something I appreciated about Bank Street Writer , and I'm happy to see a mature version of it in AppleWorks . The rule shows our tab stops and their respective styles, but doesn't show an indicator for the right margin. Editing tab stops is as easy as editing a line of text, thanks to fixed-width type. Just type a character for the tab styling you want at the place you want it, including centering and decimal alignment. The bottom displays a running line count and where the cursor currently sits. The bottom left has what seems like pretty useless information to me, "Type entry or use Apple commands." as a gentle reminder of how to use the program, I guess? The bottom right shows how to reach Help. That gives the UI four lines, with 20 devoted to writing. It is enough, though I do wish for a live word counter. AppleWorks 3 brings integrated spell check, with reviews of the time saying it is better than any of the third-party spell-checkers that preceded it. That's a strange note, because I find its usage convoluted. Invoked by , to "verify" the document, the "Options" it presents are obtuse and unintuitive, but all its asking is if we want to replace words one at a time or in a list, and do we want a post-verification summary of everything changed? If so, how do we want the summary presented, printed or on screen? The "Summary" is where we finally can see our word count, plus a list of all the weirdo words we used and how we corrected them individually, along with the occurrence count for each word. I'm not clear what I'm supposed to get out of knowing I used the word "turducken" 8 times. Maybe this is to help encourage the writer to mix things up? I refuse. As is typical of word processors of the day, a vast amount of the program's features are related to printer formatting and control. I do not own a printer, and "printing" to an ASCII file strips away most of the fun stuff, like bold and superscript. The spreadsheet portion is both amazing and boring. It's VisiCalc , with minor changes to meet the keyboard shortcuts (i.e. no slash menu command) and basic UI chrome of the suite. It does almost everything VisiCalc does, on a much larger worksheet and without quite as strict a memory barrier. If you have the RAM, even in this economy , you can eat VisiCalc's lunch. It also has many of the same frustrations as VisiCalc , including row-based vs. column-based calculation order, which can require a full document double-calculation to synchronize early formulas with later cell values. We're also stuck with the archaic (even by this time) "one column width for all cells," rather than per-column width adjustments. There are no graphing tools, and what options are available are so limited as to be effectively useless in a Lotus 1-2-3 world into which AppleWorks 3 was born; we'll need to rely on third party solutions to this problem. Copying complex formulas with relative cell references still requires manually selecting "relative" for every single reference in every single cell copied to a new position. Just a quick note to those cloning popular apps: you don't need to clone the terrible parts. Still, VisiCalc once turned the Apple II into a must-purchase investment and birthed an entirely new genre of business software. Now, that watershed event is collapsed into one subset of bullet points in a longer feature list on the back of the product packaging. The flat-file database module is uninspired, which is surprising to me because it is the module upon which this entire program was built. With a great word processor, and effectively a full clone of VisiCalc included, I expected similar depth from the database. Setting up fields and records, searching, sorting, and filtering, are all simple enough to accomplish. Its adherence to AppleWorks's common keyboard commands makes it pretty trivial to search for records, edit text, delete records, and print. Yet there is so much it doesn't do, I can't help but feel disappointed. To start, fields cannot be assigned value types, something even dBASE on CP/M could do years earlier. Everything is a string, without even the crudest form of data validation. AppleWorks 4 would gain the ability to at least set a field to be a number vs. text, though still no Booleans or field widths, for example. When setting up a contacts database, a common use case is to have a "notes" field to remember things like family members, follow-up discussion topics, and the like. Thanks to a limit of about 70 characters per field, no single field in AppleWorks can hold that much information. I guess the solution is to set up half a dozen "memo" fields, just in case? That kind of workaround feels quite silly to me. The remainder of the module is focused on generating "reports" and "layouts," the difference being "for printing" or "for screen." This has its place, perhaps to show a subset of data and focus on the important stuff, or to copy some data over to the word processor. Otherwise, there's just nothing to get excited about here. I can't even pretend to be excited for the purpose of writing a fun blog. Moving on! The early days of productivity took some time to figure out how best to handle object selection and manipulation. On 8-bit systems especially, modality ruled the day. Today we've settled on the paradigm, where we select an object to manipulate, then choose the manipulation. Modality-driven software was the opposite. First, choose an action, then choose the object to affect, i.e. . To move a paragraph, we first select , to indicate our intention to "move." The system switches to a text selection mode, allowing us to highlight the text we wish to move. Finally, we position the cursor at the new location for the text, hit , and the text is moved. Its backwards, relatively speaking, but easy enough to adjust to. Thanks to AppleWork s integration, we can copy/paste between any two Desktop documents. It's a little strange, because copy and paste are both considered "copy" operations, via . We copy to the clipboard from application A, then we copy from the clipboard into application B, using menu prompts along the way to specify the direction of the copy. Copying has its quirks, though. Copying between two documents of the same type behaves as expected. Across application types, copying into the word processor fares best, and most closely meets expectations. But we still encounter broken formatting, like how a line of text that spans multiple cells in the spreadsheet will break apart into discrete, tab-width sized chunks in the word processor. What we actually want is to do is "print" to the clipboard, as backward as that sounds. Depending on the module, different print options become available, to define what part of the data we want to print, and how to format it. Once done, printing to is an option. The end result is much cleaner data that is closer to WYSIWYG over the normal clipboard copy command. "Print to clipboard" is the only useful cross-application option, in my testing. As the connoisseur of modern computer software that you are, I know you know that however good something is, it could always be a little better. Most software requires us to beg and plead for the developer to make the improvements that will ease our mortal burdens. Some software embraces the notion of "extension," allowing itself to be a mere vessel for thoughts yet unthought, ideas yet unrealized. When you bought AppleWorks , it never occurred to you that you'd even want it to do outlining, until you saw ThinkTank and now you kinda wish you had something like that. AppleWorks is shockingly accommodating to our wishes. One of the core features of Lissner's engine is a crazy-efficient memory management assembly routine. This gives, what by all rights should be "not enough computer," the superpower of fast app switching. Especially with a ProDisk (hard drive) attached, AppleWorks can become essentially a graphical shell for the Apple II. A number of companies took advantage of this and created various add-on packages for AppleWorks . Beagle Bros, JEM Software, Pinpoint Publishing, and PBI Software collectively published over 100 extensions. They had a lot of competitive overlap, but provided add-ons like: This was all well and good, but there was no official plug-in architecture for AppleWorks ; it was every developer for themselves. Installation routines were effectively "patches" to the base application, and there was no guarantee that patch A wouldn't step on patch B's toes during the installation process. It put some burden on the end-user to detangle things like installation order, or even compatibility with other add-ons of interest. Beagle Bros wanted to solve that problem once and for all. TimeOut was their solution to a simple, universal plug-in architecture for AppleWorks . Developers targeting TimeOut could rely on it to do the down-and-dirty interfacing with AppleWorks's memory management routines, without having to worry about patching or memory collisions. For the end-user, adding new TimeOut modules to their system was as simple as copying a single file to the AppleWorks directory. This worked like a champ, and eventually led to Beagle Bros being the contractor for AppleWorks 3 . As the main app developer, they could steer its growth to align with their own vision, and so TimeOut became part of the foundation of AppleWorks proper. This then made it super simple to produce updates of significant value to the end-user, because those upgrades already existed in the form of TimeOut extensions. And so, AppleWorks 3 received its first built-in spell checker as a result. As time went on, the outliner, expanded Desktop, better clipboards, and the like all became built-in features to AppleWorks . The big one, from my perspective, is TimeOut UltraMacros. I have it installed as an extension here in AppleWorks 3 , and in AppleWorks 5 it became a built-in feature. With it, we get a kind of mini-programming language, which includes if/then/else statements, variables, loops, value comparisons, and more than enough to fill out a 110-page manual. Macros work across all AppleWorks apps, and since the language is identical across apps, if you know how to script one, you know how to script all of them. Still a few more years until the 6502 reaches its maximum potential. With the powerful combination of AppleWorks + UltraMacros + StoryWorks , we have a veritable turducken of creative utility. In later writeups espousing the benefits of the AppleWorks plug-in system, it was suggested that future application authors would be writing exclusively with AppleWorks as the hosting application in mind. I can see the appeal. Unfortunately, StoryWorks didn't get that memo, so the workflow between AppleWorks and StoryWorks isn't the seamless experience I would have liked. They are two completely separate programs; quit one to enter the other for each phase of the writing/debugging development cycle. Let's all take a moment to re-appreciate our multitasking operating systems. Now, I've never written a Choose Your Own Adventure (CYOA) before. In browsing the books of my youth, I can see how AppleWorks's integration could be useful in designing such a thing. I can keep multiple documents open simultaneously, and jump quickly between them with . Here's my plan. With the spreadsheet, I'll keep a list of page branches. This will sketch out the core content on each page, with page numbers showing where each choice leads. In the database, I'll set up a template for the "programming" information for each page, based on the spreadsheet sketch. The database's 70-character limit on fields means I can't type the entire story into the database, but I can rough out a plot point or two. Then, in the word processor, I'll merge the database into a template to save me from tedious, manual boilerplate formatting. Once merged, if I've done my job right, that should build a skeletal game I can dry run through StoryWorks . "It's so simple, I can't believe people struggle to write games," the author scoffed. OK, look, I was out of pocket and I take it all back. In the previous section I was a naive child, unaware of his own ignorance. That was "Then Me" and he was a fool. "Now Me" is wiser, less prone to shooting his mouth off about things he doesn't understand. Making games is hard, I can admit that now. My biggest stumbling block is conceptually very simple: I don't know how to write a CYOA. What seemed like a clear plan in my mind turns out to not be that in practice. I thought, "Start with a rough sketch, start filling in details, repeat until done." would work. So, fine, all this really means is that I'm not a game designer, a fact I actually already knew. Looking through websites about how to create such games, everyone seems to have a different take. Spreadsheets as an organizational tool comes up a lot, as does sketching out decision trees, for which the "Paint" add-on might work, but I have a better plan. For my first attempt, I simply don't have the wherewithal to create something this intricate. My goal is to push the software, not my brain, so I'm going to be a thief and steal someone else's idea. Or rather, I'm going to steal someone else's structure. The original Choose Your Own Adventure series, the "fourth best selling children's series of all time," was established by Edward Packard in 1976. Relaunched by Chooseco in 2010 , the classics have been reprinted in handy box sets, and even new adventures have been published. If you've ever wished for a Cthulhu-themed CYOA, aimed at readers 9 - 12, your very specific prayer has been answered. The original CYOA books are also available on the Internet Archive's "Open Library" program. My intention was to get "inspired by" those, but instead I'm going to swipe the basic decision tree from one, and write my own story. There is a spy adventure called "The Deadly Shadow," by Richard Brightfield, that I will lean on for the heavy lifting of providing the structure for my story about low-rent super agent, James Epoxy. Ah, let's just go ahead and call it a parody at this point. There's no point in kidding myself that I'm about to do anything original. In working through my failed attempts, I encountered some frustrating limitations of AppleWorks integration. In AppleWorks , we have an option to set numbered "markers" throughout a word processing document. These are invisible tags that can be jumped to, by number, for quick navigation to known locations. StoryWorks relies on these to delineate story "segments." Just as AppleWorks can jump to markers, so too does StoryWorks use the same principle to jump to segments. The difference is that StoryWorks will only display text up to the next segment marker, effectively splitting the text into something akin to HyperCard "cards." And so, like HyperCard , StoryWorks also calls a collection of cards a "stack." In setting up my word processing template, I added segment markers where appropriate. Merging in the test data went smoothly, the result of which was a new document written to disk as ASCII. Maybe you already smell the trouble? ASCII conversion stripped out all non-printable tags and markers, rendering it StoryWorks- incompatible. That means adding StoryWorks markers must be done post data merge. Luckily, I have UltraMacros installed, so complex, repetitive, menu-driven marker creation is a simple self-defined keystroke away. In fact, because macros are just "a list of keyboard commands performed in order" I should be able to semi-automate find-and-replace placeholder text of my own design with a real StoryWorks marker. If I run that a few times I can prep my document for StoryWorks ingestion lickety-split. Does anyone still use that term? Where did that phrase come from, anyway? One thing I've quickly learned is that debugging is tricky. StoryWorks will parse the AppleWorks document and list its errors, but in practice I found that an early error can cascade, generating pages of errors. All we can really do is fix the first one and test again, hoping the later ones will disappear. So the proper strategy is to test early, test often. Before there is anything resembling a cohesive narrative, we need to make sure the skeleton of the project, as merged from the database, compiles and works. For this purpose, it actually does help to have even truncated page descriptions populated from the database, just to get a yes/no understanding if navigation is working. First, here's the template I've come up with. and placeholders in the template align with field names in the database. During a merge, possible field names are presented in a list for insertion, making it impossible to accidentally mistype one. means "don't insert a blank line if this database entry is blank." will always insert something , even if there is no entry in a record. Here's what a representative record looks like. You can see the field names are referenced in the template, above. As I mentioned earlier, the merge gives us an ASCII text file. That file, before adding true markers via macros, looks like this; note the text I use as a placeholder for where the macro should place a real page marker. I see that the mail merge did not follow through on the promise of skipping blank entries. I also see that, despite showing me centered text on screen in my template, that was stripped from the document as well. I have quite a bit of cleanup to do to tighten up these layouts. This is becoming work. Next, I'll run the macro I created to automate marker insertion. Macros are action-for-action transcripts of the keyboard actions required to accomplish a task, including cursor repositioning or text cleanup to prepare for the next step. Simple macros can be built simply, because macros are defined in the language of the AppleWorks user . I cannot praise this approach to macro scripting enough, and I will bring it up every opportunity I can. The more I encounter it, the more I honestly feel something fundamental has been taken from users over time. It's one thing to allow someone to record their actions blindly, with a slightly patronizing "There, there, don't worry your sweet little head about how this all works." It's quite another thing to tell a user, "Hey, you know all those tools you've been learning? You can use those exact same skills in powerful new ways." It would be relatively simple then to teach someone how to wrap an existing macro in a loop or decision construct ( UltraMacros can do these things and more), building up a foundation of self-confidence. It must surely be far more accessible than whatever Google is proposing. I mean, speaking of Cthulhu! After running my macro to the end of the document (holding down the macro shortcut auto-repeated until finished), here's the final result. I have to to "zoom" into the document and reveal the hidden formatting codes, but so far so good. Everything appears to be ready for StoryWorks . Fingers crossed. Hot dang, got it on the first try! (as far as you know) Now that the basic structure seems to be in working order, the last thing to do is write an entire novel. *dry cough * What a delightful piece of software; a delicious turducken! It will make excellent sandwiches for tomorrow's lunch. Easy to learn, easy to use, and even relatively complex cross-application actions are achievable with a gentle learning curve. I found I only needed manuals and books for very rare, "Can AppleWorks even do this?" questions, and very rarely for, "I'm stumped." Of the word processors I've covered to date it's easily my favorite. Typing is responsive and editing is intuitive. I like the advanced tools, like markers, though I think other tools could be made easier to use or expanded. Keyboard shortcuts became second-nature very rapidly. Tasks I usually dread, like mail merge, are trivially accomplished. The database " is ." It exists. It's fine. Perhaps its simplicity is a virtue, to some degree? I'm not personally smitten, but I'm glad to have it, and it did solve a real problem for me with the CYOA construction. I enjoyed the spreadsheet as much as I did VisiCalc , which is to say it's very good but Lotus 1-2-3 still wins the day. Still, it's a lot of bang for your buck, considering it's only 1/3 of this package. Being able to easily share data with the other modules elevates its usefulness over VisiCalc . Its convenience gives it a clear win. Plus, it didn't include "Copilot integration" long before Microsoft removed it from Excel . Perhaps the biggest takeaway for me is the stark reminder of how much can be achieved with so little: so little CPU, so little RAM, so little hard disk space. One can't help but ponder, "If this is what 128KB can do, imagine applying the same discipline toward 1MB of RAM." Writing for inCider , Oct. 1989, Senior Editor Paul Statt compared the development cycle of Lotus 1-2-3 Release 3 (a rather notorious, oft-delayed "update") with AppleWorks 3 . Both were initially created by lone developers, Jonathan Sachs and Rupert Lissner. Release 3 of Lotus took a team of 40 developers years to write and shipped on 14 floppy discs. AppleWorks 3 was written by three Beagle Bros developers and shipped on two double-sided discs. An update to Lotus likely meant also updating one's computer to a 286 with 1MB of RAM and a hard drive. AppleWorks 3 ran on the exact same 128KB machine the version 1 ran on. Almost 40 years later, his complaint feels uncomfortably modern, don't you think? In recent news we read of ways to whittle Microsoft's 1GB weather app "down to" 130MB RAM consumption. While on the other hand, we have something like Weatherbot , occupying 6MB RAM on Windows 11 and about 2MB on System 7.5.5 (yes, it runs natively on both and more!) The call for an efficient use of system resources is clearly still championed by some, but I don't see much of a call for brutal efficiency. Using 1/100 the RAM of a Microsoft-made app is fantastic, don't get me wrong, but let's get far more ambitious. No more thinking in megabytes, think in kilobytes ; express ambition by orders of magnitude . AppleWorks singlehandedly kept Apple IIs in production use for decades , well into the GUI era, well past their "prime." We deserve such longevity again. We need it. There is a financial cudgel of planned obsolescence that beats us down for lunch money on a regular basis. We're told it's for progress, but that proposition holds no tether to reality when we can see and touch a 128KB rebuttal that proves the bully a liar. Don't worry, I wouldn't leave you hanging. The beginning of my James Epoxy CYOA; pause to read. The description of Dimitrius is straight from the source book; I did not set that up for my running gag! Ways to improve the experience, notable deficiencies, workarounds, and notes about incorporating the software into modern workflows (if possible). As with many tools of this type and era, the program's emphasis on "printing" limits our formatting options pretty drastically. We can only embed non-printing printer codes into a document, and that requires setting toggles for (say) bold to start and stop at specific points in the page. It's anachronistic in all of the non-nostalgic, annoying ways. Late in my review cycle I came across a project keeping ProDOS alive on real Apple II hardware. The last official version of Apple ProDOS was 2.0.3 in 1993, but this project is at 2.4.3 with 2.5 on the way. John Brooks has been maintaining this for years now. It includes a kind of app fast launcher called Bitsy Bye which might smooth the process of switching between apps (like between AppleWorks and StoryWorks , in my case), if for no other reason than it appears to eliminate keystrokes and simplify file navigation. 1984 rocked. AppleWin x64 1.32.00 on Windows 11 Emulating a prohibitively expensive Enhanced Apple IIe. All cards and hard drive included, maybe $8,000? ($22K in 2026) 3x machine speed and enhanced disk access Super Serial Card Mockingboard C Disk II w/two floppy drives Hard Disk Controller with 5MB ProDrive Z80 SoftCard RamWorks III (3MB) AppleWorks 3 TimeOut UltraMacros In the database, will "zoom in" and "zoom out" between the record list and an individual record. In the spreadsheet, it will toggle display of raw values or formulas embedded in the cells. In the word processor, it will reveal hidden formatting, like for printing, navigation tags, bold/italic format codes, carriage returns, and so on. Graphing and plotting spreadsheet data, for the Lotus -envious Expanding the limits of open documents, clipboard, database size A full clone of MacPaint Appointment calendar High-resolution font support The new version of AppleWin x64 makes it super simple to install a whole host of fun circuit boards into various slots. You can build the pimped out Apple IIe of your dreams effortlessly. I mostly ran at 300% CPU speed and experienced no quirks, repeating keys, crashes, or anything else unbecoming of a well-behaved computer. AppleWorks never crashed; the recent update to AppleWin x64 worked perfectly. I did have AppleWorks fail to bring in some fields when I copied from the database into the word processor. CiderPress2 works perfectly for opening Apple II disk images. Hard drive images are of file extension, and CiderPress2 can manage these, including copying file structures between disk images to cobble together your own custom disk from others. CiderPress2 can natively open AppleWorks documents to show formatted content. Personally, I found it better to print from AppleWorks to an ASCII file, then copy out the text from that via CiderPress2 . Doing so will ensure no hidden AppleWorks formatting codes are copied over. Modality The modal nature of the editing tools limits the power of macros. If we could select some text, then apply a macro to that selection, that would be fantastic. Markdown keyboard shortcuts would be a breeze! StoryWorks The biggest issue I have is how there is no option for exporting standalone stories. Being able to build a bootable Apple II floppy would turn it into a great, simple game maker; translations of existing CYOA adventures would almost be self-coding. I would also like for it to respect formatting codes, like centering. Spreadsheet Though opening DIF files is an option, I had zero luck getting it to open my VisiCalc DIF exports. It doesn't match the integration goals of the program, but I did miss the slash menu; it could have been a nice alternate UI for those transitioning to AppleWorks . Editing cells, such as to change a cell from a label to a value, always tripped me up. Considering this is version 3, well after Lotus 1-2-3 hit the scene, I do wish for better control over column widths and some concession to graph making. 3D spreadsheets, linking a cell in one sheet to an entirely different file, would be useful; this feature debuted in AppleWorks 5 . Database Basic concerns about being unable to apply value types to fields, or any kind of data input validation, are addressed in AppleWorks 4 and 5. Form layout formatting tools are overly simplistic. Being able to merge into the word processor directly from disk, rather than from the clipboard, would be helpful. Word Processor I don't have much to complain about. It is more robust than it first appears, with a gentle learning curve. I think word count should be surfaced to the main UI, and the on-screen ruler could be more informative.

0 views

How I think about reducing AI costs

AI inference costs are something I've been writing about for a while . It's clear that for many companies this is becoming a huge problem: AI spend per employee per month. Source: Ramp AI Index via a16z, 12 August 2026. I've heard from a lot of readers that reducing this cost is becoming a hot topic internally. So, here's how I think about this problem at a high level. You need to have a good handle on what is driving your bill. I've met a lot of companies who have a pretty fragmented understanding of the costs - they can be siloed over many different teams, business units and roles. At a minimum, you need to know - business-wide - how much you are spending on AI and importantly, key stats on how that breaks down. There are two main dimensions on this. Firstly, the model in use. I see a lot of people running ancient, poor value for money models. The tech debt is real here! For example, GPT-4o is $2.50/$10 per million tokens, but is drastically worse than GPT-5.6 Luna, which is 10% of the cost. It's really important to know what models you are using. GPT-5.6 Luna scores over 4x higher than GPT-4o on the Artificial Analysis intelligence index, and costs a tenth as much. The second dimension is how your spend breaks down by the three main components of token costs - cached input, uncached input and output. As I wrote recently , with agents the distribution of these costs is rarely what you'd expect. Keep in mind you need to collect this data for all your usage. This includes LLM usage via API or similar, coding agents and any other autonomous/business agents you have running. A key mistake I see is people focussing on their API spend, but not looking at their enormous coding agent spend for their dev team, which is out of control, or vice versa. Once you've got this done, the next thing is to look for obvious cost savings. As I mentioned before, it's usually quite easy to swap out legacy models with something cheaper from the same provider - though it should involve a verification stage, because you can risk regressions this way. It can result in a lot of the 'hacks' you've perhaps used for a less intelligent model backfiring. The other key thing to look for is models that are 'overpowered' for the use case you are working on. I often see teams using the 'largest' models for use cases where a much smaller and cheaper model would do. This is not as intuitive as it looks, and takes quite a lot of experience to realise what can and can't be switched out intelligence-wise. But certainly if you have large bills coming from certain workflows, it's definitely worth experimenting with them. The more "drastic" option is to switch away from OpenAI/Anthropic/Google models to a different provider that can host open weights models for you. Whether this is worth it really depends on your spend. If it's a fairly minimal level of spend, it may not be worth the procurement and data privacy reviews your company may have. But for most, this is often where the meat of the savings comes from. It's also important to say you don't need to move everything off at once. I've seen some token-hungry workflows that account for a huge proportion of spend - these can be moved off while you keep everything else with the frontier labs, reducing the amount of upfront work dramatically while maximising savings. There are a lot of companies based in the US (and Europe!) offering hosted open weights models at attractive prices, and they can offer high-quality SLAs and meet data residency requirements. Expect to have to break the myth internally that DeepSeek (for example) is "based in China". While their own API is, many providers offer that same model in your jurisdiction. While the models move too fast to keep the article up to date, there's more on my token cost optimization page if you want to get in touch and get some thoughts on which open weights models and providers may be best for your use case. The final lever for optimising your spend is going much deeper into what each of your workflows is actually doing. This is a more involved process and I'd recommend starting with your ten highest cost workflows and seeing if you can identify issues. While this can be extremely nuanced, I'll list some common failure modes below so you can see if these map to yours. Firstly, a problem I see a lot is putting far too much into the prompt. For example, including pages and pages of documents that may or may not be relevant to the user's prompt. With tool use these days it's often far more token efficient to give the LLM a tool it can use to search for relevant documents, rather than putting hundreds of pages of documents in the prompt just in case. Ironically, the second issue I see over and over again is incredibly poorly optimised tools themselves . These can be either internal or third-party, but often MCPs return tens of thousands of characters of JSON back to the agent. This absolutely burns through tokens and a few small tweaks to your tool definitions can really help. Intuit's official QuickBooks Online MCP server is a good example, and it's worth saying it isn't sloppy work - the code is careful, the tests are real, and the security thinking is better than most repos I read. It's a token disaster anyway. It ships 142 tools. Serialised, that's roughly 21,000 tokens of tool definitions going up on every request, before the user has typed anything at all. Search results come back as raw dumps. hands back every field of every invoice it finds, straight from the QuickBooks API, with no filtering and no summarising. Ask a broad question and you can pull an entire ledger into context verbatim. And returns the PDF base64-encoded as text. A 100KB invoice becomes something like 33,000 tokens of noise that the model can't read and can't compress. That one call could cost you more than the rest of the conversation put together. [1] The other category of issues is around tool failures which then drive up agentic token spend as the agent has to try and retry them or work around them. This requires good monitoring. This can be incredibly expensive for longer running agents, especially tool call failures towards the end of a run (where cache read costs start exploding). There are many other weird and wonderful ways teams manage to use LLMs inefficiently, but hopefully this gives you a good overview of the main ones I see. The final issue is that as AI is developing so quickly, you have to keep up with best practice. Best practice from even a year ago can often be actively harmful now. And with the plethora of models coming out, you need to stay on top of model trends as well. I'd recommend teams schedule at least a quarterly review of their token spend and how the model/agent landscape has changed and what mitigations can be applied. Finally, I do have some slots for companies that want to bring my experience in, and I help show them how to do this. The recent results I've got are extremely positive, reducing one company's token bill by over 50% in three weeks. If you'd like to reach out , I'd be happy to do a deep dive into your token spend. This is the part that bothers me. Almost nobody shipping tools to end users treats token cost as a design constraint. Vendors optimise for feature coverage and demo appeal, because that's what gets an integration listed and adopted - and the token bill lands on the customer, not the vendor. There's no line item anywhere that punishes them for it. Until buyers start asking how many tokens an integration burns to do its job, I don't expect this to improve. ↩︎ This is the part that bothers me. Almost nobody shipping tools to end users treats token cost as a design constraint. Vendors optimise for feature coverage and demo appeal, because that's what gets an integration listed and adopted - and the token bill lands on the customer, not the vendor. There's no line item anywhere that punishes them for it. Until buyers start asking how many tokens an integration burns to do its job, I don't expect this to improve. ↩︎

0 views
Phil Eaton 2 days ago

The road to ACID transactions in Cassandra 6

This is an external post of mine. Click here if you are not redirected.

0 views