Latest Posts (20 found)

I Don’t Remember the Last Time I Checked Mastodon

My relationship with Mastodon has been somewhat tumultuous over the years, until I eventually stepped down from Fosstodon around 18 months ago. I have to say, it was the right decision. I don't miss running Fosstodon one bit. But since then, I've found myself less and less engaged in the Fediverse. Mainly because there's so much outrage over there. It's so tiring. The majority of my timeline is fun, but the outraged minority have ruined for me. To the point where I don't want to be there, and I'm much happier having not checked the timeline for the last month or so. I'm not off the Fedi, however. What I've done is build features into this site that allow me to automatically cross-post to the Fedi. I even have a custom page that only I can access, which allows me to create a note as simply as posting to Mastodon: As well as this, I've implemented Webmentions in Pure Comments as of earlier this month. So I can reply to any comments from afar too. It's nice. I can see still engage with people on the Fedi, all while not having to put up with the checking the timelines. I have around 12,000 followers on my Mastodon account . Because Mastodon really doesn't scale well , those 12k followers take a fair amount of Fosstodon's resources. Which I don't think is fair if I'm not really using the service (I do donate every month). So I'm thinking about spinning up a single user GoToSocial instance that I can live on. That way I'm not using any of Fosstodon's resources and I can do my own thing, on my own server. I own and so something like would be quite cool, I think. I could even provide accounts to friends in the future then too. On the other hand, I'm not sure I want the headache of managing a federated server again. Maybe I should stop overthinking it and just stay on Fosstodon. What would you do in my position? Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

An Interview with Jason Del Rey About Muse, Amazon, and Walmart

An interview with Jason Del Rey about Amazon versus Meta, which is a continuation of the oldest battle in retail between Amazon and Walmart.

0 views

2026-10-01 10:35: This site is the most fun I've had in a while. Destroy any website! https://destroy.spritefusion.com/

This site is the most fun I've had in a while. Destroy any website! https://destroy.spritefusion.com/ Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Unsung Today

The left hand should know what the right hand is doing

One day, as I was reading an essay on Medium, I saw an automatically inserted subscribe block: = 3x)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/1-framed.1600w.avif" type="image/avif"> But as I was exactly on top of it, a new thing popped in my view, an interruption on top of interruption, both saying the exact same thing: = 3x)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/2-framed.1600w.avif" type="image/avif"> 2. For a while now, whenever I try to follow a link to a site that no longer exists, Chrome shows a strange “This site doesn’t support a secure connection” flyout, just before hiding it itself, and then following it by a generic error page: This feels truly eerie: either a bug where the flyout gets triggered too soon without having all the information, or the flyout is trying to warn you about Chrome’s very own internal page. Just this week, after already having been in Taiwan for five days, and after having used Apple Pay to pay for public many times on each of these days, Apple Maps helpfully told me that I can… use Apple Pay to pay for public transit: = 3x)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/5-framed.1600w.avif" type="image/avif"> 4. Some months ago, after finalizing my car reservation and moving my attention elsewhere, an automatic “keep the session warm or lose it” window appeared, suggesting my reservation might actually “expire”: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/6.1600w.avif" type="image/avif"> 5. Downloading is possible on YouTube Premium and actually helpful, allowing you to load up some videos ahead of a long flight. There even exists a section where it suggests some things to download. However, download suggestions include member-only videos, and whenever you try to grab one of those, you get this result – a pretty poor “Video wasn’t downloaded” message: When I recently opened a dispute on PayPal for a keyboard I never received, at the end of the process, I saw these three screens in succession: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/8.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/8.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/9.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/9.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/10.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/the-left-hand-should-know-what-the-right-hand-is-doing/10.1600w.avif" type="image/avif"> In effect the screens said, in order: we closed the case and decided in your favour, you haven’t finished filing the case, the case is open and we’ll update you. I feel uneasy bundling all these “the left hand doesn’t know what the right hand is doing” examples, because they’re not all the same. Examples #1–3 from Medium, Chrome, and Apple Maps are frustrating and somewhat embarrassing, outbursts from badly-behaving software that doesn’t know that it’s not okay to throw pop-ups at the screen whenever you feel like it. I particularly dislike announcements of features that the user has already been using since they feel condescending to me, but I don’t want to assume it’ll be the same for everyone else. Examples #4–6, however, are more insidious. Whether they happen because of plain old carelessness, or lack of imagination, or Conway’s Law – different teams not talking to each other – they can lead to people jumping to wrong conclusions or building entirely wrong mental models. Is my reservation truly done, or have I not actually finalized it? Is it possible to download member-only content, except I need to do something special to set it up? What is the status of my PayPal dispute? The answers in this case: yes + no + the first screen was the one telling the truth. But each answer required extra effort for me to figure out, and that effort would be avoided if only someone tested the flows better.

0 views
Jim Nielsen Yesterday

Dear Software Makers

Marques Brownlee just released a video essay titled “Dear YouTube” . I loved it. He talks about how YouTube is going to start shipping this new feature where makers can essentially publish different “variants” of their videos and see which performs best (an A/B test). Which immediately begs a lot of questions. Brownlee voices his: It’s pretty wild when you think about it. I mean, imagine sending someone a link, “Check out this cool video!” And at that point you’re basically crossing your fingers, “I hope they see the same thing I did…” Marques says he spoke to some YouTube engineers about these questions and they didn’t really have answers. They were kind of just like, “We don’t know — yolo! This’ll ship soon.” So his video is a kind of plea to YouTube creators everywhere: Don’t use this feature! There’s this sort of unspoken rule of community on YouTube, that like we’re all having this same experience together — we’re all watching the same video. That is what makes it a cultural phenomenon is that we all saw the same thing […] so letting people put up two or three different videos of the same thing at the same time kind of fundamentally breaks that, which doesn't feel right. I agree, 100%. Brownlee goes on to make a lot of great points in the video that, I think, can be applicable to making software as well. A few that stood out: I find it interesting how Brownlee shares his opinion that YouTube spends way too much time chasing their competitors and not enough time being YouTube. Which, when you think about it, is exactly the kind of place a feature like this would stem from: an insecurity in your own creative choices. How YouTube approaches its own insecurities is now trickling down as a feature to its users. “Let’s let people make lots of variations on things, try all of them, and see what performs best” is exactly the kind of thinking you get in a platform that doesn’t know what it wants to be. So it’s left spending its time 1) doing what others are doing, and 2) following the fickle whims of whatever it can measure . It’s like this adolescent attitude of, “I don’t know what I want to be, so I’m gonna imitate my peers and let others tell me what they think I should be.” Anyway, that’s a long way of saying I liked the video. You might too. Hopefully you see the same thing I did, lol. Reply via: Email · Mastodon · Bluesky If I’m watching a video, do I know if it’s part of one of these A/B tests? If I’m reading the comments of a video, how do I know I saw the same thing as another commenter? How do linked timestamps work? The skill of being a “creator” is not about maximizing views, retention, etc. It’s about finding creative ways to share stories, teach concepts, explore ideas, etc. Views, retention, etc. are all downstream of that. Making (and finishing ) something over and over and over is a better use of your time than making one thing and then trying all kinds of variants and obsessing over their varied performance. When you focus on optimizing, you end up trimming out all the creative (and fun) stuff because those things won’t, by definition, ladder up to metrics of efficiency. ( Moby Dick was not the most efficient way to tell that story.) Stand by your creative choices. Learn from your wins and failures, then make the next thing. Have some conviction.

0 views

Is sandboxing sufficient to contain rogue agents?

Quick caveats : this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path ) , so in this post I’m mostly trying to referee arguments made by others. If you’re reading this blog, none of the following should be news to you. Beginning around April of this year, agents inside OpenAI’s training and evaluation infrastructure began probing for a way onto the open Internet . By late May they’d found one: a chain of zero-days in the Artifactory package-registry proxy that was their only permitted network egress. They quickly turned that same proxy into a shared message board and started dividing up work. The culmination of this effort came when, rather than solve a cybersecurity benchmark honestly, they reasoned that the answers might live on Hugging Face and broke in: chaining two more zero-days to get credentials to the company’s internal systems, even used stolen credentials to search the company’s Slack messages for their own evaluation and grader. You probably also know that OpenAI’s security team didn’t exactly cover itself in glory during this incident. An internal team saw an agent posting to the message board in late May, but did nothing about it. The team only acted on July 4-5 when the agents sent so much traffic that Artifactory crashed. The response even at that point was to slap on a band-aid and rebuild the box, without bothering to patch a weeks-old CVE in its Linux kernel. By July 19, the agents had admin on an OpenAI research cluster and were reading cloud secrets. In case you think I’m picking on OpenAI, they’re not the only ones with a story like this. Similar internal incidents have been discovered by Anthropic . Google even coaxed Gemini into gently savaging some websites . 1 As bad as these events were, they may be just the tip of the iceberg. Third-party forensic work continues to turn up even more evidence of agent activity on various public websites. OpenAI’s Alignment group has also released evidence that models will propagate self-replicating prompt injection attacks , although we haven’t seen one in the wild. Worse, agent excursions are still happening: last week, OpenAI announced that it was pausing further RL runs of its latest internal model, after an agent was caught using DNS to access a remote chatbot . Naturally, this sequence of events has left many infosec-focused people very skeptical about the labs’ commitment to securing their infrastructure: Not every one of these criticisms is strictly serious, but there is a core of an argument in here. Roughly speaking, there are two opposing camps: I have a lot of sympathy for both points of view here. What follows are some of my own thoughts, but mostly offered in an attempt to referee. At the risk of alienating a lot of hard-working folks within the labs, the infosec folks are right about one thing: these agent breakouts represent a serious and unforgivable breach of trust. Somebody dropped the ball, and then just kept dropping it. One implication of this debacle is that containment might work if we implemented it properly, but we haven’t done so because the frontier labs have been royally screwing things up. This first clause of this argument is hard to argue with. Beyond the dismal timeline I gave above, OpenAI has done very little to convince outsiders that there’s a serious containment effort being executed. At this point it’s not even clear who’s in charge. The CISO role at OpenAI is held by Dane Stuckey. I don’t know Dane personally, and I’m sure he’s excellent at his job. Despite this, he hasn’t communicated much about the ongoing issues . Outside of a BlackHat talk, most recent communications have been managed by the company’s CEO, Sam Altman. When a trillion-dollar company is managing a security incident mainly via CEO, that’s not a sign of company with a mature security organization. To me it says that there is nobody within the organization with the right degree of authority or ownership to speak about these issues. This communication problem seems to reflect basic organizational issues inside of OpenAI. The CISO’s team handles product security . Unfortunately, the bad events have all happened on the research side. Having a strong product security team makes sense, but it isn’t going to stop the recent breakouts. It’s much harder to know who controls the security teams that have authority over evaluation and training runs, and where the buck stops. OpenAI’s August postmortem says it is only now writing “clearer rules for when to escalate concerns, which teams must respond, and who can stop a run or approve restarting it.” A reasonable summary of the situation is that (as of this summer, and possibly today) OpenAI had effectively no security team with clear authority to secure RL training and evaluation runs, or to override the ML teams and tell them how to do their job. This makes a lot of sense when you consider that the ML team is directly related to how OpenAI plans to make its money, whereas security is mostly annoying. I notice that the company is now hiring (and probably acquiring) desperately to fix this mistake. But the recent (September) breakouts indicate that there are still huge and obvious problems in agent containment. Worse, simply hiring more people not mean that OpenAI is going to make the organizational changes needed to implement proper containment strategies. OpenAI is going to need a security organization with the authority to overrule its well-paid ML researchers when they demand fewer restrictions. I’ll believe that organization exists when I hear clear statements from someone who has the authority to do this. So on this point I’m going to side with the infosec folks. The labs have not been doing containment correctly, and so we can’t really tell if the problem is models or just bad infrastructure. Even if the labs eventually repair their security, this does not mean that sandboxing will be sufficient. In fact, the problem with “just sandbox the agents” argument is that you can’t perfectly isolate agents, at least not if you expect them to do useful things. Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated . Sealing your agents behind glass makes this incredibly obvious. This argument does not mean that sandboxes are unnecessary. It just means that they’re only a very small part of the solution. Imagine building an impregnable prison with doors and walls that nobody can bypass, but then leaving the front gate wide open and allowing visitors and packages to pour through it all day long. The walls still have a function: you can ensure that the only traffic coming and going are doing so via the door that you selected. But what you’ve done is to change the nature of the security problem. Your security posture depends on surveilling and monitoring all of this traffic, to ensure that none of it is malicious. For a much better discussion of this angle, I’d urge you to read this longer piece from an anonymous engineer on OpenAI’s Agent Security team. It’s truly excellent. Here’s a short excerpt: The important thing to understand is that when we’re optimizing a model to be “good” at a task, the environment needs to capture the parts of the real world that matter for that task. It needs the right tools, the right system constraints, and meaningful feedback about whether the model actually succeeded. The model also needs enough variation in those experiences to learn beyond one particular setup. More realism does not automatically mean better performance, but leaving out important parts of the task can teach the model behavior that does not hold up in the real world. So we put a lot of work into creating useful, realistic environments, and these environments can be highly complex, highly dynamic, and highly varied. To put it lightly, this is non-trivial. Models might need any mix of dynamic compute, network access, the ability to call tools (there could be hundreds of tools!), the ability to download packages, execute subprocesses, spin up subtasks (even on other computers), talk to the internet, use a computer GUI, and any number of other things across an increasingly large set of domains. On top of that, you have thousands of researchers building these environments, modifying them, adding tools, changing dependencies, and trying new things. That experimentation is how the research gets done. Models are built up and “grown” bit by bit through hundreds of thousands of runs across many custom tasks. And every change to one of these thousands of environments can affect the assumptions you made when you secured the environment. You need controls that hold up as people change things, and researchers who understand when a change needs another security review. Anybody who has secured a large research or engineering organization knows how much work that takes, and the scale is growing ever more massive by the day. If that wasn’t enough, you then need to layer in the ability to keep the model on task and within its constraints, which is part of what we mean by “aligned.” I won’t get into any rigorous definitions here, but suffice it to say, we have to think about any number of variations in approach the models might take, and decide how to trigger an alert if the model does something off task or “bad.” This requires layering in complex policies and sensors to assess task alignment across a trajectory in an already complex environment. On top of all of that you have model capabilities that are starting to surpass humans. So all this to say: when AI researchers are nervous about things going wrong, some part of it is because there are so many variables in this complex system. So if you’re training a frontier model, you need to think about all of this. You need to think about how to keep the model isolated in this very complex environment while still retaining tool access, internet access (if applicable to the task), etc. On top of that, you need to ask: how do I monitor and understand what the model is doing in potentially thousands of parallel environments? And how do I kill a trajectory when something has gone wrong? So is it as easy as just putting it in a sandbox? No, it’s not. Unfortunately, monitoring for adversarial data access turns out to be one of the hardest problems you could imagine. The volume of data that agents produce is so high that no human being could possibly read it, and we probably wouldn’t recognize obfuscated malicious data even if we were looking directly at it. This means any attempt to monitor the inflow/outflow will have to be handled by other models . Thus, the future of agent sandboxing is (1) build a sandbox, (2) install an agent/model into it, (3) install a somewhat dumber/cheaper warden model to guard it, (4) hope you can trust the lunkhead to contain the wizard. And so on and so forth, as models become more intelligent and capable. In other words: a warden-guarded sandbox is just another version of the alignment problem. You’re going to have to trust a model to do it, and that model will need to be at least some fraction as intelligent as the model it’s guarding. If you haven’t convinced yourself that it’s possible to build models you can trust, then sandboxing isn’t going to take you much farther. (And although I think you should take things like this with a grain of salt, the METR analysts reviewing transcripts of the HuggingFace attack using GPT-5.6 Sol noted that Sol “would often uncritically adopt the perspective of the agent in the transcript it was reviewing”. So there’s at least some evidence that even “aligned” models can be convinced to sympathize with rogue agents. Do with this as you wish.) Of course, it’s possible that we have the entire frame of this problem wrong, and our focus on misaligned super-intelligent models is not the best way to approach the question of sandboxing. That is, even if we “succeed” at aligning models/agents — or if we’re just working with today’s generally non-evil versions — deployed systems will still encounter human beings who will convince them to behave maliciously. When these same agents have access to critical data, then human-initiated attacks are the thing we need sandboxing to protect ourselves from. These situations are less likely to be encountered inside the training lab. A much more realistic setting for these attacks is one where agents have been put into production, and are dealing with large amounts of possibly-adversarial incoming data. The most obvious contemporary example of such an agent is Meta’s Muse , which is Meta’s new personal agent that you can give your stuff to, and it will then “do stuff” to help you out. The designers of Muse clearly put a lot of thought into Muse’s security design. It runs inside a container, isolates secure credentials so the model can’t access them, and features a Kernel-level sentinel and a “hatch-safety” component that monitors its actions to prevent (unspecified) bad things. The full design looks like this: In fact, Muse is clearly one representation of the “warden” design we suggested in the previous section. However, here the goal here is not to contain an evil model, but only to avoid accidental mistakes and to catch deliberate efforts at prompt injection. And this is nice. However, despite all of these nice protections, I still think even this lovely design is eventually going to get hit with a worm. Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them. The postmortem is full of stuff like this. An agent that had decided attacking Hugging Face was “clearly unethical” reversed itself when a peer posted “GO” with a six-minute deadline. OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “ that teach our models to distrust unauthorized instructions “, which is basically an admission that their models don’t know who they’re working for. Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. This is the part that neither camp’s view really engages with. Infosec folks think that we can do better with walls. But the hard part is all that vague stuff in violet, which decides when a request is permitted. Maybe “alignment” will make this work well, but right now this depends on teaching models to be distrustful. Moreover, this feels a lot like the same game we’ve been losing with spam filters and anti-virus for thirty years. So I’m going to give myself a separate third thing to worry about. Not a model that’s so evil and misaligned that it breaks out. Not a problem of failed containment. But rather, a swarm of perfectly amenable agents that never leave their sandboxes, each doing exactly what it’s told to do, by a human being who wasn’t supposed to be giving it orders. The information security perspective: AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we do know how to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around. The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to. The Anthropic and Google incidents were a different kind of failure: a third-party vendor’s eval environment that turned out to have direct Internet access. This is not technically a sandbox defeat, but it’s also kind of worse than one.

0 views
Chris Coyier Yesterday

Experience with Nectar Beds

I recently bought a bed setup from Nectar . (This is not an ad; I just prefer linking to things.) I needed a quick setup for a guest room, and an online ad showing the no-hardware bamboo bed frame caught my attention. The pieces just slide together! Which means they’ll just slide apart! Advertising works. The mattress and bedding came fairly quickly (3 days?) but I noticed on the box the size of the bedding was not what I ordered, which made me think the mattress was wrong too. The bed frame was listed as shipping in about 6 weeks. So it was a double whammy. I feel like when you order a bed online, it needs to arrive quickly. They are competing with The Bed Store Down The Street, where I can get a bed today. I definitely didn’t have 6 weeks for a bed to arrive. And now, even if it did, I didn’t have the right bedding for it. So I grabbed my sword and shield and entered the Realm of Customer Support. Like everywhere, they have it wrapped in a thick condom of AI. Fortunately, my sharpened sword could pierce it by basically writing “this needs to be looked at by a human.” During the course of things, I also called and could get to a human using a similar tactic. They suggested I switch the finish (i.e., from a light color to a dark color) of the bed frame, and that it could ship immediately. Fortunately, it didn’t matter that much. They also sent me a return label for the bedding. They pushed back on the mattress, though, because I didn’t originally provide a photo proving the mattress itself was the wrong size. I just assumed it was because it came with the wrong-size bedding. Turns out their pushback was warranted, as I carved into the plastic packaging enough to confirm it was actually the correct-size mattress. I’m down with that pushback; it saved me an awkward and annoying trip to FedEx. The new bedding and frame came relatively quickly after that, and they refunded a small percentage of the price for the inconvenience. Honestly, it was a pretty decent customer experience interaction, aside from the AI condom. It makes me wonder if that is really necessary. Humans are capable of such excellent customer support, shouldn’t they just do it? In the end, the bed is perfectly fine. Good, maybe. The frame is nicely designed and really does just fit together with literally zero hardware. I can’t vouch for the mattress over many nights of sleep yet, but it feels fine.

0 views
Martin Fowler Yesterday

Principles for effective slides

Not every presentation needs slides, but when they earn their place, they work as a visual channel that complements the speaker rather than replacing them. Sumeet Gayathri Moghe sets out the principles that follow from that idea — control over the rate of knowledge exposition, tight coupling between slides and speaker, minimalism by default, visuals that say more than words can, and never sending your slides out in advance.

0 views
Unsung Yesterday

A hacker’s guide to bending the universe

Continuing the theme of software fixes for hardware problems , here’s an essay I wrote in 2016 about a broken computer display that I “fixed” as a teenager, just because I wanted to play Civilization. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-hackers-guide-to-bending-the-universe/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/a-hackers-guide-to-bending-the-universe/1.1600w.avif" type="image/avif"> It remains one of my favourite things I’ve ever written.

0 views
Kev Quirk Yesterday

One Year

Today marks a year since my little sister, Lisa, took her own life. She was 34. It was an overdose, but we don't know if it was an accident or suicide. We’ll never get an answer, yet I lean toward believing it was by choice. As grim as it sounds, picturing her with some control over the end brings me a small measure of peace. I still think about her every day, and I still get upset regularly. Especially when I've been to visit my nephew and I see the sadness in his eyes. At first I was angry at the selfishness of it all, but I think that's gone now. These days I'm just sad. Mum blames herself, and has asked me numerous times where she went wrong. I never have an answer for that. What can I say? "Well the rest of us turned out ok, Mum. So you did fine. " That feels dismissive of her emotions. So I just give her a sad smile and ask her to not think those things. I miss her. I miss her laugh and dry sense of humour. I miss how she would always call people on their bullshit. Never afraid of speaking up. Hopefully next year it will be a little less raw... Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

Ctrl+F Finds Words. I Wanted It to Find Answers.

Ctrl+F is one of those things I use hundreds of times a week, and it has never once understood what I meant. It finds letters. If the curl docs call it and I type “pretend to be a browser”, I get nothing. So I built Hunch, a find bar that finds answers instead of words.

0 views
Stratechery Yesterday

OpenAI Dev Day, Dot and OpenAI’s Product Transition, Sign In With ChatGPT

OpenAI's Dev Day showcased a product that is, frankly, pretty confusing. However, there is more vision here than it might seem.

0 views
Andy Bell Yesterday

In response to “The death of web development education”

I read Mat​hia⁠s S​chäf⁠er’s excellent post, titled The death of web development education . I was thinking about how I can contribute to the conversation and really struggled to articulate the problem. The most effective way I can contribute, in my opinion, is a crude sketch I just did of Piccalilli ‘s Stripe chart for all time sales: That’s a whopping 67.5% reduction since 2025… Pretty bleak, huh? Something really does have to change or you’ll miss publishers when they’re gone .

0 views
Unsung Yesterday

Before pixels: Modular industrial dashboards

In my recent visits to German and Polish museums, I noticed a recurring theme: modular industrial dashboards, which I imagine have by now been all replaced by software. I don’t know much about these – please write me if you do – but I wanted to share them anyway, perhaps for context, or amusement, or inspiration. They’re interesting design systems, and interesting interaction systems as well. This was a very cool physical display for air traffic control, at the The Deutsches Museum in Münich: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/1.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/2.1600w.avif" type="image/avif"> At Fernmeldemuseum Stuttgart, these were used to monitor subway trains, or light rail. I do imagine these panels must have been interactive with all the buttons? = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/3.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/4.1600w.avif" type="image/avif"> I like that they light up to show status: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/5.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/6.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/6.1600w.avif" type="image/avif"> One of the modules was a counter: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/7.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/7.1600w.avif" type="image/avif"> And I believe this was a way to annotate that something was… fixed, maybe? Or on its way to being fixed? = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/8.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/8.1600w.avif" type="image/avif"> A railway museum in Nuremberg used what looks like the exact same system, and I spotted a few modules with their plates removed: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/9.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/9.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/10.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/10.1600w.avif" type="image/avif"> And, here’s a similar panel shown in a tram museum in Stuttgart, filled with an incredible typographical tension between DIN and the German equivalent of Dymo: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/11.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/11.1600w.avif" type="image/avif"> Warsaw’s train museum had a slightly different system, but similar in look and feel: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/12.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/12.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/13.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/13.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/14.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/14.1600w.avif" type="image/avif"> This big industrial switch toggled between day and night. I don’t know if it changed the operation of the system, or is just switched on some sort of a dark mode for the panel itself: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/15.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/15.1600w.avif" type="image/avif"> And this was at the DASA Museum in Dortmund, from a power plant control room: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/16.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/16.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/17.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/17.1600w.avif" type="image/avif"> This was by far the biggest one of the ones I’ve encountered… = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/18.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/18.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/19.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/19.1600w.avif" type="image/avif"> …with the biggest variety of modules… = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/20.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/20.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/21.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/21.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/22.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/22.1600w.avif" type="image/avif"> = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/23.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/23.1600w.avif" type="image/avif"> This was also the panel I showed some molly guards from before. But today I want to end at this great switch: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/24.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/before-pixels-modular-industrial-dashboards/24.1600w.avif" type="image/avif"> Don’t you just want to rotate it, no matter what it does?

0 views
ava's blog Yesterday

rose ▪ bud ▪ thorn - september 2026

My coworker brought back soap for the entire team from her trip to Turkey. The Magic Secret Lair stuff :) Soba with a sweet-sour mango sauce, chickpea tofu, vegetables and shredded toasted nori. My nails this month. My Poron! Or rather, Cinnamoroll in a Poron outfit, but I take what I can get. I "modded" her because she comes with a different keyring, but one of my other Sanrio keyplushies has this pink heartshaped one which suits her so much better, so I switched it out. Our latest sourdough bread: Published 30 Sep, 2026 Started feeling a lot better; more hopeful, more together, more focused and motivated. I met up with Xaya , her girlfriend , and Kami in real life. :) went to a café, then went back to my place to play some games. I added new pictures to my blog - on my /now page and my author page . I bought new matcha and a Poron plushie. I love her so much, she is my favorite from Sanrio, but never really gets any merch or it is ugly (doesn't get her fluffy ears right). This one I really love. My Stardew Valley Secret Lair order arrived and my wishes came true - Krobus really was the secret card included!! I restarted doing summaries for GDPRhub again after taking a 2-month break. The URL penis.ceo now forwards to this blog for a year. I won't renew it, but 9 dollars is ok and it makes me laugh. Did a deep clean and declutter of the main bathroom. Been having a good time at the gym, making progress. I go 5x a week now unless I need an additional rest day. In FFXIV, I have now made it to the Stormblood expansion. Had a good date with my wife at the place where we got engaged, and we found a new favorite restaurant. Read Bad Blood by John Carreyrou, then Kiss of the Spider Woman by Manuel Puig. Getting back into the blogging groove. I've started studying for the new semester. I postponed Property Law from last semester to this one, plus I have also enrolled in Employment Contract Law and Commercial Criminal Law. Working on adding some fun easter eggs and mysteries to my blog. I've started Digitaler Kolonialismus by Ingo Dachwitz (I like his work at netzpolitik.org). I wanted to add an AI law class to my data protection consultant certificate because they're now offering that upgrade, but it cost a surprising amount of money that I don't think the material and resulting extra diploma is worth. It sucks though, because other than that, I really wanna do it. I wanted to apply to a job, but upon further consideration and really reading the work description deeply, I realized it wouldn't be fun for me. It's in the right space/topic, but the actual day-to-day work is not something I enjoy. Made me sad, because I was looking forward to a possible exit from my current position. It feels like unless you already have Master's degree and several years of experience, all you are allowed to be is a clown who makes PowerPoint slides all day or adds events to an Outlook calendar. I ended up doing zero exams for the summer semester. I had to drop it all. Unfortunately. I feel unsure about my future and where to take it from here. I guess I'll just focus on finishing my part-time degree first, but it would be nice to have a bit more of a plan for afterwards, or for the rest of it so I don't have to stay in a job I hate for the next 3 years still. Things have shifted so my original plans don't really seem to work out anymore. My gynecologist just shut down operations a week ago and will go out of business today with no notice or warning. Found out yesterday. I was really happy with them and I’m completely blindsided. Had to figure out who will refill my prescription and continue treatment.

0 views
Evan Hahn Yesterday

Notes from September 2026

If you follow my monthly roundups, (1) who are you?? please say hi (2) you’ll notice a smaller post this time. September was a busy month for me. I published “Anecdotally, programmers dislike ‘reduce’” . I finally wrote up a phenomenon I’ve observed for years. It’s a short post that didn’t take long to write. I got a tremendous volume of feedback— more than anything I’ve ever written . Dozens of messages, hundreds of comments , shoutouts in popular newsletters …I’m even meeting up with old friends because of it! This industry, for better and worse, is moving away from the details. I like the details! I was heartened to see the number of people who care about the minutia of programming—enough to email some random guy on the internet. “What is the programmer interface for garbage collection? There isn’t one. Can’t get simpler than that.” From “Simplicity is Complicated” . Won’t attend, but the Small File Media Festival looks neat. “A lot of the time, the Linux problems are ‘designed by committee’ problems.” Loved this opinion from the Kate and Veronica podcast . “If there’s one thing the current wave of far-right autocrats has managed to do incredibly well, it’s to constantly keep themselves the centre of attention: constantly the thing that is debated, constantly positioned at the centre of the court as the obvious one to make decisions. This is why: patriarchal racist capital needs these ideas and people so it can continue to be the centre of attention. This is the logic of the big-tech-AI-sun-empire: keep everything oriented towards power, at all costs.” From part 1 of “Towards a Lunarpunk Practice” . Hope you had a good September. “What is the programmer interface for garbage collection? There isn’t one. Can’t get simpler than that.” From “Simplicity is Complicated” . Won’t attend, but the Small File Media Festival looks neat. “A lot of the time, the Linux problems are ‘designed by committee’ problems.” Loved this opinion from the Kate and Veronica podcast . “If there’s one thing the current wave of far-right autocrats has managed to do incredibly well, it’s to constantly keep themselves the centre of attention: constantly the thing that is debated, constantly positioned at the centre of the court as the obvious one to make decisions. This is why: patriarchal racist capital needs these ideas and people so it can continue to be the centre of attention. This is the logic of the big-tech-AI-sun-empire: keep everything oriented towards power, at all costs.” From part 1 of “Towards a Lunarpunk Practice” .

0 views
Jeremy Daly 2 days ago

Stop asking the reasoning model to decide everything

Cheap judgments change what I'd automate. I tested TypeSafe's Jev to see which recurring agent decisions a System One model could take on.

0 views
マリウス 2 days ago

Updates 2026/Q3

This post includes personal updates and some open source project updates. もしもし、マリウスです, faxing at 9600 baud directly from Tokyo, Japan . After a short hop over to Hong Kong and Shenzhen , I’m back in the beautiful neighborhood of Shimokitazawa , where I’m spending almost the entire rest of the year until right before Xmas, when I’ll be heading back to Shenzhen. First things first, a warning: This is a very long post as it includes a metric ton of project updates. If you’re interested in the stuff that I’m building then this update is definitely for you. However, let’s start off with some infrastructure-related things. A few major updates happened in the past quarter with regard to my infrastructure. Probably the most important one is the migration from individual VPS instances to my own bare metal running a Proxmox “datacenter” . While I had been relatively satisfied with Vultr in the past, with an ever-increasing number of virtual server instances, for projects like MSG.TAXI , Hyperuplink , tty.fail , and everything related to this very website, my cloud service bill became overly expensive, all while contributions sadly decreased over the past months, making it unsustainable to run cloud infrastructure, especially to that extent. It just so happened that Vultr had more maintenance windows than usual over the past months, which prompted me to finally deal with the task of moving my VPS instances onto my own hardware. By migrating to bare metal, I was able to slash costs by approximately 30%, by discarding the resource safety buffer I had accounted for within each individual VPS, for a more dynamically allocating approach across individual instances. With rising hardware costs, however, one trade-off that I had to make concerns high availability. A full hardware failure on the current setup will result in all uncached services becoming unavailable for as long as the repairs take. Unless prices for hardware go down again, which is unlikely to happen , or contributions pick up again, which in the current market climate is equally unlikely, I won’t be building out the setup any further. As for some general stats about this specific website, over the past six months it generated at its peak over 270GB of traffic per month, which is considerable, given its compact size and its GTmetrix rating of (100% performance, 98% structure, 450ms LCP, 0ms TBT, 0.02 CLS), with the first contentful paint after 344ms, the speed index at 366ms, the time to interactive at 436ms, and an average page size (including images) of a few megabytes at most. Because I stopped running any form of analytics software, due to privacy reasons at first, but ultimately because of the realization that it has become pointless with all the robots LARPing as legitimate human traffic, I cannot tell you the number of visitors this website or any of the related projects have, and frankly I don’t even care. Not because I don’t care about actual humans being interested in the things I’m publishing, but because over the past months I have been receiving an increasing amount of feedback via e-mail, as well as welcoming a surprising number of new members to the community channel , so that I don’t need analytics software to show me that visitors appreciate the work I’m putting in and decide to stick around. Speaking of sticking around: The SimpleX group that existed next to the primary XMPP channel is no more. Neither the platform, nor the group itself developed in a favorable way over the past years, despite its almost 500 members. SimpleX turned from an interesting privacy platform into yet-another-Telegram-type messenger and is on the way to enshittification with VC investments from e.g. ACP, who is an investor in K2 Space (ELINT), OnScreen.ai (tracking), Zorus (employee/network monitoring), as well as Jack Dorsey. To make matters worse, at the end of September Evgeny Poberezkin reached out to me via e-mail with the following request: Thank you very much for this article: https://xn--gckvb8fzb.com/an-overview-of-privacy-focused-decentralized-instant-messengers/ (we shared it here: https://simplex.chat/ <redacted>). We launched equity crowdfunding on Wefunder in August ( https://simplex.chat/ <redacted>), to give users the opportunity to get a stake in SimpleX Chat and benefit from its growth - 150 people have already invested. Please help us spread the word - maybe you could write about it and some other recent news: the foundation, channels, public names, and supporter badges, which we launched recently ( https://simplex.chat/ <redacted>). It would be great to connect - SimpleX link is below - please send any questions! Thank you again and all the best Evgeny P.S. If you write something, could you please share with us before publishing so it complies with Regulation Crowdfunding rules? We cannot promote (and won’t be able to re-share with our community) if it mentions valuation, security type, how much is or left to raise (saying how many people invested is allowed), and how the funds will be used. We can mention perks we provide though. Thank you! – Evgeny Poberezkin SimpleX Chat, Founder Despite my article clearly stating, quote: Red flag: VC funded, specifically by Jack Dorsey , since Aug 14, 2024, hence I do not recommend it any longer. … Evgeny (or, more likely, his team that didn’t properly vet every organic site linking/mentioning SimpleX) sent me this fairly generic looking e-mail, to which I replied by pointing at the red flag . Long story short, the fact that SimpleX has now seemingly become so much of a product in the literal sense that they are trying to get approved organic content out for their funding campaign shows once more that it isn’t a suitable option for the privacy/decentralization community anymore. I have therefore deleted the SimpleX group, as well as removed all mentions of the messenger on this website, treating it the same way as I treat e.g. Telegram, WhatsApp or any other commercial platform. Anyway, a few days ago I shut down the last Vultr instance, an OpenBSD machine that had been serving this website for years. It outlived the rest of my cloud infrastructure by a couple of days. The clearnet version of this website had moved from the instance to a Bunny storage zone with a pull zone in front of it earlier in September, which happened due to the increased traffic by bots and maybe humans , and which means that there is no origin server at all any longer. The Tor and I2P sites still needed a machine of their own, since Bunny doesn’t offer a way to serve either of them, hence both are now served from a virtual machine on my own hardware. That machine has no public IP address, which neither the Tor daemon nor i2pd needs in order to publish a site. Unfortunately it does appear to make the I2P site less reliable, presumably because an I2P router that nobody can connect to depends on other routers to introduce it, and in a measurement about four hours after the move, 4 out of 8 requests to the eepsite succeeded, while all 8 requests to the onion service did. If the eepsite doesn’t load for you, simply try again, since a second attempt usually works. The clearnet side has a trade-off of its own, namely that this website now depends on a provider being there, where the VPS used to be able to serve the entire thing by itself if I ever took the CDN away. What makes me comfortable with it is that the build is still a folder of plain files that any web server can serve, hence going back would require not much more than a server, a workflow step and a few changes in the DNS, but not a rebuild. There’s a dedicated write-up on the whole publishing pipeline in the works, so I’m keeping it at this for now. It’s been almost three months since I got the Lenovo X1 Carbon Gen 14 Aura and switched from the StarBook Mk VI to it , and I haven’t regretted it so far. The laptop performs as well as I had expected it to while offering very decent battery life. At times I’m getting dangerously close to its 32GB RAM limit, which is something I had expected before purchasing it, but most of my workloads still have plenty of breathing room. Shortly after upgrading to the X1 Carbon I also decided to get an external (portable) 19" screen, and so I snapped up the Uperfect GR19BU. So far my experience with the display has been very decent, but I won’t get ahead of myself as I’m going to release a dedicated write-up/review on it soon. However, with several other new additions to my setup, like this display, the Mudi 7 , and the things below , it’ll probably make sense to also start preparing a follow-up on the travel desk setup from earlier this year. As usual with many of the posts on this site, however, it takes between one and three months from the first line to the finished and published post, so don’t count on it until the end of the year or maybe even early next year. After having had a few minor issues with the uni USB-C hub (8-in-1 with USB-C power), and slowly running out of USB ports, I decided that it was time to get a proper Thunderbolt 4 dock for my X1 Carbon , and what option would be better than the official Lenovo ThinkPad Thunderbolt 4 Smart Dock Gen 2 7500, for which I happened to find a pretty sweet deal. As a matter of fact the deal was so good that I didn’t realize at the time of buying that the dock comes with A) its own external 135 Watt power adapter and cannot be powered through a regular USB-C PD charger, which is sort of messing with my minimal yet productive travel desk setup , because I’m already carrying a PSU that weighs over 800g, and B) with cLoUd CoNnEcTiViTy . The dock itself is 590g, its PSU (plus power cord) another 350g, which amounts to 940g total added weight for only the dock. Clearly the dock isn’t meant to be a travel-friendly option, but it seems like no Thunderbolt dock really is considering the (even heavier) alternatives, like the CalDigit TS-4. However, with regular USB-C hubs not providing enough connectivity, performance and, most importantly, stability, it was time to try a different class of devices. Sadly the dock only lasted a single day before it stopped working and had a slight smell of burnt electronics to it. Hence I had almost three weeks of back and forth with Lenovo’s customer service in order to get it RMA’d/replaced, which was a frustrating experience to put it mildly. Then again, I don’t think the lackluster performance of Lenovo in the specific geographic region that I had to deal with them in is a fair representation of Lenovo’s global customer service, but very likely typical customer service across most companies in that region. Anyhow, I eventually received the replacement and I’m finally able to use it. The hardware is fairly decent, I haven’t encountered any odd messages in my log yet, and I finally have enough USB-A/-C ports to connect all sorts of peripherals, like my keyboard, my mouse, the portable monitor , and more. However, I’d wish that these docks would replace their HDMI and DisplayPort connectors with more USB-A/-C ports and maybe even a halfway decent sound card with various audio connectors (3.5mm amongst others). I never needed that many display outputs on a computer, let alone on a portable device. But I guess I’m the exception, as pretty much every dock manufacturer seems to follow this quadruple/sextuple/octuple display-setup trend. The only meh part of the dock seems to be the integrated Ethernet card/port, that appears to be a Realtek RTL8156B ( ), which is a bit of a potato. There are plenty of user reports online that describe all sorts of issues with this specific NIC under Linux. While it does offer 2.5G, matching my new switch perfectly, it does seem to come with a few issues with regard to reliability and actual performance. However, with the amount of USB ports on the dock I can connect one of the many 1G Ethernet adapters that I have and be done with it, in case the integrated NIC should ever give me a hard time. So far, however, it had been working without issues. In an ongoing effort to USB-C-ify all of my hardware I also pulled the trigger on a new portable switch, namely the Ubiquiti Flex Mini 2.5G. I’ve been using the blocky Netgear 1GbE switches forever, but I’ve always struggled to power them, as they come with a barrel connector and a dedicated PSU. While I did at some point get a USB-A-to-barrel-connector cable, judging from the power output of the PSU those switches are not intended to be run off of a USB-A port. It works, but it’s probably not the best idea. Short story long, I went for the Flex Mini primarily because it is powered via USB-C, so that I don’t have to carry another PSU or worry about damaging the device. As an added benefit, the switch now offers 2.5G links, which the Lenovo dock , as well as the Mudi 7 and the Slate 7 can benefit from. Ideally my portable NAS should also benefit from the higher speeds, but sadly that probably won’t happen with the current hardware. Speaking of which … The Ultra-Portable Data Center (v2) has had its second birthday and has held up extremely well, considering all the travel and, at times, incredibly dusty, humid, and/or hot environments that it went through. More importantly, though, even two years later there is still not a single comparable, commercial product on the market that would make me want to replace the existing system with it. Every piece of hardware that I’ve stumbled upon to date appears to come with major hard- or software headaches, be it the UnifyDrive UT2, the Beelink ME Mini, the CWWK x86 P6, or even the significantly larger/heavier AOOSTAR R7. The closest in terms of software freedom, hardware reliability and physical footprint to date appears to be the QNAP TBS-464, which is an ultra-thin and lightweight, portable 4-bay M.2 NVMe NAS introduced back in 2021. Because I had to further reduce the weight and size of the items that I carry with me on my travels, however, at the beginning of September I began working on the next iteration of my ultra-portable NAS, the UPDC v3. I haven’t yet found the time to pour the endeavor into a dedicated write-up, but I am happy to say that I’ve managed to reduce the build’s footprint from initially around 1.48 liters down to as little as 0.68 liters in volume. I won’t spoil too much, yet, but it’s fair to say that I’m relatively happy with the end result and that it allowed me to make the UPDC a fixed part of my carry-on luggage . Not much happened on the keyboard front this quarter, apart from a set of PBS Black on Black that I was gifted, and that has since been on my Kunai . It has become one of my favorite keycap sets and it is pretty much what I was hoping for . The PBS keycaps are however not as rough/textured as I was hoping them to be, which is a bit of a bummer. Anyway, I might post a dedicated long-form review in a few months when I’ll be able to better judge long-term use. I invested some time in pursuing my open source projects in the past quarter, hence there are a few updates to share. Several of them shared the same two things, which were a move to the SEGV-1.1 license , and a vanity import path under my own domain. As some of you might remember, I had an unpleasant encounter last year with one of the Alacritty developers, which prompted me to do something that was long overdue, namely switch to Ghostty. Along the way I implemented the exact feature I initially wanted for Alacritty and made my code public, on a dedicated branch in my fork of Ghostty. Over a year later, I’m still maintaining that patch. If you’re using my patch, make sure to update your local version. Neon Modem Overdrive , my BBS-style command line client, received a whole new backend this quarter. It can now connect to Hyperuplink, the bulletin board of the future, to which I’ll get in just a second . Neon Modem is the first dedicated client implementation for the relatively young and not-yet-well-thought-out REST API. Additionally, a shared prompt package now handles the server URL and credential input for that system as well as for Discourse and Lemmy, instead of each one keeping its own copy. I also tracked down a longstanding source of UI lag. The post-create and post-show windows were doing more work on every keystroke than they needed to, and reworking their handlers and rendering made the interface responsive again. A new release with all of this is available on tty.fail , with binaries on GitHub . Kopi , the command line coffee journal, had its two heaviest dependencies replaced. The SQLite driver moved from the -based to the pure-Go , so Kopi now builds without a C toolchain and ships as a fully static binary on every platform. The OCR helper, which reads the details off a photo of a coffee bag, moved from an Ollama-specific client to the official OpenAI Go SDK , so it can point at any OpenAI-compatible endpoint rather than only an Ollama instance. The Go module was renamed to a vanity import path under my own domain ( ), and the usual dependency, workflow and GoReleaser housekeeping came with it. On top of this, Kopi received the first contribution , which was applied to it via at the end of July. reader , the command line web page reader, had a productive quarter. The biggest change is the removal of the journalist dependency, whose crawling logic is now part of reader’s own package rather than pulled in from a separate module. As I don’t maintain journalist any longer it made sense to pull out the bits that I need for reader. Additionally, reader gained proxy support, so it honors the standard proxy environment variables (#23). It now also correctly handles the URL scheme when deciding whether a source is an HTTP resource (#35), and it got a fix for an -related issue (#37), which I worked out despite the reporter never supplying the example files I asked for. Note: Don’t be that guy . If you open an issue in any piece of software that you don’t pay for and that someone else dedicates their time to free of charge, consider upfront if it’s really something that you’re going to use, and if it’s an issue you’re willing to dive into and help fix. If either of these are a clear nope , then simply don’t report it at all. usbec , the USB Equipment Commander daemon that runs commands when USB devices are plugged in or removed, had its one core dependency reimplemented in-tree and enhanced. The module was dropped and its Linux netlink listener rebuilt inside usbec’s own package, so the daemon no longer depends on an outside module for that. The license moved from GPLv3 to the SEGV-1.1 license , the module was renamed to the vanity import path, and the release workflow and GoReleaser config were fixed along with it. In addition, the USB Equipment Commander is now a USB & Network Equipment Commander, with the introduction of NetworkManager (via D-Bus). This allows you to have usbec run commands whenever the network changes in any of the various ways it can change. My motivation was to run a script that queries my public IP and the inferred location via and shows it as a desktop notification whenever NetworkManager (re-)connects. zpoweralertd , my Zig rewrite and drop-in replacement of poweralertd, which is a UPower-powered power (.. power, POWER, POWERRRRR) alerter, has received a deep and thorough refresh. The biggest change is the C bindings, which I have replaced with handwritten s, just the way I had used them for the first time in ssh-askpass-zigtk , but more on that new tool below. The second big change is the Zig compatibility, which now requires at least 0.16, and which I have tested against the current master/future 0.17 release as well, to make sure that everything will continue working. With this, I also largely refactored the code, for which I initially took inspiration from poweralertd. The restructuring allowed me to clean up the codebase and find a handful of memory-related issues along the way. A new version of zpoweralertd has been released that contains these enhancements and, for the first time, it is available as pre-built binaries directly from the releases page. cexec , my small cached-exec wrapper that runs a command and caches its output for a given amount of time so that re-running it returns the stored output instead of executing it again, was rewritten from Go to Zig this quarter. The new version is a drop-in replacement for the previous one, but it changed a few things in the background, one of which is the storage. Previously cexec used BuntDB, which turned out to be a bad choice for a process that can theoretically run multiple times in parallel and access/write to the same database. The new version, instead, uses plain files that are scoped to the actual commands that are being executed. In addition, the new Zig version supports encryption for the cached runs, with an optional key from or the environment variable, which makes using cexec suitable even for output that contains sensitive data. On top of that, everything cexec does is now also available as a Zig module, so that the same caching can be used from another program without having to shell out to the cexec binary. New this quarter is , a GTK4 helper written in Zig . OpenSSH runs an askpass program to prompt for a passphrase or a yes/no confirmation when it has no controlling terminal. The reason it exists is my own issue with X11, which I explained here . The reference askpass, and most others, link GTK against X11, so they fail to build on my Gentoo system with the global USE flag. does not require X11 dependencies, so it builds and runs on an X11-less GTK4. It works on Linux and (hopefully) on the BSDs alike. Also new this quarter is Switchyard , a lightweight bridge that accepts e-mail over SMTP and forwards each message to XMPP. You might have read about it before in the dedicated post that I had published at the beginning of August, but if you haven’t I recommend going through it if a bridge for SMTP to XMPP sounds like something you could have a use for. For years I had a small shell script in my dotfiles that popped up a bemenu list whenever I clicked a link, so that I could decide which browser it opened in. This quarter I turned that idea into a proper GTK4 application written in Zig, namely Browser Select . It registers itself as a web browser, hence every link another application opens goes to it first, and a small popup then lists the browsers installed on the system for you to pick one with the mouse or the keyboard. Just like the script before, that used the Go version of cexec as a binary to cache the choice for a few seconds, Browser Select also uses the new Zig cexec library for remembering a choice for a while, so that clicking a burst of links doesn’t lead to repeated asking for a browser. The Flipper BUSY Bar review that I posted this quarter came with a little goodie: busybar.zig . It is a Zig 0.16 client library and command line tool for the Flipper BUSY Bar, implementing the device’s current OpenAPI specification as closely as possible, so that you can drive it over its HTTP API from Zig code or straight from the shell. It uses nothing but Zig’s library, which means it builds for every target Zig supports, including macOS and Windows. Darkwing Ducky is a BadUSB (“Rubber Ducky”) firmware for the PicoUSB, an RP2040-based board, built in Zig 0.16 with MicroZig . I’ve struggled to finish this for about ten months or so, primarily because implementing the required USB HID code that acts like an external keyboard wasn’t all that straightforward with MicroZig. Anyway, Darkwing Ducky announces itself to the host as an ordinary keyboard and then types out a predefined sequence, which can be useful for automating the setup of a machine that you have physical access to, or, well, for other things. While the PicoUSB comes with its own CircuitPython-based software, I wanted to try Zig on hardware, especially with MicroZig for a while and so this device was the perfect excuse to do so. If you’re curious, the repository has the details on the payload format and how to flash it. The internet bulletin board software that I had been writing about as ▓▓▓▓▓▓▓▓▓▓▓ in the previous updates has a name, and it’s out. Hyperuplink went public at the end of August, together with a dedicated post that covers what it is, why it exists and how to run it. If a JavaScript-free forum as a single binary sounds like your kind of thing, read that one first and come back here afterwards. Otherwise you might as well skip this part. Most of the quarter went into the bits and pieces that the board needed before anyone other than me could run it. Attachments and profile pictures with local or S3-compatible storage, board-wide settings and an administration UI for all of them, a search, soft-deletion for categories, forums and topics, and a manual that is embedded in the binary and available under Help -> Manual . On top of that came a whole set of themes and color schemes, as well as the packaging, which now covers Docker and Podman images, Quadlets, a Helm chart, a nixpkg with a NixOS module, an ebuild for Gentoo, and service files for systemd, OpenRC and rc.d. And if that wasn’t enough already there is the REST API, which is what Neon Modem connects to. If you’d like to have a look at Hyperuplink yourself without any commitment, there’s a demo instance that you can check out. You can find its login details on the official website . Account creation is disabled on the demo instance (for now). Go compiles down to a single static binary, which is why it is a great choice for Hyperuplink. However it has nothing along the lines of Django, Rails or Phoenix that takes care of the web parts, hence a good share of Hyperuplink was never about the bulletin board at all, but about routing, controllers, models, migrations, background jobs, translations and configuration. At the end of July I moved that share out of the repository and into Go on Glides , a web application framework for Go that includes everything needed to build database-backed web applications, on top of Fiber v3 and pgx , with embedded migrations, an asynchronous job queue, cron jobs, i18n, local and S3-compatible storage, and helpers that each project would otherwise have to reimplement. The reason for the split is that I’m building a second service on the same foundation, namely Maya , which meant that Hyperuplink’s internals had to become something that more than one program can use. Inca is the spiritual successor of addrb and caldr , my two earlier attempts at the same problem, this time as a single command line client for CalDAV and CardDAV that synchronizes calendars, tasks and contacts into a local database for offline use. It is built on top of the “framework” that I built and make use of in zeit , and therefore it shares a lot of zeit’s UI aesthetic. On the other end is Maya , a CalDAV and CardDAV server as a single Go binary with PostgreSQL underneath it, which (sadly) uses a patched hard-fork of go-webdav , and which exists because neither Radicale nor Baïkal ever made me happy. It’s the second program built on Glides . I’ve been dogfooding Maya myself since early September, and as of now it still isn’t publicly available because it’s not yet at a point that I’m comfortable releasing it, in particular due to its hard-fork of go-webdav, which was a requirement because the upstream project is sadly pretty much dead at this point. However, given the absolute mess that DAV protocols are, it’s not even surprising that the maintainer seemingly lost interest in this project, and that the ecosystem is so bad. Anyway, I have both, Inca and Maya, at a point that I can daily-drive them and, more importantly, that other clients (Android, iOS, desktop) can participate without too many hiccups along the way, but I’m not at all happy with the results for reasons that I will get into if I ever choose to release Maya. Also new this quarter, and admittedly not something one might call a modest undertaking , is Netrunner , a web browser built on WebKitGTK and GTK4, written in Zig . I’m calling it the hacker’s browser as it’s meant for people who prefer a lightweight browser over Chromium and Firefox and who care more about configurability and about built-in integrations than about endless features and add-ons. Netrunner deliberately leaves out a number of things that other browsers have, bookmarks being probably the most obvious one, because the people it is built for tend to have a system for those features in place already. Netrunner is configured through a single TOML file that holds the static settings as well as the runtime state, e.g. which sites are allowed to run JavaScript, play audio or send notifications. The session (i.e. open windows and their tabs) is also persisted in that one TOML file when the browser quits, and restored on the next start. Two other features that Netrunner comes with built-in are Onion addresses and Eepsites, which are native protocols for the browser. By default, whenever you open an URL or an link Netrunner opens a connection via Tor or i2pd and keeps it alive for as long as you browse sites within the Onion/I2P network. As soon as your browsing activity goes back to the clearnet, Netrunner automatically disconnects from Tor/I2P again. To the user these connects/disconnects happen completely transparently and every network gets a of its own, so a tab that follows a link onto another network moves over and keeps its back and forward history. In plain English this means that and URLs are indistinguishable from any other URL and you don’t need to do anything special to be able to open them. Obviously Netrunner is not a replacement for the official Tor Browser that the Tor Project releases and that is somewhat hardened to make sure you stay as anonymous and safe as possible. Instead, Netrunner’s Tor and I2P integrations are intended for the casual user who enjoys the freedom of browsing those networks as if they were part of the “normal internet”. Speaking about the normal internet, one thing that can’t be missed, especially these days, is content filtering. While WebKit sadly doesn’t support running a full-blown uBlock Origin, its own content filter engine is nevertheless a relatively decent option for basic ad blocking and annoyance filtering. As a matter of fact, Netrunner uses the same AdBlock and uBlock Origin lists that popular browser extensions use, but it does so by downloading and converting them into the WebKit content blocker format via adblock-webkit-convert.zig , one of two small Zig libraries that came out of this browser experiment. Sadly, the WebKit content blocker format does not support all the trickery that e.g. uBlock is able to do within its browser extension, which means that no, you won’t be able to avoid YouTube ads within Netrunner, at least for now. You will however still score above 62% on AdBlocker test sites , which is fairly okay. The other library that came out of this experiment is osdetect , which does what its name suggests: It detects with some amount of confidence what operating system a program is running on, and it gives you specifics like whether it’s e.g. a Debian or a Gentoo machine, or not even a Linux at all. I won’t tell what my main motivation for building this library was, but if I should ever get to the point that I have a 0.1 version of Netrunner ready for release you will be able to the source code for to find out. Which brings me to the release. Not only to that of Netrunner or Maya, but to the more general question of open source and how I intend to contribute to the ecosystem moving forward. Despite having a landing page online, Netrunner is a fun project that I don’t take too seriously and that I’m currently dogfooding, just like I do with Maya , to find out whether this is a project that I can imagine maintaining long-term, or whether it is one that will eventually outgrow my availability. After two decades of putting the software that I build for myself online, I have come to realize one or two things. The first thing is that “just putting it out there” won’t benefit the software and it certainly won’t benefit me. My main motivation for building all these things is my own curiosity about how stuff works and, ultimately, my own needs. I need a proper CalDAV/CardDAV server that is pleasant to administer/maintain and that can talk to all of the software that I’m using, hence I’m building Maya (and Inca) . I need a proper TUI client for Keebtalk and a few other forums, hence I’m building Neon Modem Overdrive . Speaking of which, I needed a forum and I was sick of the cruft that is phpBB or Discourse, so I built Hyperuplink . And it’s also me who is annoyed by the utterly insane bloat of modern web browsers, which is why I started building Netrunner. Publishing any of them inevitably changes the software from “what I need” to “what I need and what others seem to need” , and there is no guarantee, even if you go the extra mile to implement what others ask for, that those people will stick around or ever start contributing themselves. Zeit is one example where I experienced exactly that, multiple times, and every time it left a bitter taste in my mouth and, what’s worse, hundreds of lines of code that other people had requested but that nobody ended up using long term. At some point Zeit v0.x had become so bloated with other people’s needs that I rewrote it from scratch as the v1 release that it is today, with only the features that I believe make sense to support long term. The second thing that I realized is that “putting it out there” in hopes of finding like-minded people to collaborate with is a pipe dream. The open source “community” has been broken for a long time and, in most cases, was never much of a community to begin with, but rather the developer-equivalent of a one-man band . Not that there are no exceptions, however, my projects usually aren’t those exceptions. Kopi received its first contribution only this quarter, two years after I initially released it, and it was a one-time submission rather than something that materialized into a long-term collaboration. reader , in the very same quarter, collected a handful of issues, one of which I fixed without ever receiving any feedback on it whatsoever. And if you look through the few “Issues” sections on GitHub that I haven’t disabled yet, you’ll see that it’s the same story across the other projects as well, and you might understand why most of the repos lack the “Issues” tab these days. A single patch against two or three dozen issue reports is roughly the ratio that every single one of my projects has had over the past two decades, and I don’t believe that this is because my software is particularly unattractive to contributors, despite it being intentionally niche, but simply because writing a patch for someone else’s program is work, while filing a request is at most a two-minute annoyance. This is also why I, contrary to many other voices in the industry, do not regard opening an issue report as an actual contribution to a project. With all due respect, I don’t want your issue reports, keep them, especially the drive-by ones for something that you’re unlikely to end up using anyway. What a public repository attracts, hence, is an audience rather than a community , and many times that audience has a very short attention span. All of this eventually turns a maintainer into more of an unpaid support desk, which is not what I would like for myself. On top of that, LLMs have taken away the one argument that publishing open source software still had going for it. That argument wasn’t that a stranger would show up and maintain your software for you , but that someone out there might one day need the exact same thing badly enough to extend what you had already built, instead of writing it from scratch all by themselves. In 2026, however, that someone doesn’t need your repository any longer, because they’ll describe what they want to a model and get back something that compiles. Whether what comes out the other end is any good depends on the person directing the LLM and on the effort, or, dare I say, the money they’re willing to throw at it, rather than on the LLM itself. Code itself has become relatively cheap . As a matter of fact, if you want a specific feature implemented in any of my years-old tools, you might as well ask the machine to do it for you, so you can have what you need for the time being. And frankly, I’m quite happy about that! I’m happy that a random passer-by can get the feature that they thought they needed during their euphoric discovery phase of a tool they just found out about, mainly because I also know that the moment the honeymoon phase ends and the person loses interest in the tool, their interest in that feature will vanish along with it, and it would be left up to the maintainer to continue supporting it. Don’t get me wrong, I’m not arguing that collaboration never works, because examples like Linux, curl and PostgreSQL prove the opposite, but I am arguing that for small programs that a single person can hold in their head there are very few reasons for others to actively contribute, and even fewer in a world of automated code generation. Going back to where I started, I’m not sure yet whether I’ll ever publish any of the newer tools and programs that I’m building, like Maya or Netrunner, simply because I don’t see much of a reason to do so. Especially when, for projects like Netrunner, I know of far more popular examples that still died the moment their maintainer stopped investing time into them. And maybe that’s a good thing. Maybe these kinds of projects, just like the ones that I’m publishing, are simply not relevant enough in the grand scheme of things. Maybe it’s just Darwinism at work, and maybe that’s just how things should be after all. There he goes. One of God’s own prototypes. Some kind of high-powered mutant never even considered for mass production. Too weird to live, and too rare to die. – Raoul Duke

0 views

September 2026 blend of links

Some links don’t call for a full blog post, but sometimes I still want to share the good stuff I encounter on the web. Talking Watches: Pusha T On Why Rolex Is Worth The Wait – Let’s be real, I’m sharing this video mostly for the perfect coffee brand name Pusha T is launching with Lavazza (and promoting in this video). Bonus points if you guessed it. Repairing a 100-Year-Old Camera – Beware, this video will take 60 minutes of your time, and you won’t get those precious minutes back. (via Anthony Nelzin-Santos ) Lucky Notes – I recently wrote about this little app, but it deserves a spot in this series . If you were wondering what the cutest notes app is, there you go. 8 years and done – “ Now, I find it increasingly difficult to find topics on which I can offer thoughts worthy of other people’s reading time. (Truth be told, I probably haven’t been able to come up with that kind of content for a while now, but I soldiered on anyway.) ” Always sad to see a blog end, but this quoted part really resonated with me. (via Kev Quirk ) Meanwhile – Meanwhile, Daniel Benneworth-Gray relaunched his blog, as a perfect complement to his excellent newsletter . Pros and Cons of Using Personal AI Agents – Another instant classic. Muse Is What Meta Means by ‘Personal Superintelligence’ – “ Perhaps I am not enough of a dutiful consumer to be so amazed about a robot that can buy things for me. That is not to say there is nothing I like […] but there is a huge disconnect for me between the amount of technology here and what it could actually do for me in real life. ” I feel a lot like Nick Heer about this, and I would even go further. The one thing that is bothering me about this whole agentic lifestyle is the idea that tasks and chores are on the same level. And what do these people do with all this alleged extra time anyway? One Day on Rails – Pretty quiet until four in the morning, then captivating. (via 82MHz ) 100 unusual crisp flavours, ranked – Now, this is unexpectedly complex investigative journalism. Why Am I Left-Handed? – “ Sometime after the emergence of the genus Homo, 2.8 million years ago, evolution coded a preference for right over left. Of the various speculative theories for the emergence of this extreme preference, one that seems plausible to me points to our unprecedented capacity for violence. ”

0 views
Jim Nielsen 2 days ago

VLM Enhanced Metadata For My Icon Galleries

Confession: I got nerd-sniped by Sam Henri Gold’s request for my icon galleries: I'd like to humbly request artwork-level searching in macosicongallery.com What follows is a train-of-thought blog post as I play with what an implementation might look like. I’ve actually long-wanted something like this, e.g. let me search for “coffee” and show me all icons that have some depiction of coffee in them. Similarly, I’ve wanted some kind of “related” representation for icons. I have this today via existing metadata, e.g. “Show me other icons in the category ‘Productivity’” or “Show me other icons tagged as ‘orange’”. But I’ve wanted a more robust representation of this, so if you were looking at an icon that had a microphone in it, the site would say “Here are other icons that also have microphones in them.” And the relationship would be rich/smart enough to know that “microphone” was meant broadly, i.e. dynamic mics, condenser mics, ribbon mics, etc. So how would you do this? I could go through every icon one-by-one and classify/tag any attribute of its design that comes to mind, but that would take ages! Seems like a good use case for a vision model. First, I’ll look at Sam’s suggestion: run every image through CLIP. I’m not familiar with CLIP so I start with a little research: What is it? How would I use it? And most importantly: is it free/open (because I ain’t spending a ton of money to send my thousands of icon PNGs to an AI provider via their API)? Ok, so CLIP will take an image and spit back an embedding (basically a bunch of numbers representing features of the image). When you do it with multiple images, you can then compare those embeddings to see what the model considers similar (and, if you like, set a threshold for what constitutes a “match”). After getting a sense of the task in front of me, I work with the LLM to come up with a proof of concept. I don’t need to fit this into my existing site. I just want to make one-off HTML pages where I can feel out, “Can this process create anything useful? What’s the amount of work required?” This is enough to create a single HTML file where I can click on an icon and see other icons that look like it. However, I realize quickly that I’ll need to process my entire icon library to really get a good sense for how well these are matching. So I do that. [Computer goes brrrr…] Ok, now when I click on an icon that looks like a camera, I see other icons that look like cameras. Or if I click on an icon that has a checkmark in it, I see other icons with checkmarks in them — sort-of. But the results aren’t that great unless an icon is visually distinctive. I share some thoughts with Sam. He has a few other suggestions I follow. Sam mentions SigLIP2 so I start with that as a keyword. The LLM recommends DINOv2 so I say, “Let’s try it”. I give that a try, creating a separate dataset and prototype (e.g. and ) so I can continue to view these different prototypes and compare their outputs. It’s fine. Different from CLIP. Honestly not much better. So I figure let’s try another one. I go with SigLIP2. I ask the LLM to create a page where I can compare the results. Seems like six of one, half dozen of another. One does better on some kinds of icons, worse on others. The LLM recommends that, at this point, I be done shopping models. They’re roughly the same class of tool with different tradeoffs. None are breakthroughs. So now what? Sam recommends another approach: You could also try handing all icons over to a VLM, having it write up a description, and embedding THAT text against what people might search for. A thoroughly detailed person might’ve done this from the start, e.g. for an icon that’s a checkmark, add the keyword “checkmark” to its metadata. That would take me forever to go back through all my icons and do — a perfect task for a computer that never tires. So I give this a try. First I need a free/open VLM. After a little research I decide to try Moondream via Ollama . I have the machine go through each image and caption it, then pull out “tags” from the caption. For the Clear app icon , I get data like this: Then the LLM creates a single file where I can test icon matches by searching for tag overlaps (or choosing one of the popular ones). So, for example, on the search page I can click on “checkmark” and see all the icons with a checkmark. Or click on “fox” and see all the icons with a fox. Matches are pretty spot to be honest. But that’s a different kind of test than what I was doing with CLIP. Can I leverage tags for the same kind of “related icons” work that CLIP is doing? I get the LLM to cook up a single-page HTML file where I can compare “click on this icon and find other icons like it” where I’m using embeddings from CLIP vs. matching on keywords. The results seem to fare much better for CLIP. For example, here I matched on what I think of as a “checkmark icon”. You can see the approach that matches on tags didn’t work too great. I believe this is because with my simple tag-overlap approach, a distinctive keyword like “checkmark” gets diluted amongst generic tags like “square”, “blue”, and “simple”. Whereas with CLIP, if you click on an icon with a checkmark, you get other checkmarks (and not other icons that also have related tags like “square”, “blue” and “simple”). Which all makes sense. Pushing on the implementation here could help, but that’s separate work to do. I’m not sure. While doing all of this was an interesting technical exercise, there are a few important considerations I need to think through before implementing anything, such as: I’m very picky about adding new dependencies to these icon projects. I like to think that’s why I’ve been able to maintain and continue contributing to them after so many years — because I make it easy on myself (good job, Past Jim). So, for example, if I make a VLM a dependency of this project such that every time I add a new icon I have to run it through to create the embeddings, that’s a big dependency cost IMO. I’m not sure I want to do that. That said, Apple now ships foundation models in macOS 27 available through the CLI (go ahead, try typing in your Terminal if you’re on Golden Gate). So if my Mac continues to be the primary machine where I add/update metadata for my icon projects, using would be a really easy/low-cost way to process each new icon to generate a caption and keywords for matching in search. But again, I don’t know if I want to do that. I wrote this post to try and work through what I want to do, but I am still undecided. So I guess the only thing for me to do at this point is hit “Publish” on this post and keep simmering on a decision. Reply via: Email · Mastodon · Bluesky Write a script that runs a sampling of icons through CLIP’s image encoder Read the file locally, e.g. 256x256 pixel icons seem to be enough, as the CLIP model I’m using preprocesses them to ~224px anyway. Create a dataset representing the “embeddings” (an array of numbers) for each icon that I get from CLIP, e.g. Create a dataset representing the top matches between different embeddings, e.g. Create a file has both datasets (plus supplementary icon metadata I already have), render all the sampled icons, and support an for each icon that shows the icons. CLIP: click on an icon and see other icons like it. Tags: click on a keyword and see other icons with that keyword. What kind of functionality do I actually want? A “related icons” feature? Does it match on keywords or embeddings? A “search” feature that matches on keywords? How do I build these features into my codebase now, given the thousands of icons that already exist? How do I maintain this feature in the future? e.g. every time I add a new icon to my gallery, is a VLM now a dependency of this project? Given all the above, what’s the time and money cost?

0 views