Latest Posts (20 found)
iDiallo Today

The Front Page of the Internet Is Up for Grabs

Last week I opened Edge to test my Internet connection, because that's the only thing I use it for, and noticed something new. Right beneath the address bar was a new bar with a button offering to let the browser launch automatically whenever I start my computer. That was strange. Yes, I open the browser the moment the computer starts, but why would they want to do that for me? And just now, I opened Chrome, and saw the exact same message but for Chrome: It looks innocent enough. We all start browsing the web right away when we turn on our computers, so what's the big deal here? Well, the big deal is that everyone wants to be the front page. Defaults matter. While we all want to believe we have agency, the numbers don't lie. Most of us never change the default settings . Just last year, Google paid Apple close to $20 billion just to remain the default search engine on Apple products. While we're all distracted by the AI race, the one race that actually matters is who remains the front page of the Internet. Microsoft is trying to claim that spot by loading its browser at Windows startup. Now Google is trying to win that space back by doing the same with its own browser. On my Mac, Anthropic has been very aggressive about trying to load Claude automatically. I suspect every company will be pursuing this aggressively in the near future. The quality of your product is secondary to the exposure it gets. Not even benchmarks matter if you can't be on the user's front page. Google still holds the front page of the Internet, and that makes up for its AI model not being at the top of the benchmarks. But don't let them take this space from you. They are all fighting for your attention. I urge you to simply consider pressing the little X button they tuck into the corner. It's your computer, use it the way that suits you best.

0 views

Initial thoughts on the EU KIDS Act

The European Commission has proposed the EU KIDS Act . It is yet another set of Internet/web regulation proposals to examine, for jurisdictional overreach, lack of common sense in terms of material scope, and so on. It is only a legislative proposal at the moment, so it may not become law, and it may not become law in this form. Based on a quick skim, this is indeed another fine mess, full of unrealistic expectations. The proposal covers a lot of services: I have read this from the perspective of online social networking services, thinking predominantly about Mastodon and other fediverse services. At least code forges are out of scope (“open-source software-developing and-sharing platforms”). And Wikipedia seems to have its own bespoke exemption (“not-for-profit online encyclopaedias”). Small, low risk services are in scope. The covering material specifically notes: small and micro enterprises are not exempted from this Regulation, since they may equally provide harms to minors. It would undermine the objective of this proposal to exclude them from scope Wow. I wonder if the drafters will realise just how harmful this is. In terms of territorial scope, it is broader than the EU GDPR, and indeed the UK’s Online Safety Act, purporting to apply to providers of services irrespective of where they have their place of establishment where they offer those services to recipients of the service that have their place of establishment or are located in the Union I wonder if anyone working on this stopped to think about the boundaries of their laws, and whether they really think that they can impose obligations on people in other countries, merely because that person is running a service which happens to be available to people in the EU? Do I, as someone who runs my own fedi server, where people in the EU can read my toots and respond to them from their own instance, fall into scope? I do not know. Providers of online social networking services … shall not allow a natural person below the age of 15 years to create an account with that service or to access that service by means of an account, created for, or attributed to, that person, where the service poses a risk to the privacy, safety or security of a minor below that age. (Article 6(1)) The tests for “poses a risk” set an incredibly low threshold, and include: enables recipients who access the service through an account to transmit content in real-time to an indeterminate number of other recipients of the service, including through live streaming of audio-visual content enables recipients who access the service through an account to contact, communicate and otherwise interact with other recipients of the service not part of the recipient’s pre-existing connections or subscriptions So a “papers, please” web would become the norm, according to this. For example: When creating an account for a minor pursuant to paragraph 2 of this Article, the provider of online social networking services … shall take measures to establish whether the person creating the account is the holder of parental responsibility over that minor in accordance with Article 26 and verify that the recipient of the service has reached the age of 13 years in accordance with Article 28(1). (Article 6(3)) Article 26 sets out how the European Commission envisages this working, but, wow, I just don’t see it. Harking back to (what should be the exceptionalism of) broadcast regulation, there’s another banger: Providers of online social networking services… shall put in place effective measures to ensure: time-limited access for minors on their service; interruption of usage by minors on their service. Such measures shall be designed in a way that protects school time and core sleep hours of minors. (Article 9) Sorry, I have to turn off my fedi server now, because a child in a different timezone might be heading off to bed and my toots might be distracting… Some of the proposals seem to relate to core browser functionality: Providers of online social networking services … shall put in place measures to ensure that settings are set by default to a high level of privacy, security and safety of minors. To ensure compliance with this paragraph, such providers shall, by default, turn off at least the following settings: other recipients of the service shall not be able to download or take screenshots of contact, location or account information of minors or of any content uploaded or shared by minors on the service; I have no idea how the drafters of this expect the provider of a social media service available via a web browser to restrict screenshots of everything posted by a user. It is not within their gift. The only way to make this work would be either to force all access to be via an app (which would be daft), or preclude child access (which has age verification challenges). Some of the use restrictions would seem very challenging: Providers of online social networking services … shall put measures in place that ensure a high level of privacy, safety and security of minors as regards contacts between minors and other recipients of the service. Those measures shall at least ensure that: other recipients of the service are not able to initiate direct contact with the minor, if the minor has not pre-approved such contact So a 17 year old here posts something interest. No-one is able to interact with their post, unless the 17 year hold has “pre-approved” it. Oh, don’t worry, you won’t be able to see their post anyway: by default, other recipients of the service not previously accepted by the minor shall not be able to access account information of the minor or content uploaded or shared by the minor on the service online social networking services; video-sharing platform services; software application stores online games; operating systems; AI companions; general conversational chatbots. time-limited access for minors on their service; interruption of usage by minors on their service. access to microphone and camera

0 views
Unsung Today

“Insert your favorite nursery rhyme.”

Make Some Noise is one of many fantastic improv comedy shows on Dropout . I noticed one of the ongoing themes is “tech gone bad,” so I compiled a short list below. I just find these really funny. Video call with shitty wifi : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt1-play.1600w.avif" type="image/avif"> Logging in with 30-factor authentication : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt2-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt2-play.1600w.avif" type="image/avif"> A cutscene in a video game that’s glitching out : = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt3-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/insert-your-favorite-nursery-rhyme/yt3-play.1600w.avif" type="image/avif"> In case you’re curious, the improv actors are: Zac Oyama, Josh Ruben, and (occasionally) Brendan Lee Mulligan. (If you’re a Dropout member, these are the original, slightly longer landscape videos inside the respective episodes: wifi , 30-factor , game .)

0 views

I don't like LLMs

I have a lot of mixed feelings about AI and LLM technology. I’m fascinated by its effect on our profession, excited by the potential gains in productivity - and thus the products we could rapidly build. On the other hand, I’m fearful of the damage AI might cause: agent swarms taking over our virtual and physical infrastructure, designing bio weapons. But, back on my first hand, LLMs might also design miracle cures, and come up with clever ways to raise our prosperity. Fundamentally I don’t think we have a choice about riding on the AI technology train. It’s a wild ride and I just hope we’ll get through it OK. But as I mull on this more, I realize that among this mix of contrasting feelings, there is one emotion that dominates - one that comes from my direct interactions with LLMs. I don’t like them. They talk to me in this grating LLM-voice, an uncanny valley of talking to a real human. They confidently bullshit me - often giving me useful, helpful answers. But also just making stuff up with the same assurance - and with only a veneer of fake remorse when I call them out on it. That’s not enough to make me feel we should avoid them. As Jessica Kerr put it “not only are they useful, it is irresponsible not to use them…. They’re more thorough, as well as faster.” This contradictory reaction comes through in polling , where people say they find these models are useful, but also that they think they will be bad for society. Much of this may be because LLMs are young - we haven’t trained them to grow up yet. Maybe I’ll like them once they mature. (I hope we get to find out.) But I’m not encouraged when I think of the kinds of environments that cultivate them. I’m wary of the Silicon Valley brogrammer subculture, and these LLMs are their products, so naturally lean toward their world-view. When we think of AI agents, we shouldn’t anthropomorphize, treating them as conscious beings with their own will. They are (software) machines, developed by people working in corporations. While the agents’ behavior aren’t explicitly programmed, they are nurtured with the values of their creators. One of my most successful life-hacks is to avoid people I don’t like or don’t trust. I decline to interact with them socially, and make a deliberate effort to avoid working with them too, even if they are doing much that is beneficial. I feel that hanging out with pleasant, capable people, the people with integrity, has made my life a far better one. Hence my visceral dislike of interacting with an LLM that’s not just making a pretense of being human, but also posing as the kind of human I walk away from.

0 views

The Avocadinator

I think I bought this thing at some grocery store endcap or something. The magic is that it can do all three advocotasks: Amazing! All in one! My review: I actually kinda like it. I do reach for it when I have this job. But I’ve shown it to several people who have almost viscerally bad reactions. Like, they can’t explain it; they are just: nope. I literally handed it to someone who was about to undertake the job of slicing avocado, and they were like  no thanks, got my knife here. The downsides are: Cut avocado in half (little knife thing) Take out the pit (little indent with teeth thing) Cut avocado halves into slices (grate thing) The avocado has gotta be pretty ripe for it to work nicely (which it should be anyway if you’re about to take on this task). The device is immediately filthy and much harder to clean than a knife/spoon. It just has to go in the dishwasher immediately.

0 views

Five Striking Cycling Bridges Nearby

Our province, Limburg, is known in Belgium for its cycling-friendly environments. This is not just a statement uttered by its inhabitants but also actively pushed forward by politics as one of the many strong points of Limburg: hence Fietsparadijs Limburg ( Cycling Paradise Limburg ). I am privileged to be able to make good daily use of these cycling paths that normally are neatly separated from the usual road traffic. But it’s not just the greenery or the maintenance of (most) cycling paths that’s striking: it’s also some of the bridges that take you across a national road or a canal that stand out. I thought it would be fun to make a few snapshots to highlight a few of these bridges. During the period 2007 to 2024, fifty (!) bridges across the Albert Canal—Belgium’s most important waterway from Antwerp to just above ALiège—were rebuilt to be able to support four-layer container ships, as the previous bridges were too low. See Upgrading of the Albert Canal , an official site of the Flemish Waterways, for more information. One of these nearby recently restructured bridges used to carry trains across the canal. Naturally, it was inaccessible by bikers. The new bridge now carries both trains and humans on bikes across, connecting two sides of the waterway in new ways. It might not be a big deal as the previous (car and cycling) bridge is just about a kilometre away, but this is very convenient for my daily commute. Oh, and if you take a break on top of it, you can peek down inside the courtyard of the adjacent prison. I’m not sure you can throw smartphones and other shady packages that far though. The new and higher train/cycling bridge above the Albert Canal. Left: train tracks. Right: cycling track. Note the waterway below the bridge visible on the lower right. A good three kilometres north, you’ll bump into one of the six canal locks needed to accompany for the difference in elevation of 56 metres between Antwerp and Liège. This one, in Diepenbeek, has a lift of 10 metres. The lock control tower on the upper left is not new, nor is one of the many windmill in picture, but the winding cycling-exclusive bridge standing on top of big concrete pillars is. It’s part of the cycling network highway known as the Fietssnelwegen or F network that allows faster e-bikes to speed through at ease without hurting normal bikers like me. Normal traffic can pass on top of the tarmac-laid lock as usual, but for bikers it’s a little bit safer now. Except that you still have to cross that same lane at the end of the bridge, so I’m not sure what they were aiming at here. It’s certainly striking though, especially because of the weird arcs. The cycling-exclusive bridge over the Diepenbeek Albert Canal lock. Still in the rough area of Hasselt/Diepenbeek/Genk, we find a bridge that’s not really a bridge but rather a way downwards . In the Provincial Domain Bokrijk, you can cycle straight through the water and meet the ducks and swans eye to eye instead of from the usual height. Too bad the sides aren’t made from some kind of translucent material. Our kids love this, and it’s a real tourist attraction. The 'bridge' below instead of above the water? A bit farther from home, a train ride of forty minutes away, you’ll find the city of Leuven with its many iconic historic buildings including the station. Yet the building itself isn’t the most striking aspect of the station: the three-sixty degrees cycling bridge is, taking cyclists from the station square up towards the business complex nearby. Or how about going even further up, turning right, and cycling above the train station and its tracks to get to the other side of Leuven? As a traveller waiting for your train to arrive at your track, you can literally look up to see cyclists passing by! The only downside of this bridge is the pedestrian stairs cutting through the centre that I used to take to walk towards the station. It’s quite crowded and crossing this bike line can be a challenge. I waited for five minutes before I could take this picture with only a few cyclists in the far distance. The 360 degree bridge at the station of Leuven (taken from the business district). Back to our own university college campus. In order to connect the campus with the other side of the busy national road that leads to more parking options and another connecting road to the northern part of Limburg, 1500 tonnes of steel was brought in via the aforementioned Albert Canal to create the Campus Bridge. This dramatically reduced conflicts of the intersection for pedestrians and cyclists and should find a connection with the F cycling highways nearby—the road and infrastructure up to the bridge is currently being rebuilt. The bridge looks more impressive from a certain height as explained on the agency for roads & traffic website . Even though in the photo below it looks like a classic bridge spanning at most 50 metres to cover five car lanes, you’ll be cycling for almost half a kilometre (410 metres) to cross it! The Campus Bridge connecting two sides of a busy intersection in an area with many schools. The photo was taken from the right side where you can take a big ramp up in two parts including a U-turn to reach the top of the bridge, only to cycle back down in a circular style on the far left, reaching the campus buildings of the business department. Fortunately, pedestrians can take a shortcut via two connected staircases in-between the turns. There are even cosy looking benches placed halfway towards the top to enjoy your lunch on a sunny day. Related topics: / bike / By Wouter Groeneveld on 17 September 2026.  Reply via email .

0 views

An Interview with Joanna Stern About the iPhone Duo and AI for Normal People

An Interview with Joanna Stern About the iPhone Duo and AI for Normal People

0 views

Companies don’t make decisions

I feel kinda stupid for even saying this, but it’s probably worth saying anyway: companies are not people. Companies don’t make decisions. Companies don’t decide who they want to donate money to. Companies don’t feel shame. Companies don’t learn from their mistakes. Companies don’t have political orientation. Because companies are legal entities, they're a bunch of papers, letters and numbers. That’s what companies are. When you’re mad at a company, figure out who is actually responsible for the things you’re mad about, and be mad at them. Because people can be held accountable. There’s a reason why, in 1789, it was people’s heads that were rolling on the ground. Thank you for keeping RSS alive. You're awesome. Connect via email :: Sign my guestbook :: Support for 1$/month

0 views
Jason Tucker Yesterday

Inside Quartermaster: A Conversation with Lewis Egley

I've been running Quartermaster on my phone for months now, watching it grow from a solid Sonarr and Radarr companion into something that's trying to handle my entire homelab in one app. When I reviewed it back in the spring , it was already doing more than most apps in this space attempt. Since then, developer Lewis Egley has shipped an enormous amount of work, and the app is now getting a Docker-based companion tool that changes what it's capable of entirely. I sat down with Lewis, virtually and asynchronously (Lewis as you will find is British and I'm located in California so there is an 8 hour time difference and Lewis doesn't sleep much) over Discord, to talk about where Quartermaster came from, what Companion actually does under the hood, and where he sees this whole thing heading. He was candid about the late nights, the pressure of shipping fast in a space that's gotten suspicious of anyone who does, and what it's like going from a private side project to something with a real user base watching every move. If you want to follow along or dig into any of this yourself, the app is on the App Store and the GitHub repo for Companion is open source under MIT. There's an active community over on Reddit and in the Discord server, which is honestly the best place to catch build announcements before anyone else does. Walk me through your path in tech, how'd you go from where you started to building Quartermaster? It's a very good question and sort of cliché if you will :) The world of computers has always interested me, ever since I was young. I started making little programs as a child using Python, and weird little websites I'd show my friends. At the time, people my age didn't understand, but I did, and that was enough. In regards to my professional career, I started as any other young and hungry tech advocate. I got my study on! I majored in Computer Science and got my diploma, and then I started out in the working world. My first ever job was IT 1st line, a helpdesk position. Looking back, I hated it, but it's the only way in. A true rite of passage for anyone who wants to start. Over the years I started picking up skills and moved quickly into cloud infrastructure, building out cloud environments for businesses up and down the UK. That opened the door to DevOps and ultimately software engineering. I built Quartermaster as a personal project and ran it locally for a few months. Then I was speaking with a few guys on a tech forum and they really wanted to test it. They gave me improvement ideas, I built them, and so on and so on. The app launched on the 3rd of July and I haven't looked back since, just a lot of late nights, haha. Is Quartermaster your full-time thing now, or are you building it around a day job? Believe it or not, Quartermaster is not my full-time job. It is, and probably always will be, a side project. I love my full-time job, the companies I get to work with and the people I meet. It has taught me everything I know and I just don't think I'm ready to turn my back on it just yet. What pulled you into the homelab space specifically? Are you running a big stack yourself, and if so what's in it? One word. Community. The ideas people come up with, and everything in between, is honestly captivating. It's odd, because it's essentially its own ecosystem, and once you're in it, you're in it. I've been homelabbing for around three years and every day I'm amazed by what people do. No matter where you look there's something out there for everyone, a dashboard here, a service there. There just isn't anything quite like it. The stack I run is small compared to some. I have a soft spot for networking, so it's Unifi throughout, switches, CCTV, the gateway. For media I run a 16GB ECC box with a 12TB array, which covers around 60% of what Quartermaster supports. Obviously I have to spin the rest up locally for proper testing. Any other apps, side projects, or open-source work people might know you from? No, lol. Long-time lurker, first time going public. I'm a very honest guy who keeps himself to himself, so opening the floodgates on a project like QM as a solo dev has been a very big learning curve. Your day job took you through cloud infrastructure and DevOps before software engineering. How much of Companion's design, the proxy layer, the tiered access, the whole approach, came out of instincts from that world rather than from building consumer apps? Almost all of it, if I'm honest. The consumer app side I had to learn as I went, but the Companion architecture came out of habits I already had. The clearest example is the Docker access model. A consumer app instinct would be a settings toggle, switch on Docker management, done. What I built instead is a ceiling you set at deploy time. The Compose profile you start with decides the maximum access the process can ever have, and every profile boots into read only regardless. The owner can raise the active mode from inside the panel, but only up to the ceiling they already committed to on the host. You cannot escalate past it from the UI, because the UI was never given the authority in the first place. That's just least privilege, and it's the sort of thing you internalise building cloud infra for other people's businesses. Nobody gets prod write access because a toggle in a dashboard said so. The socket proxy is the same idea. Companion doesn't get the raw Docker socket, it gets a keyed proxy in front of it with its own scoped surface. The other bit that came from that world is being blunt in the docs about what the thing can actually see. Docker inspection exposes container metadata, mounts, network layout and anything sitting in environment variables. A read only mount still means Companion can read that file. That is in the README because someone running this at home deserves the same honesty I'd give a client during a design review. Do you see Quartermaster becoming a full home lab platform in its own right, or is Companion still in service of the mobile app? I don't know yet in all honesty, and I'd rather say that than invent a roadmap. Companion came about because I hit the ceiling of what a phone can do on its own. Some things need something running on the box with real permissions, and no amount of clever iOS code gets you round that. So it started very much in service of the app, and it does that job well. Where it goes from here is the exciting part, because it isn't entirely my call anymore. Open sourcing it under MIT means people are already running it for things I never designed it for, and that's brilliant to watch. It's the whole reason I love this space. You put something out and the community immediately finds uses you'd never have thought of on your own. I'm still careful with the word platform, mind. It carries promises about support and longevity, and I want to earn those rather than claim them. But if it keeps growing the way it has, and I can keep it secure and well maintained, then I'd love to see where it ends up. Right now I'm just enjoying building it. You clearly thought hard about the trust boundary here, the proxy, the hard ceiling, the tamper-evident config. Was there a specific moment or scare that shaped that approach, or was it just discipline from day one? Discipline, mostly, though it came from a very specific realisation rather than a scare. When I sat down and properly thought about what I was asking people to install, it focused things fast. Companion holds service credentials and can be granted Docker control on the host. Get the owner session and you don't have a media dashboard, you've got someone's box. Once that landed, a lot of the design decided itself. That's why the setup transfer is short lived and single use. It's why pairing makes you compare words on both screens before you approve anything. It's why the default install is read only, and anything beyond that has to be a deliberate choice you make at the host level rather than a toggle in an app. The bit I'm most pleased with is that you don't have to take my word for any of it. There's a verification script in the repo that runs the boundaries against a live instance. If you're going to ask people to trust something running on their own hardware, they should be able to check rather than just believe you. Homelabbers are all over the map, some just want Home Assistant, Unraid for media, and Immich for the family photos. Others are running full racks with oddball switch configs, learning Proxmox, juggling Tailscale, Twingate, ZeroTier, Netbird, Netmaker. That's a massive range of technical comfort under one roof. How do you think about serving both ends of that spectrum with Quartermaster, and does that split change how you approach onboarding and documentation? The way I've approached it is that the complexity lives in their setup, not in mine. I don't try to guess how someone's network is put together. I just give them enough ways to describe it that whatever they've built, they can point QM at it. So you get a local address, an optional remote one, a home WiFi pin so it forces local when you're in, reverse proxy domains, Cloudflare Access header sets, custom headers, basic auth, and certificate pinning for self signed setups. Tailscale isn't a special mode, it's just an address, because to the app it should be. Someone on a Synology with three apps uses one field of that. Someone with a rack and five overlay networks uses most of them. Same app, same flow. Onboarding does branch, though not by skill level. It branches by how you're arriving. There's a manual path, a Companion path where the server exports your whole stack and the phone imports it in one go, a demo mode, and a restore path. Then you pick your services off a list, and it walks you through connecting them one at a time, most depended on first. There's very little discovery, which is deliberate. Where I'd say I've genuinely got it right is what happens when a connection fails. The Test is the moment everything hinges on, so that's where the work went. Four stages, real error codes, the server's own refusal text surfaced verbatim rather than swallowed, and one tap fixes for the specific traps, certificate trust, SSH host keys, Synology OTP, a port left prefilled on a public domain. You can't save a service that hasn't passed a Test, so nothing sits there claiming to be connected when it isn't. You mentioned there's very little service discovery by design, and that onboarding branches by how someone arrives rather than their skill level. That's a strong opinion, a lot of apps in this space lean hard into auto-discovery as a selling point. What made you bet against that? Mostly that on iOS it doesn't work as well as the marketing suggests, and the failure mode is brutal. The Local Network permission is a one hot prompt. You get it once, and if the user denies it there is no API to ask again. They have to go into Settings and find it themselves. So if that prompt fires during a background dashboard load rather than something the user deliberately did, the request races the prompt, fails as a generic "couldn't connect", and you've burned your one chance. The app actually fires a throwaway probe when onboarding opens specifically to get that prompt out of the way at a moment that makes sense to the user. Then there's the yield. Bonjour browsing on iOS needs you to declare every service type up front in the app's config, and most selfhosted software doesn't advertise over mDNS anyway. Radarr doesn't. Sonarr doesn't. SABnzbd doesn't. So you'd be building it for almost nothing. Subnet scanning is worse. Hundreds of connection attempts, straight into iOS's six per host limit, and a conversation with App Review about why your media app is port scanning. And a big chunk of these users are on Tailscale or behind CGNAT, where sweeping the local subnet finds nothing that matters. So I moved discovery to where it can actually work, which is a box that already has Docker. Companion reads the container list, works out what's running, and reads the API keys straight out of the config files on disk. The *arrs keep theirs in config.xml, SAB in sabnzbd.ini, and so on. The user types nothing. That's real discovery, it just isn't happening on the phone. Plex is the other exception, and it's the same principle. It's account side, not network side. You sign in and Plex tells you every server on your account with every route, so the app picks the best one. Where a proper API exists I'll use it. I'm just not going to guess at somebody's subnet. A lot of the competition in this space, LunaSea, Zagreus, Ruddarr, Unraid Deck, are unitaskers, they do one thing really well. You told me you're not using AI to build any of this, it's all hand-written. That's a lot of surface area to cover solo without leaning on tooling to move faster. What made you go all-in on the everything-app bet instead of picking a lane and doing it well? Because the unitasker advantage disappears the moment you run more than one thing, and almost everyone in this space runs more than one thing. The bet wasn't really "cover everything". It was that the hard part isn't the services, it's the layer underneath them. Getting a connection to somebody's box is genuinely difficult. Dual local and remote addresses that switch on their own, self signed certificates on iOS, reverse proxies, Cloudflare Access, custom headers, per host concurrency limits, a transport that won't follow a redirect and replay your API key somewhere you never configured. That's about thirteen thousand lines of plumbing, and I only had to write it once. Everything after that is a service riding on top of it. A new integration inherits the URL assembly, the auth handling, the certificate pinning, the reachability engine, the four stage connection test, the add form, the backup and restore, all of it. I won't pretend it's cheap though. A new service is still six to eight thousand lines by the time you've done the client, the screens, the mappers and the demo data (needed for Apple review). There's no plugin system, no manifest, nothing schema driven. It's all compiled in. That's linear cost forever and I knew that going in. What I got for it is that each of the 59 services has a real screen and real actions rather than a generic card. The schema driven approach would have let me claim two hundred services that all look identical and none of which actually does anything. I'd rather have 59 that work. Unraid Deck does almost exactly what Companion does, direct Docker control, SSH, local push notifications, no hosted relay, but it stays narrowly focused on just Unraid. Companion's trying to sit underneath a much bigger app that spans dozens of services. What's the trade-off there, and do you think staying narrow like that is actually the smarter long term bet? Narrow is smarter if you're solving one platform's problem. I'd not argue with that at all, and I think Unraid Deck is doing the right thing for what it's for. Companion isn't Unraid specific. It's a Node app talking to the Docker API, so it works against Proxmox, Synology, bare metal, whatever. Unraid is just the best documented install path because that's where a lot of people are. But that generality has a cost. Deck can make assumptions about paths, about the array, about the whole shape of the system. I can't make any of those, so I end up handling Linux host paths generically, Synology's volume layout, socket proxying instead of a raw socket mount, and so on. The trade off is what Companion is for. It isn't primarily a Docker manager, it's the thing that makes setup disappear. It reads your config files, finds your keys, works out you're running three Radarrs on different ports, and hands the phone an encrypted bundle so you type nothing. That job only exists because there's a big app sitting on top of it that needs 59 services configured. A narrow tool doesn't have that problem to solve, which is exactly why it can stay narrow. How do you decide what to actually support in the mobile app? I noticed Companion now has a marketplace, is that the mechanism for covering more of these different setups without you having to hand-build support for every single one in the mobile app and just let Companion handle it? It's a catalogue generated from Companion's/QMs own supported list of what it knows about. There are exactly three hand reviewed Compose starters in there. Everything else is either generated from a template or connect only. And no, it isn't a way round hand building support. I wish it were, but I'd be lying if I said otherwise. The app's list of service types is compiled in. It's a fixed union in the source, it's a key in dozens of exhaustive lookups, and it's the discriminant the app stores things under. There's a hard boundary on import that drops anything not in that list, with a test asserting it does. So if Companion learns about something new tomorrow, it can detect it, deploy it and show it in the web UI, but the phone will not treat it as a service. No card or even screen and certainly no status probe. That's a build time thing and no amount of Companion architecture changes it. There's one escape hatch, and it's honest about it all, the Companion control plane in the app shows containers, stacks, images and events generically, without knowing what any of them are. So something new turns up there the moment you deploy it. But it turns up as a container, not as an integration. How I actually decide what to support is much less interesting than that. It's what people ask for in the Discord, weighted by whether I can get a real instance running to test against. That last bit matters more than people expect. I pulled up the actual Marketplace in Companion. Right now every one of the built-in entries is first-party, curated by you, and the catalog banner says every deployment still starts with a review. But there's already a Sources tab sitting there, empty at zero right now, that lets people point Companion at external Portainer v2 template feeds, and it says outright those entries aren't reviewed by you. That's a real two-tier trust model already built and waiting. Was that the plan from day one, community sources living alongside your own with clearly different guarantees, and how worried are you about someone pointing that at something malicious? I wouldn't say the exact shape of it was mapped out from day one, but the separation was deliberate once the Marketplace started taking shape. There had to be a hard distinction between something included by me and something a user has added themselves. If it appears in the built-in catalogue, I'm putting my name against it. I've looked at it, I know where the image comes from, I know what it mounts and I know what permissions it asks for. An external source cannot inherit that trust just because Companion happens to be displaying it. And yes, a malicious feed is a real risk. A Compose template can ask for privileged mode, mount the Docker socket, mount the host filesystem or run an image that does something completely different from what its description claims. There is no honest way for Companion to turn arbitrary third-party Compose into something safe automatically. That's why external entries are labelled differently and why deployment still starts with a review. Companion shows you what is actually going to be created before anything runs. It doesn't quietly install something because it appeared in a feed. But I don't want to pretend that removes the responsibility from the person deploying it. If you add an untrusted source and then approve a template with access to your host, Companion can't make that decision harmless. The aim is to make the trust boundary obvious and make dangerous requests visible, not to claim I can sanitise arbitrary code from the internet. You mentioned there are only three hand-reviewed Compose starters in the Marketplace right now, everything else is generated or connect-only, and even if Companion learns about something new, it can't become a real integration in the phone app because the service list is compiled in. So what would "other devs contributing to the Marketplace" actually look like in practice, more Compose starters and Sources feeds on the Companion side, or is there a future where the compiled-in service list itself opens up somehow? Right now it means the Companion side. Other developers can publish Portainer feeds, build Compose starters and help Companion understand how something is deployed or detected. If something is going into the first party catalogue, I still need to review it before it gets that guarantee. Alternatively, somebody can maintain their own source and users can choose to add it with the external-source warning attached. What it does not mean today is dropping a manifest into Companion and suddenly gaining a full Quartermaster integration on the phone. The phone app is deliberately much stricter than that. A service type is compiled into the app and tied into the connection layer, storage, backup and restore, onboarding, demo data and native screens. Opening that at runtime would either require an enormous plugin system or reduce integrations to generic cards driven by a schema. Neither gives me what I want from Quartermaster. So in the near term, community contribution means deployments, templates, detection and knowledge on the Companion side. For the app itself, people can request integrations, help me get a test instance running, document strange API behaviour and test the result, but the actual implementation still has to become part of the app and go through a release. I won't say the service system can never open up, but there isn't a hidden plugin SDK around the corner. If I ever do it, it has to produce something that feels like a real Quartermaster integration rather than a web form wearing an iOS skin. A while back a security researcher found serious auth bypass vulnerabilities in Huntarr , and the maintainer banned people who raised concerns about it. It's become the go-to cautionary tale people point to for vibe-coded apps cutting corners. Has that story stuck with you? Does having something like that out there change how you think about your own pace, or your own security process? It's stuck with me in the sense that I've watched how the community reacted to it, and that told me something useful. People in this space care about this stuff. They will look, they will test, and they'll say something when they find it. What it reinforced more than anything is how you handle it when someone does come to you. If a researcher turns up at my door tomorrow with a genuine finding, that's someone doing me an enormous favour for free. The right response is thanks, then a fix, then telling everyone what happened. Closing the door on that is how a small problem becomes a permanent reputation. On pace, it hasn't changed it exactly, but it's part of why the security model went in early rather than later. It's much harder to bolt a trust boundary onto something after people are already running it. Do it up front and you're building on top of it rather than retrofitting round it. There's a lot of noise right now about "vibe coded" apps, and Huntarr's become the example everyone points to, real security holes, a maintainer who didn't seem to understand what he was shipping. Where do you actually sit on that? Do you use AI tools in your workflow, and if so, how do you keep the line between using them well and leaning on them too hard? I sit somewhere that probably annoys both camps. I don't use AI to build Quartermaster. Not the app, nor the Companion. Where I do use it is checking a GitHub URL for a new integration in case an API's changed and I'd otherwise miss it. That's a research task and it's genuinely useful for that. But I think the "vibe coded" label gets used as a lazy shorthand, and it points at the wrong thing. The tool isn't the problem. You can write every line by hand and still ship something dangerous if you've never had to think about a trust boundary. Equally you could use AI sensibly, review properly, and be completely fine. The real question is whether you understand what you're shipping. That's it. Can you explain what you decided, what it exposes, and what happens if someone attacks it? If you can't answer that, it doesn't matter who or what wrote the code. I'll use it to research and to check my blind spots. I won't use it to make decisions I'd struggle to defend afterwards. I've watched you build this out in real time, and you've even put a verification script in the repo so people can check the security boundaries themselves instead of just taking your word for it, which is the exact opposite of what happened with Huntarr. But I know "vibe coding" accusations get thrown around fast, especially with that story still fresh in people's minds. Does being anywhere near that comparison worry you? Has anyone actually come at you with that kind of accusation, and how do you handle it when the work itself is the real answer? It doesn't worry me, no. If someone wants to lob that at me they're welcome to, but they'll be arguing with a repo that's sat there in the open. The security model was designed deliberately, because I've spent years doing exactly this kind of work for businesses and I know what happens when you get it wrong. Least privilege, a ceiling you set at deploy time, read only by default, a proxy in front of the socket rather than handing over the socket itself. None of that is accidental and I can explain every bit of it. But the strongest answer isn't me telling anyone that. It's that I'd rather nobody took my word for it in the first place. That's the entire reason the verification script exists. Run it against a live instance and it tests the setup and authentication boundaries for you. If you think I've cut corners, that's the quickest way to prove it. The repo's open, the app's on the store. Have at it. You've shared some pretty stark numbers on what the hosted notification relay was actually costing you to run. What made you consider dropping it entirely, and how did you land on moving notifications into Companion as a local relay instead? Did that same cost reality shape the new pricing structure for 1.2, and how did you balance that against keeping your lifetime purchasers taken care of? The relay costs got out of hand, plainly. We're talking millions of requests a day through Cloudflare, and that's not sustainable for one person funding it out of pocket. The fix is Companion. Notifications are moving to a local relay, which means people will be able to set their own alerts and run the whole thing on their own box rather than everything routing through infrastructure I'm paying for. That's better architecturally as well as financially, which is the nice version of a cost problem. It's also more in keeping with the whole point of self-hosting. Your stuff, your box. The 1.2 price change is partly those notification costs and partly everything else that's landed since launch. Companion, the custom built macOS version, all of it. I've kept the increase as small as I can and it'll still come in cheaper than the alternatives. The important bit: every current subscription stays at the price you're on until you cancel, and everyone who bought lifetime keeps lifetime. That was never up for discussion. People backed this thing early, before it had proved anything, and they're not getting punished for it. You mentioned QM is a side project on top of a full-time job, and you shipped an enormous amount of work in eight weeks. What does a week actually look like for you right now, and how close have you come to burning out on it? Honestly? Weekends are soaked. During the week it's evenings after work, and more often than not I'm going to bed around four in the morning. Then up for the day job. That's just been the rhythm since launch. My partner hasn't always been thrilled about it, and I don't blame her. But when I show her what's actually happening, the numbers, the people using it, the messages coming in, she's amazed. That helps. It's my passion, and having someone who gets that even when it's eating the weekend makes a big difference. As for burnout, it's real and I won't pretend otherwise. I make mistakes when I'm running on not much sleep. Things slip through that wouldn't if I'd been fresh. That's genuinely where the closed beta community has been unreal. They catch things and they tell me straight when something's broken. When you're a solo dev at two in the morning, having a group of people who care about the project as much as you do is worth more than any amount of extra hours I could put in. I'd have shipped a much worse app without them. You said this is the first project you've taken public after being a long-time lurker. What's that jump been like, going from something you built quietly for yourself to a real app with a user base and an active Discord community? Terrifying, then brilliant, roughly in that order. Building something quietly for yourself is safe. Nobody sees the bad decisions, nobody files a bug at midnight, and if it breaks it only breaks for you. Putting it out means every choice you made alone is suddenly up for discussion by people who know what they're talking about. That first week I checked everything constantly. What I didn't expect was how generous people would be. I'd braced for criticism and what I got was people wanting it to work. Detailed bug reports, feature ideas better than mine, folk explaining their setups so I could understand why something wasn't behaving. The Discord has genuinely made the app better in ways I couldn't have on my own. The other thing is that it's changed how I build. When it was just mine I could cut corners and know where the bodies were buried. Now there are people relying on it, and that raises the bar on everything. Not in a stressful way particularly, more that it makes you take it seriously. I'm still a fairly private person, so going from lurker to having a community around something I made has been a big adjustment. But I'd not go back. Whatever I thought I'd get out of shipping this, the people have been the best part of it by a distance. I'd just add a closing note on that. Quartermaster is ultimately just code running on people's devices. What actually drives the late nights, the new features and everything in between is the community. They keep me going and I couldn't ask for a better one. There's a stereotype that British people apologize for everything, even when it's not their fault, always "ever so sorry" or "terribly sorry" for the smallest things. You've said sorry in Discord more than once for delays on something you're building for free in your spare time. Is that just very British of you, or do you genuinely feel like you owe people an apology? How do you deal with the pressure of a deadline you set for yourself versus one the community's expecting, and which one do you think has more pressure? Bit of both, honestly. It's very British, lol. When someone clicks purchase, or takes time out of their day to message me about a problem, I think they deserve the best service I can give them. That's not me being self flagellating about it, it's just how I think it should work. They've given me something, whether that's their money or their time, and a delay means I've not held up my end. As for which deadline has more pressure, it's mine by a distance. The community are generous. They're patient, they tell me not to worry about it, and nobody in the Discord has ever given me grief over a slipped date. The pressure is entirely internal. I set a date, I want to hit it, and when I don't, that's on me. Which is probably exactly why I keep apologising for things nobody's actually asking me to apologise for. You mentioned "the custom built macOS version" when talking about what's driven the 1.2 pricing changes, so it sounds like Mac already got real dedicated attention rather than just running the iPad build. What does "custom built" actually mean there, is it a fully native AppKit app, or an iPad app with Mac-specific work layered on? And what about Android or Windows, are those actually on the table, or does it make more sense to lean into a solid web app for anyone outside the Apple ecosystem instead of chasing every platform natively? It's the iPad app running on Apple silicon through "Designed for iPad", with a significant amount of Mac specific work layered on top. It's not Catalyst and it's not AppKit. Shipping it is a checkbox. Making it feel like a Mac app was not. There's a whole separate shell for it, a sidebar, a today column, a desk bar, sat behind a flag with a kill switch. Detecting you're even on a Mac needed a small native module, because React Native reports Catalyst as false in that mode. Then there's window minimum sizing, hardware keyboard shortcuts, and disabling keyboard avoidance in three places because the system still fires keyboard events for the input accessory strip with a hardware keyboard, and the maths for a movable window is completely wrong. There's a good story in there about boring choices too. The sidebar was originally a real UISplitViewController via an experimental router feature. It misbehaved three times in one day during beta (as you know), so I ripped it out and replaced it with a plain flex row. It's less clever and it works. Android and Windows, honestly, not soon. There's more Android scaffolding than you'd guess. Five of the twelve native modules already have real Kotlin implementations. But there are over a thousand icon call sites using SF Symbols with no Android fallback, seven native modules with no Android side at all, and the whole visual language is built on iOS 26 glass. The business logic would port cleanly, because the connection layer and the data mapping are deliberately framework free and tested under plain Node. The presentation layer would be a rewrite. I might even need to bring on a dev to help me, the project has grown that large. 1.2 shipped later than planned, but the delay was actually Apple's review process, not you. What's it like waiting on something completely out of your hands after you've done everything you can on your end? Honestly, worse than when the delay is mine. If something is broken, I can fix it. If testing turns something up, I can stay up late and deal with it. Once the build is submitted and sitting with Apple, there is nothing useful left for me to do. You end up refreshing App Store Connect even though you know perfectly well that refreshing it isn't going to move the queue. The frustrating part was that everybody could see me saying 1.2 was ready, but nobody could actually download it. From the outside, a delay is still a delay, even when the finished build is just waiting for somebody else to press the button. Apple review is part of shipping an iOS app, so ultimately it is still something I have to allow for. The lesson for me is probably to stop treating submission day as release day when I talk about dates publicly. I can control when I finish a build. I cannot control exactly when Apple lets it out. You just announced QM Reader for 1.3, a full request/read/listen experience for books. That's a real expansion beyond media and homelab management. What made now the right time to go there, and how does it relate to the book and audio tools already in the Marketplace, Audiobookshelf, Kavita, Shelfarr? It looks like a jump from the outside, but from inside the app it feels like the next missing piece. Quartermaster already understands the different parts of that stack. Shelfarr is about finding and requesting something. Kavita is about the library and reading side. Audiobookshelf handles audiobooks and playback. They are good tools, but today those experiences still live next to each other rather than feeling like one journey. QM Reader is the layer that joins them together. The idea is that you should be able to find a book, request it through the service you already run, see when it arrives, then read or listen without the experience suddenly falling apart into three unrelated integrations. It isn't about replacing Audiobookshelf, Kavita or Shelfarr. They remain the servers and the source of truth. Quartermaster is the native experience across them. Now feels like the right time because the underlying work is finally there. The connection system, multiple instances, secure credential storage, downloads, media handling and Companion discovery have all already been solved in other parts of the app. I can build Reader on top of those foundations instead of creating another isolated app and asking people to configure the same stack twice. It is a bigger piece of work than adding another dashboard, definitely, but it also feels like the first time Quartermaster can turn several integrations into one complete feature rather than simply putting them beside each other. With Pulsarr, PeaNUT, Uptime Kuma, QM Reader, and a pricing review all confirmed for 1.3, that's a lot on the plate right after shipping the biggest release yet. How are you pacing yourself into this next stretch instead of just running the same sprint again? By trying not to pretend all five of those things are the same size. Pulsarr, PeaNUT and Uptime Kuma are contained integrations. They still need proper clients, screens, tests and demo data, but they inherit the connection and onboarding work that already exists. QM Reader is the large piece. The immediate priority after 1.2 is stability. There is no point announcing the biggest release I've done and then sprinting past every bug people find because I'm already chasing 1.3. I'll deal with the real-world feedback first, then build the next pieces behind TestFlight rather than trying to land everything at once internally. I'm also being more careful about dates. I am very good at turning my own estimate into a deadline and then treating that deadline like somebody else imposed it on me. Nobody in the community is asking me to work like that. That's entirely self-inflicted. So the plan is to keep moving, because I'm obviously not very good at sitting still, but to sequence it properly. Stabilise 1.2, land the smaller integrations, give Reader the time it actually needs and only call 1.3 ready when it is ready. Whether I manage to behave that sensibly for the entire release is a separate question. What are you most excited for people to actually experience once this update is in their hands? The moment Companion makes the setup disappear. People have heard me list service counts and talk about Docker discovery, but the number isn't the interesting bit. The interesting bit is scanning once and watching the app fill itself with the services you actually run, including their addresses and credentials, without spending half an hour copying API keys out of config files. That's the point where Companion stops being another thing you have to manage and starts earning its place. I'm also excited for people to see that 1.2 isn't just a collection of integrations. The app, the Mac experience and Companion now feel like parts of the same system. You can still use Quartermaster directly without Companion, exactly as before, but if you do run Companion the whole thing becomes much more joined up. I've been looking at individual screens and broken beta builds for so long that I'm probably too close to it now. I want to see what it feels like when somebody encounters the complete flow for the first time. Is there anything I haven't asked that you think people should know, or that you've been wanting to say? Probably that Companion being here does not change the original promise of Quartermaster. It is optional. The phone can still connect directly to your services. There is still no Quartermaster account, no analytics profile and no requirement to send your homelab through infrastructure I control. Companion is there when running something locally makes the experience better, not to drag a self-hosted app towards somebody else's Cloud. The other thing is that beta genuinely means beta. I've tried to be clear about where the boundaries are, what has been verified and what still needs real-world testing. I would rather tell somebody a feature has a limitation than quietly let them discover it after trusting it with their server. And, genuinely, the amount of useful feedback people have given me has changed the app. A lot of what is in 1.2 exists because somebody joined the Discord , explained their strange setup and then patiently tested three builds while I got it right. It might be one person writing it, but it has not been built in isolation. Where should people go to follow along, report bugs, or get involved, Discord, GitHub, somewhere else? Discord is the main place. That's where I post development updates, TestFlight builds, feature previews and the occasional message apologising for a delay nobody has complained about. It is also the quickest place to report a Quartermaster bug, request an integration or show me a setup that behaves differently from mine. GitHub is the better place for Companion issues that need logs, reproduction steps or a technical discussion around the implementation. The Companion repo is open, so people can inspect it, run the verification tooling and raise something properly if they find a problem. For security issues, please report them privately first rather than dropping exploit details into a public channel. I will take them seriously, confirm what is affected and be open about the fix. The website and App Store have the polished version of what Quartermaster is. Discord and GitHub are where you see how it is actually being built. What stuck with me most out of all of this wasn't the security model or the thirteen thousand lines of connection code, it was watching someone build in the open and mean it. Lewis is still answering bug reports in Discord at hours most people are asleep, still apologizing for things nobody's upset about, and still building toward something bigger than what fits in one app. I'll be watching what QM Reader turns into, and I'd bet a lot of you reading this already have Quartermaster running somewhere in your stack. If you don't yet, now's a genuinely good time to start. I'd love to hear in the comments what you think about Quartermaster, how you are using it and how friendly we are in the Discord there.

0 views

Love for tiny utility programs

These days, while I’m having more and more difficulties appreciating software in general, I’m getting a lot more joy and delight from tiny apps and little utility programs that offer very few features. In the past couple of days alone, I’ve discovered — or rediscovered in some cases — a few of these that I thought would be worth your time. The first one is History Book , from Zhenyi Ta n (of And a Dinosaur’s fame) that I use as my new read-later app. The twist is that it is not really a read-later app. The app’s purpose is actually to be an alternative to Safari’s browsing history. It vastly improves on it by making the content of the pages visited searchable as well, and not only the title of the page as it is mostly the case with browsers’ browsing history logs. By default, the app will save only pages where an article is found, but it can also be set up to save all pages automatically, whether after a few seconds or after a waiting period. There is a third way to save URLs in this app: by clicking a button the old-fashioned way. This is what I do, and, despite a few bugs, I adopted the app very gladly. You see, I love an app like GoodLinks , but somehow, I find it quite difficult to manage, and maybe over-complicated for my humble needs? Between the read/unread status and the different tag folders, I end up being lost on what I’ve read, what I want to keep, what I need to archive, and what I can delete. In History Book there are just folders, and that’s it. Barely any other setting is available, and I love this purposefulness in an app. Also, and you must have seen this coming if you’re a regular reader , it’s superfast and light (only 2.9 MB). I just wish it had a few comfort updates: Great icon, though. The second app I want to recommend is also from Zhenyi Tan. If this app is harder to recommend, even harder to find a good use case for it, Lucky Notes is probably my favourite app of the year. For starters, this has to be the cutest app for the Apple ecosystem: the icon alone makes it worth keeping the app in the dock, but the overall vintage Mac look is delightful. Zhenyi Tan describes Lucky Notes as an app to save text snippets, and the app is actually a companion app to the Lucky Safari extension , which greatly improves Google Search result pages. Indeed, notes in Lucky Notes are only searchable via the companion extension. The app on its own is more than bare-bones, it’s lacking a lot of basics, even a settings window. But I don’t seem to mind, that’s the crazy part. I’ve spent the last couple of days trying to find a reason to add this app into my setup because it really is an app crush: I want to use it. I’m currently using it as a separate note collection from the ones I store on my Mac via BBEdit, and the ones I keep in Apple Notes as an easy way to have them shareable and synced between my devices. If this app had the following features, I’d consider committing to put all my notes in this: Hard to write a post about tiny utility programs without mentioning one of Sindre Sorhus’s apps ; the third and final app I want to recommend is Plain Text Editor . *2 I’ve recently added this editor to my Mac setup because if there is one thing that’s missing from BBEdit , it’s a proper full-screen distraction-free mode. Plain Text Editor provides me with that, and, most importantly, just that. I’ve liked this app for a while now, but if it’s limited as a main writing app, it’s brilliant as a companion app for first drafts. I used it to draft this very post. The only thing that’s missing is a line-height setting. The app could also use a few fixes when it comes to tiny visual bugs. That’s it: I think any extra feature more would make this program too complex and probably would disqualify it for this “tiny app” category. These three apps, History Book, Lucky Notes, and Plain Text Editor made me realise that I have a soft spot for simple, bare-bones apps. Piezo , Yoink , Tot also come to mind. What I like the most in these apps is the minimal number of things to configure. Coming from a new BBEdit aficionado , this may sound strange, but this feels like plug and play, for apps. Download, launch, enjoy. Despite my strong enjoyment of digging into settings, this is very refreshing and somehow reassuring: reinstalling each of these apps on a new computer is a breeze. What about you, dear reader? I’m very curious to know which tiny little program you’re found of, which are the ones you recommend and why. Please let me know which is your latest tiny app crush. It’s worth noting that in the past couple of days alone History Box received a few updates fixing a lot of my initial issues, forcing me to gladly update this draft a few times. By the time you read this, it’s possible all the things I listed will be covered.  ^ Among the other numerous Sindre Sorhus apps, I like Aiko : I purchased it a while ago to use on my work computer to replace MacWhisper (a bit too “power user” for me), and it’s brilliantly efficient and simple. I also recommend Hyperduck .  ^ The ability to assign a specific folder to the extension button, separate from the background history saving if enabled Extra metadata on saved articles (not just the time it was saved, but the publication date for instance, author, reading time, etc.) A proper export system, just in case. The current solution documented is a bit messy Extra Mac functionalities like the ability to hide the toolbar, a status bar to preview URLs, better default keyboard shortcut support like Cmd + F for search, etc. *1 a search box without the need for the Lucky extension a text size/line length setting a better resolution system for syncing conflicting versions of a note an option to only use windows instead of tabs (which would feel even more classic Mac-ish). a proper export/backup system. It’s worth noting that in the past couple of days alone History Box received a few updates fixing a lot of my initial issues, forcing me to gladly update this draft a few times. By the time you read this, it’s possible all the things I listed will be covered.  ^ Among the other numerous Sindre Sorhus apps, I like Aiko : I purchased it a while ago to use on my work computer to replace MacWhisper (a bit too “power user” for me), and it’s brilliantly efficient and simple. I also recommend Hyperduck .  ^

0 views
Martin Fowler Yesterday

Fragments: September 16

Reports of agentic hacking continue, in this case it happened back in May and it seems OpenAI did not disclose that they were responsible. Simon Willison sees two options: Both of these are bad! Given this incident, the Hugging Face situation, and the Wiki attack, the obvious question right now is how many more incidents like this are out there waiting to be discovered? ❄                ❄                ❄                ❄                ❄ Dave Farley: Stop asking the sci-fi question: ‘Is it conscious?’ Start asking the engineering question: ‘Is this a powerful, unpredictable component being put somewhere consequential, and where’s the feedback that tells us that it’s safe? ❄                ❄                ❄                ❄                ❄ Nate Silver is known for his forecasts, but to do them he writes a lot of code for his models. He’s found agentic programming capable of doing miraculous work . In spending so much time with the LLMs, I’m super attentive to improvements in their capabilities. And these changes tend not to be so linear. Instead, they improve in step functions, almost as phase changes. Suddenly, the models just start doing things capably that they were screwing up before. In my experience, there was a big leap forward when reasoning models first came out in late 2024/early 2025 — enough that they were occasionally useful for tasks involving data and not just words — and then another one this past winter. The most recent changes I’ve noticed, however, have had less to do with intelligence and more with persistence. Consider the Hugging Face attack. Although these agents showed remarkable intelligence, they weren’t really super-intelligent - but they were super-persistent. This is a common theme of AI in its various forms: Game engines like AlphaGo Zero start out by basically making random moves — but by playing against themselves millions of times, they eventually far surpass human capabilities As we try to figure out what kind of regulations we need to keep AI under control, we need to remember that we should design our guards around super-persistence as much as worrying about super-intelligence. ❄                ❄                ❄                ❄                ❄ “Uncle Bob” Martin has made many posts on X during the last few months about his programming with LLMs. His approach has been to build a firm harness to keep them under control, so they create software that is maintainable as well as functional. Sadly the posts have been frustratingly light on detail. But now it seems that lack of information may not matter And while I was heads-down getting that to work, the agents got a LOT better. So much so that when I came up for air, the need for my harness was obviated. Indeed, the need for any but the most liberal of harnesses may be obviated. ❄                ❄                ❄                ❄                ❄ Some tidbits that struck me from Ezra Klein’s recent (recommended) interview with Matt Sheehan on the interplay between regulation of AI and competition with China . When American policymakers are like: Where do you start? — I sometimes say: Well, you start by starting. You learn how to regulate things, you learn how to legislate on them by regulating and legislating on them. After the Hugging Face and Wiki attacks OpenAI were still unable to review their previous logs and determine that they had previously attacked RubyGems. They knew about the attack on RubyGems and made the decision not to reach out to the RubyGems team about it. While Chinese models have made some surprisingly remarkable gains in the slipstream of US frontier models, the US still has 8 times as much compute available to it than China - which is a material gap. People in the US worry that regulation will slow down the US model builders, but these rapid recent gains in China have occurred under much more regulation Americans say that when they set up a hotline to talk to Chinese leaders in a crisis, the Chinese don’t pick up the phone. But this misunderstands the Chinese system. Individual Chinese, even powerful ones, aren’t given individual decision-making power. They operate with committees and documents. So the Americans are better off sending a fax than trying to call an individual Like so many things, effective regulation needs regular practice

0 views
Unsung Yesterday

“59.94Hz is standard on televisions in North America.”

A big screen carries with it different UI expectations, and how that manifests itself in Apple TV’s Settings app, is that there’s always room for an additional hint about the thing you’ve just selected. So yes, I have been exploring Apple TV’s Settings after the recent update like any normal person would, and I have noticed how well-written those are. They help navigate difficult technical things, and they’re not afraid to get conversational, or offer advice, or even commentary… but they always seem to stay succinct and on-point. I think they’re worth studying. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/1.1600w.avif" type="image/avif"> 59.94Hz is standard on televisions in North America and other regions using NTSC. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/2.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/2.1600w.avif" type="image/avif"> Automatically select the best audio output. Some TVs require 16 bit. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/3.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/3.1600w.avif" type="image/avif"> Enjoy movies and music without disturbing others. Experience softer sound effects and music, but keep all the detail of the original sound level. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/4.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/4.1600w.avif" type="image/avif"> Apple TV will re-encode audio to send compressed Dolby Digital 5.1 to your speakers. ¶ Use only if your equipment doesn’t support Atmos or multi-channel PCM. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/5.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/59.94hz-is-standard-on-televisions-in-north-america/5.1600w.avif" type="image/avif"> Switch Control allows you to use your Apple TV by sequentially highlighting items on the screen that can be activated through an adaptive accessory. À propos Unsung’s first post ever , whoever works on these did not abuse the privilege. It would be so easy to just keep writing, but if a hint is not necessary, it just doesn’t appear: …with one exception, which is the screensaver section. Here, the hint strings feel redundant, mostly repeating what the label already said:

0 views

Data Broker Radaris Loses Domains in Privacy Fight

The consumer data broker Radaris.com has long had a reputation for ignoring requests to remove personal information from its vast empire of people-search services online. That reputation caught up with the company recently in a lawsuit alleging Radaris violated a New Jersey privacy law that provides for hefty fines against data brokers that publish personal information on state law enforcement officials. In the face of repeated stonewalling and prevarication by attorneys for Radaris, the judge in the case ordered that radaris.com and more than a dozen other data broker domains be transferred to the plaintiffs. The radaris.com website, prior to the domain transfer to Atlas. In February 2024, Radaris was sued by Atlas Data Privacy Corp , a company that has been pursuing data brokers alleged to be violating a New Jersey statute called Daniel’s Law . The statute allows state law enforcement officials, government personnel, judges and their families to have their information completely removed from commercial data brokers and people-search services, and provides for fines of $1,000 per violation against companies that ignore removal requests. Less than a month after Atlas sued Radaris, KrebsOnSecurity published a deep dive into the Radaris co-founders — Igor and Dmitry Lubarsky (also spelled Lybarsky) — Russian-born brothers living in Massachusetts who operate a dizzying array of people-search companies as well as a number of Russian language dating services and affiliate programs. Attorneys for the Lubarsky brothers threatened to sue for defamation if the story wasn’t removed and an apology issued. Their attorney asserted that our reporting was wildly inaccurate, and that the true owners of the company were Ukrainians living in Ukraine. The Lubarsky brothers Dmitry or “Dan” (left) and Gary/Igor. KrebsOnSecurity doubled down and showed how the Lubarsky brothers built and operated Radaris and other data broker companies using a fictitious CEO’s name . Our follow-up story noted that Radaris’s attorney — a lawyer with the Boston Law Group named Val Gurvits — admitted his clients had invented the CEO pseudonym “ Gary Norden ,” and that Radaris also had issued multiple press releases over the years that quoted the fake CEO while seeking money from potential investors. Attorneys for Radaris waited until the last minute to appear in court and contest what was all but certain to be a default judgment in favor of the plaintiffs, and then told the court that Atlas had failed to serve the real owners and operators of Radaris and several of its sister data broker companies. Atlas re-filed the lawsuit in June 2025, this time dramatically expanding the number of Radaris family data brokers accused of violating Daniel’s Law. Matt Adkisson , president and CEO of Atlas, said Radaris turned to a tried-and-true playbook: Delaying in court until the last possible minute, and playing shell games with Radaris’s true country of origin and the individuals listed as owners and operators of these sites. “We refer to this period as their island-hopping phase. Privacy policies changed constantly, and new entities kept appearing from places like the Marshall Islands, the British Virgin Islands, and Seychelles,” Adkisson told KrebsOnSecurity. “Behind the scenes, it felt like a shell game. Defense lawyers told the court that certain entities merely operated the domains and were the proper parties to sue. But by the time a judgment neared, those entities would be discarded and new entities would appear. Meanwhile, the lawyers claimed the other entities that actually owned the domains should not be held responsible.” Adkisson said when the defendants updated their terms of service to state that Radaris was suddenly managed by a company in the Marshall Islands, Atlas hired an investigator in that country and soon learned the brand new entity that Radaris claimed was managing the company didn’t even exist yet. Mr. Gurvits stepped forward as Radaris’s attorney in a class action lawsuit the company temporarily lost in 2017 because it never contested the claim in court. When the plaintiffs told the judge they couldn’t collect on the $7.5 million default judgment, the court ordered the domain registry Verisign to transfer the radaris.com domain name to the plaintiffs. Mr. Gurvits appealed that verdict, arguing the lawsuit hadn’t named the actual owners of the Radaris domain name — a Cyprus company called Bitseller Expert Limited  — and thus taking the domain away would be a violation of their due process rights. The judge in the 2017 case ruled in Radaris’ favor — halting the domain transfer — and told the plaintiffs they could refile their complaint. Soon after, the operator of Radaris changed from Bitseller to Andtop Company , an entity  formed  (PDF) in the  Marshall Islands in Oct. 2020. The plaintiffs never re-filed their lawsuit. A mind map of various entities tied to Radaris and the company’s co-founders. Click to enlarge. “That seemed to be their modus operandi,” said Raj Parikh , a partner at PEM Law in New Jersey who handles most of the Daniel’s Law litigation for Atlas. “In the past, they won by attrition. Plaintiffs’ attorneys tired of the procedural games and just gave up. That strategy worked for a decade, and it probably would have worked in this case too, since any financial recovery from foreign actors will be difficult. But we were acutely aware of the threat this website posed to law enforcement officers and other public officials in New Jersey, and decided early on to commit whatever time and resources were necessary to remove that threat.” On August 26, the judge in the New Jersey case found the defendants were given multiple chances to appear and defend the claims against them but had failed to do so. Mr. Gurvits declined to comment on the case, saying it had been assigned to another attorney, a Mr. Victor Worms . In response to questions, Mr. Worms asserted the New Jersey court transferred Radaris.com to Atlas as part of a default judgment against Radaris.com, which is not a legal entity. “We have made a motion to vacate that default judgment on the grounds that it is void since a non-entity has no legal capacity to sue or be sued,” Worms replied. “We also intend to pursue all appropriate appeals because we believe the transfer of Radaris.com amounts to a forfeiture in violation of various constitutional principles.” While radaris.com still comes up prominently in results when searching online for U.S. residents by name, the domain no longer sells detailed personal dossiers on millions of Americans. Its homepage now displays a notice from Atlas, as well as links to our previous reporting on Radaris. Atlas told KrebsOnSecurity that it has obtained more than 10,000 emails and documents in the course of litigation, and that those messages confirm our previous reporting on the owners and operators of Radaris and its myriad companies. Atlas said the emails clearly establish that the nominal legal vehicles — Radaris America, Inc. ; Bitseller Expert Limited ; Digital Orbit Corp ; Core Solutions Group Inc ; Lucky Solutions Inc ; Virtura Corp ; Veripages Inc. ; Nuform Solutions Inc. ; Growth Data Advisors Inc. ; Property Experts, Inc — are all administered by the same three or four people from the same mailboxes, share one bank or payment card set, and are all managed from one virtual office address. “The corpus establishes, with documentary evidence generated independently by banks, payment processors, hosting providers, registrars, software-as-a-service vendors and the operators’ own systems, that radaris.com and at least twenty-five other people-search websites are one operation run by a small Boston-area group whose administrative, financial and technical functions sit on the difive.com mail domain and its successors (centerex.com, scienteco.com, eprofit.com, realmo.com, pub360.com),” reads a summary shared by Atlas. Atlas said the emails show Radaris.com earns approximately $42,000 a month, while Veripages.com earns around $45,000 monthly via its partnership with the Lifetime Value Company , a marketing and advertising firm whose brands include PeopleLooker , PeopleSmart , NumberGuru , and Bumper , a car history site. According to Atlas, the emails also showed the Radaris family of websites earns as much as $25,000 each month from their partnership with Onerep , a company that claims to help people remove their information from people-search sites. In March 2024, KrebsOnSecurity revealed how the Belarusian founder of Onerep had launched and operated dozens of people-search sites over the years and was continuing to operate one of them (Nuwber), effectively spreading the disease and selling the cure. The domain radaris.com now redirects to this notice from Atlas about the court-ordered domain transfer. All told, the New Jersey court has so far transferred 14 domain names from the Radaris family of companies to Atlas. Radaris.com now redirects to a notice of the court-ordered domain transfer. The Radaris family of companies is still potentially facing fines of $1,000 per alleged violation of Daniel’s Law. For the time being, however, Daniel’s Law is facing a constitutional challenge from virtually all of the 150 other consumer data broker firms being sued by Atlas. The data broker industry responded by having at least 70 of the Atlas lawsuits moved to federal court, challenging the New Jersey statute as overly broad and a violation of the First Amendment. The U.S. Court of Appeals for the Third Circuit has not yet issued a decision on the constitutional challenge, but either way the case is widely expected to be appealed all the way to the U.S. Supreme Court. Meanwhile, at least 14 other states have now passed laws modeled after the New Jersey statute, with more states considering similar measures. However, West Virginia’s Daniel’s Law was ruled facially unconstitutional under the First Amendment by a federal district court in August 2025. Justin Sherman is a privacy expert and author of the forthcoming book “The Middlemen,” which examines how the data broker industry powers modern surveillance. Sherman said federal lawmakers have long faced intense lobbying by the technology industry against more restrictive U.S. data privacy laws, but that many powerful industries are now working against passing comprehensive data privacy legislation. “These days at the federal level, add in the intense amount of lobbying against these laws from social media companies, big tech, cryptocurrency firms, and now AI proponents in the mix who claim that limiting their data scraping is somehow going to collapse the whole U.S. economy under Chinese rule,” he said. Sherman said people-search companies will continue to thrive unless and until Congress enacts meaningful consumer privacy and data protection laws that are relevant to life in the 21st century. That’s because virtually all state privacy laws exempt records that might be considered “public” or “government” documents, including voting registries, property filings, marriage certificates, motor vehicle records, criminal records, court documents, death records, professional licenses, bankruptcy filings, and more. At least 25 states have passed or implemented laws requiring age verification for residents seeking to access adult content online, but there is no federal law that limits how the companies that are scanning everyone’s drivers license can use, share or keep the data provided. Had such restrictions been enshrined in law, we may have avoided the recent breach at IDScan.net , which exposed the drivers license information on more than 153 million Americans when the records were briefly turned into a point-and-click identity theft service on the dark web. “The average person can look at Daniel’s Law and have a perfectly normal reaction, which is that everyone should be covered, not just police and judges,” Sherman said. “But we don’t need more wake-up calls. We’ve had eight million wake-up calls already on the need for better privacy laws. The lack of comprehensive federal privacy law is not for a lack of knowledge, and anyone claiming otherwise is either not reading the news or kidding themselves.”

0 views
dfir.ch Yesterday

Living Inside the Shell: zsh Modules on macOS

When investigating shell-based activity on macOS, it is tempting to focus on the usual suspects: , , , , and similar utilities. But itself provides considerably more functionality than simply executing commands. macOS uses as the default interactive shell, and zsh ships with a module system that can extend the shell with networking, file manipulation, extended-attribute access, and other functionality. Functionality commonly associated with separate utilities can instead be performed by builtins inside the already-running process. As a result, there may be no corresponding , , or process for an analyst to find. Detection gaps can arise when detection logic relies primarily on process execution and command-line telemetry.

0 views
Chris Coyier Yesterday

Which Rude is it?

The commerical airport we use here in Bend, Oregon is actually in Redmond, Oregon. Flights from here generally depart very early. It think it’s because they need to make it to bigger airports to make connections to further-away places. Flight typically depart at 4:30-6:30 AM. They want your bags an hour before departure, and the airport is 30 min from Bend, so you gotta be out the door sometimes at 3:00 AM meaning ungodly 2:30 AM alarm clocks. That’s the extreme case though. If you aren’t checking a bag and you’ve got a 6:00 AM flight, maybe you’re leaving the house at a spicy but tolerable 4:45 AM. That was too much preamble for this, but now you know. The one giftshop/coffeeshop in the airport opens at 4:00 AM. One person opens it up and starts selling things to the couple hundred people milling around in the one terminal preboarding area. This shop sells all the normal stuff you see in airport giftshops like cheezy Central Oregon sweatshirts and magnets, cold beverages and string cheese, magazines, and the like. They are also, and perhaps mainly, a coffeeshop. People stand in line to buy coffee. It’s early in the morning. You can’t bring in liquids. It’s damn coffee time. Right in the heat of the morning airport action, there might be 20-30 people in line. It’s a whole thing. Now we’ve arrived at my point. What do you order from this one person working at this coffeeshop at 4:00 AM? You can’t help but be aware there are 20 people behind you in line and how there is one person taking orders and making the coffee drinks. Right?! You could order a latte, which will take like 3 minutes to make. Or you could order a drip coffee in which this person hands you a cup in 3 seconds. My brain is built such that I cannot possibly order something that will take this person a while to make. Like the words would be unable to come out of my mouth. Even if a cortado sounds really good right now, actually , I can’t do it. I can make an active choice to get a perfectly fine drip coffee and get this line moving and get all these strangers-yet-neighbors their coffees too, or I can cause a big ol’ hitch in the giddyup. I hope I’m not trying to grandstand how perfect I am. I’m showcasing one part of how my brain works. I really don’t like inconvinencing other people. I notice, because it seems like plenty of other people don’t. People order cappaccinos and flat whites and all that shit without abandon. The line takes forever. It just is what it is. And we come to why I titled this The Rude Trifecta. These mocha-ordering fellow humans must fall into one of these categories: I actually don’t know how it would break down if there was a way to figure it out, but I suspect it’s a fairly even mixture. Like for some, it just doesn’t cross their mind that it’s any problem at all to order a 3 minute drink. It’s a coffeeshop and they ordered a coffee. Maybe if they thought about it for far too long like myself, they could see the problem, but that’s not their normal thinking pattern. For others, they couldn’t give any less fucks. Again it’s a coffeeshop and they ordered a coffee. They stood in line like everyone else. Yeah, it might take a while, but it’s their turn and they are going to use it. Put whip cream on it motherfucker. The last one is very similar to the above, but it’s more intellectual. Again it’s a coffeeshop and they ordered a coffee. This is not a rude action. It’s not on them to dechiper what is and isn’t rude on a menu , or to personally shoulder a understaffing issue. They might go so far as to think it’s actually rude in the other direction , where self-censoring an order doesn’t give the business the appropriate feedback on their operations. That’s why if I was with a friend and they did it , I’d be totally fine with it. I can’t do it. I can’t ask them to get me the americano. But their actions are their own and this isn’t a situation where I cast any judgement. I mean assuming it’s #3 and not #2, that is. Speaking of airports and flying, this is why I literally cannot recline my seat if someone is behind me. It takes up their room. Can’t do it. Reminds me of a recent-ish Marcel post : I feel like the neighbor: They don’t know that it’s rude They don’t care that it’s rude They disagree that it’s rude doesn’t know it’s rude doesn’t care it’s rude disagress that it’s rude

0 views
Stratechery Yesterday

Salesforce AI Force, Agents as UI, The Race to Headless

Salesforce is abandoning UI as a moat, which is a very smart move because it's disappearing for everyone.

0 views
Andy Bell Yesterday

An RSS feed for a daily album recommendation

My music collection is ever-growing and I have a habit of listening to the same stuff on repeat. Mostly, it’s overwhelm-related because as you can see by this album covers view , there’s a lot to choose from. I thought RSS could be useful here, so I wrote a little feed that picks a random album, each day, then serves that as the single RSS item. This works great because my RSS reader, Feedbin will happily build those up if I don’t get around to checking. Try it for yourself . You might discover something you love!

0 views
Farid Zakaria 2 days ago

Visualizing Nix closures

tl;dr seenix.dev lays every byte of a Nix closure out on a map, one pixel per byte, and lets you zoom from a whole NixOS system down to the hex of . Try hello , firefox or a GNOME desktop . Nothing runs on a server. With the advent of LLMs I keep tugging at any crazy question I ask myself. I know there is the anti-AI crowd and they will happily proclaim anything pursued in this vein as “slop” but I am feeling fortunate to be able to explore these questions. My recent itch was to ask “what does a Nix closure look like?” and to answer it in a way that is interactive and visual . I wanted to see the bytes, not just the store paths. I had come across binvis.io on Hacker News and I found it a compelling way to look at data. I personally never found a need for it, but I found it fascinating none-the-less. 1 The timing for this itch was perfect. I noticed a trending thread on X where a Python binary seemingly includes and . 🤷 I built that tool. You can check it out at seenix.dev . It is a single-page web app that runs entirely in your browser, with no server. It fetches the narinfos of a closure and lays them out on a map, one pixel per byte, and lets you zoom in to see the bytes themselves. We can visualize the closure of that binary, , and see if it really does include those two packages. Turns out it does not. The closure is 41 store paths and 234 MiB, with no and no among them. Turns out those dependencies are build-time and are not included in the final runtime closure. We can visualize much larger closures. Here is a GNOME desktop: 1,324 store paths and 5.3 GiB, each colour one package. That picture needed zero NAR downloads. It was laid out in 3 ms from the narinfos alone. 🤯 The “trick” I learned to make this visualization possible, is the Hilbert curve . A Hilbert curve is a single, unbroken line that folds back and forth such that it completely fills up a flat square. It is a fractal . Every store path in the closure is sorted by name (the root first) and their NARs are concatenated into one long line of bytes. The Hilbert curve folds that line into a square, so byte n is pixel n along the curve. The Hilbert curve has two properties that lend itself nicely to visualize binaries and as a result Nix closures: Bytes that are near each other in a file stay near each other on the map. A NAR is a single contiguous range of bytes, so a store path is a single contiguous region on the map. A file inside that store path is a smaller contiguous region, and a section inside that file is smaller still and so forth. Squares are just byte ranges. Here’s a tiny 4×4 map. Each number is the byte that lands on that pixel: That means we can easily place a store path on the map by knowing its starting byte and its size. That’s what makes the map cheap to draw. 2 The layout only needs each path’s , which every narinfo carries, so the whole map exists before a single NAR is downloaded. Hovering already tells you which store path you are pointing at, its size, its retained size (the bytes that would leave the closure without it) and a “why is this here” chain back to the root. As you zoom in, the NARs on screen are fetched from the cache and the color fills in. Here is ’s closure, most of which is glibc: Blue is printable ASCII, red is high bytes, green is control bytes and black is . The speckled top is machine code. The big solid blue area at the bottom is glibc’s locale data, which is plain text. Keep zooming and every pixel becomes a byte you can read. Hovering names the file inside the NAR, and for ELF files, the section. That is of , in your browser tab, fetched from cache.nixos.org , without any server . 😈 Does everything need a purpose? Sometimes something is fun to make and to use with no real purpose. For fun, I even added a Save PNG button, and it saves the view at the canvas’s full resolution. The ultimate ricing of your NixOS system: a pixel image of your desktop closure. Can Omarchy do that? 😎 Anything you can export works: Drop the file and see the map. You can provide additional Nix binary caches to fetch NARs from as well. The source is at github.com/fzakaria/seenix . Go look at something big. Build without purpose. Have fun. Aldo Cortesi’s writing on visualising binaries is a great resource on this.  ↩ This is why the world is always a power of four bytes. hello’s closure is 36 MiB, which fills a bit over half of a 64 MiB square, and the rest is drawn as background.  ↩ Aldo Cortesi’s writing on visualising binaries is a great resource on this.  ↩ This is why the world is always a power of four bytes. hello’s closure is 36 MiB, which fills a bit over half of a 64 MiB square, and the rest is drawn as background.  ↩

0 views
Sean Goedecke 2 days ago

Jev means structured output is interesting again

I don’t write blog posts about new models. That’s Simon Willison’s beat, and he’s very good at it. But I want to write about Jev , which is a different kind 1 of AI model: a “System One” 2 model. As it turns out, it’s not that different from an ordinary LLM with structured output, but the interface it uses is very cool and I hope it becomes more widespread. Ordinary LLMs take in some human-language prompt and produce some human-language output. They do so autoregressively : first they produce one token, then the next, then the next, and so on. This makes them extremely flexible, since they can do literally anything a computer can do. But it also makes them slow and weird. Slow, because they have to run a whole new generation pass per-token, and weird, because the space of human language is so broad that you can get really odd behavior from a model trained on it. Jev takes a human-language prompt, but it does not produce human-language output. It only produces structured output. So far, so ordinary: LLMs do this already . But it turns out that if you build a model that only produces structured output, you get some interesting and desirable properties. Jev is always really fast. The fastest response time is around 70ms instead of a couple of seconds for normal LLMs. Even better, the slowest response time is only 500ms. Because Jev only does structured output, it isn’t autoregressive: it can produce answers to many questions in parallel in a single forward pass. When a LLM is producing structured output, it has to produce the tokens ”{”, ” ”, “answer”, ”:”, and so on with successive forward passes 3 . Jev does it all in one go. The most compelling example of Jev’s speed is that the model can play Doom . You can feed a text-based representation of the current game state into the model, combined with a set of choices like “should the trigger be held down”, “what should the current goal be”, “given that the current goal is X, what keyboard input should be pressed”, and so on, and it works — latency is low enough and the system is smart enough that the model plays well in real time. Of course you could train a neural net to play Doom already. But Jev is a general intelligence: just like LLMs can do your taxes, perform mathematics research, fix your Python environment, and write you a poem, Jev can do many other tasks besides playing a single video game. Current LLMs can play Doom too (albeit slowly). But as Nelson Elhage famously said , fast software doesn’t just mean we can do the same tasks faster, it means we can do entirely new kinds of tasks. What kinds of new programs can we write by injecting 100ms worth of dirt-cheap intelligence at various decision points? To me, this is the most exciting thing about Jev. Fast structured output could be a genuinely new computational primitive for intelligence. So far we’ve built a lot of programs on top of autoregressive token generation, and they all look like fancy chatbots. Leaning hard into structured output might conceivably unlock a bunch of non-chatbot use cases for AI. My biggest problem with Jev is that I think fast structured output is already available . Structured output from LLMs is only slow because (a) nobody really cares about it 4 , and (b) the people who do care about it want big JSON blobs, so it’s typically implemented with “grammar-constrained decoding” : the LLM outputs autoregressively as normal, but the logit sampler discards tokens that don’t fit the structured output (e.g. if there hasn’t been a ”[”, you can’t output a ”]”). If you want fast, parallelized structured output against limited choices, you don’t strictly need to do autoregressive generation at all. You can simply prefill the response with and generate one token 5 , restricted to the user-provided choices. Since LLMs ingest all input tokens in parallel, this is way faster than generating the entire structured output. Multiple choices can be batched into the same forward pass via ordinary inference batching. This doesn’t let you do long-form structured output, but in return you get most of 6 Jev’s “secret sauce”: the speed, the consistency, and the parallelism of a System One model. People have already started trying this after today’s Jev announcement, and it seems like it’s working OK 7 . In other words, I suspect Jev does not have a substantial technical moat, and their claimed “Reinforcement Learning for Calibrated Decisions” is not a brand-new scaling axis. It will probably be pretty easy for any other lab to replicate, or for individual programmers to retrofit existing open-source LLMs into a fast Jev-like model. However, I suspect Jev is still going to be better than most versions of “Qwen-32B-System-One” or whatever. Being able to fine-tune or optimize the model on just structured output is probably a meaningful advantage. I doubt Jev is ever going to be as smart as frontier LLMs. Not being able to use test-time compute at all 8 is a big disadvantage, and will likely cap this kind of model around the strength of non-reasoning LLMs. In practice this shouldn’t matter too much for low-latency applications, but you shouldn’t see this as a new scaling axis or a way to produce more intelligent models. Jev’s developers claim it is immune from hallucinations. To me, this seems like a semantic dodge, since Jev can absolutely still pick the wrong choice (e.g. calling the sky “red”). I suppose that’s technically just a mistake , since the model is picking a user-provided choice instead of inventing something new out of whole cloth. Still, all of this is also true about regular LLMs with structured outputs, and it doesn’t make Jev any more reliable in practice. It’s unclear to me how much of Jev’s value is in the model itself, compared to the inference strategy of only generating one token per question. The data and demos in the announcement look to me like they could have been generated by plugging any Terra-sized model into a single-token inference stack. However, the people involved are credible, and I’m sure the model is good — I just wish they’d provided some comparisons that didn’t force the LLM to unnecessarily produce a blob of JSON token-by-token. Overall, I am happy that Jev exists and I hope it succeeds. I hope we do see some real competition in the fast-structured-output space, and that it motivates the big labs to release official versions of their own models that are fine-tuned for this. GPT-5.6-Terra-System-One would be a very interesting model to build AI products on top of. I did write about Thinking Machines’ “interaction models” , which are also a fast-enough-to-be-meaningfully-different paradigm for AI inference. They call Jev a “System One” LLM, after Daniel Kahneman’s partially discredited Thinking Fast and Slow , where he divides human cognition into a lightning-fast System One and a slow-and-reflective System Two. If you’re thinking “wait, couldn’t you just aggressively prefill a regular LLM and only produce one constrained token”, keep reading. Not counting tool calls, which are built-in in a way that structured output isn’t. What if some of the user’s choices are longer than a single token? I haven’t tried this myself, but I’m sure you could translate them into a single token, or train the model to output “1/2/3” under the hood instead of the choice content, or generate only the first token of the choice if it’s different, or some other clever trick I haven’t thought of. Jev claims that their generated probabilities are “calibrated”, but I haven’t seen anything to suggest that these aren’t just regular logit probabilities. Maybe there’s some clever training they do to encourage accurate logprobs in uncertain situations (e.g. getting the model to produce when predicting a coinflip, etc)? If so, I wish they’d written more about that in the announcement. I tried it myself with and got a 2x-3x speedup compared to non-prefixed structured output. I suppose they could do some looped-transformer thing where they loop some fixed amount of times, but anything that looks like reasoning would make the model latency slow and unpredictable, defeating the entire purpose. I did write about Thinking Machines’ “interaction models” , which are also a fast-enough-to-be-meaningfully-different paradigm for AI inference. ↩ They call Jev a “System One” LLM, after Daniel Kahneman’s partially discredited Thinking Fast and Slow , where he divides human cognition into a lightning-fast System One and a slow-and-reflective System Two. ↩ If you’re thinking “wait, couldn’t you just aggressively prefill a regular LLM and only produce one constrained token”, keep reading. ↩ Not counting tool calls, which are built-in in a way that structured output isn’t. ↩ What if some of the user’s choices are longer than a single token? I haven’t tried this myself, but I’m sure you could translate them into a single token, or train the model to output “1/2/3” under the hood instead of the choice content, or generate only the first token of the choice if it’s different, or some other clever trick I haven’t thought of. ↩ Jev claims that their generated probabilities are “calibrated”, but I haven’t seen anything to suggest that these aren’t just regular logit probabilities. Maybe there’s some clever training they do to encourage accurate logprobs in uncertain situations (e.g. getting the model to produce when predicting a coinflip, etc)? If so, I wish they’d written more about that in the announcement. ↩ I tried it myself with and got a 2x-3x speedup compared to non-prefixed structured output. ↩ I suppose they could do some looped-transformer thing where they loop some fixed amount of times, but anything that looks like reasoning would make the model latency slow and unpredictable, defeating the entire purpose. ↩

0 views
Chris Coyier 2 days ago

The Four Tiers of Tab Importance

Arc is the greatest web browser ever, and has been tragically moved-on-from by The Browser Company of New York-come-Atlassian. I’ve been back on it the last month or so though. It’s still very usable as they keep the Chromium version updated. I just really like it. It’s so good. My second favorite is Zen because of how well it follows in those Arc footsteps. But I’m attempting a jump over to Dia , the sorta-kinda-Arc-replacement, as it seems like that’s where the effort is focused. But is it?! I don’t see a ton of action on Dia either, to be fair. But they have seemed to bring some of the great some from Arc over to Dia, so I figured it was worth a shot. There is already a bunch of paper-cutty stuff I don’t like, but I gotta give it some time, so I won’t dig into all that just yet. Right now I’d just like to explain one thing I think Arc really nailed : Tab Heirarchy. It’s sort of like a 4-tier system. These favicon-only buttons are tabs that persist across all spaces. Their position and ubiquity make them, perhaps, the highest tier tabs. At one point I had it in my head that Arc “kept these tabs hot” meaning if you clicked onto one of them, it was already rendered, so you felt no delay as that page loaded. Not super sure that’s true, but it would be cool if it was (and worked so well it was obvious). The icons are a little small which reduces their prominence a smidge, but I’d still call them the top. The Problem in Dia: Dia has these, but there are Profile-specific, which to me ruins the heirarchy. Why have them at all if they don’t have the ubiquity? I really don’t know what to call these, but they are also high on the hierarchy and probably equal to those pinned tabs in importance. But they don’t persist across spaces — they are very space-specific. They’re below the pinned tabs, but above (separated by a little line) the regular tabs. These tabs sort of behave like bookmarks, which is a fantastic feature that I’ve really grown to love. You can just close them and instead of literally closing and disappearing from the sidebar, they just reset to their main URL. Closing them is just like resetting them. You can remove them, of course; it’s just a more explicit action. These are great. The Problem in Dia: None. Dia has these and they are fine. The tabs below the little line are regular tabs. They are remarkable for their unremarkableness. They are just tabs. You open them and close them and behave exactly how you’d expect a tab to be. They do have one notable feature: Arc has a setting to auto-archive these tabs after a set period. It’s like a “save you from yourself” feature. I have mine set to 30 days, as I actually don’t like this feature. I keep a tidy browser anyway and don’t need to be saved here. I know some people really like it though, people that I assume also have Roombas. The Problem in Dia: Dia just doesn’t sync these?! WTF?! It syncs literally everything else but just stops short of syncing your normal tabs. Perhaps the lowest on the hierarchy are “Little Arc” windows. It takes some serious getting-used-to in Arc that you don’t open multiple windows. You just have the one browser window. It’s weird to have multiple windows. It lets you, but it probably shouldn’t. Instead, if you need a 2nd window for a sec, which is legit, you just open a Little Arc, which is this very transient browser window with none of the Arc UI around it. You do your little thing and close it. Or, you “promote” it to a regular tab with the one prominent button a Little Arc has. Little Arc is what Arc uses to open links from other apps. Like if you click a link in your email app, it’ll open in a Little Arc first. I love this. Chances are, these are ephemeral browser “tabs” I just need to look at for one sec, then whisk away. If not, I’ll just promote it. The Problem in Dia: Dia just doesn’t have these ephemeral windows. Booooo. This is the #1 loss I feel in Dia. Both Arc and Dia have this nice feature where you basically ⌘-T to make a new tab, and type in what you’re looking for. But it doesn’t just do one thing. It’s got a menu of choices. Of course, the top choice needs to be right most of the time, and it usually is, but options are nice. The Problem in Dia: It’s just not as good as Arc was. For one, it really wants to hijack many would-be web searches for “Chat” instantiations. So it answers with some ambigous LLM instead of searching. I use AI, but I literally never want this in Dia as I’d rather just use an LLM of my choice. Dia also isn’t as good at commands. It change change color scheme, it can’t open browser extensions, it doesn’t have splitting commands, lots of missing stuff. I mentioned this above briefly, but I’d like to mention again: Your profiles sync. The pinned tabs in those profiles sync. But not your other tabs. This just sucks. I use multiple computers, I want all my tabs to sync. The Problem in Dia: Normal tabs don’t sync. Both Arc and Dia have splitting, meaning you can see two websites side by side, which is so good it gets copied . Friggin love it, use it constantly. This is one of the ways “just having one browsing window” works so well. You probably have a system for this if you’re a non-Arc/Dia user already with windowing apps that help set multiple windows where you want them. I actually like just having it done right within one browser window. It just feels good. The Problem in Dia: It’s not as good in Dia. You can’t drag two tabs on top of each other to split. The command bar doesn’t have a command for splitting. I set up a key command for it which is OK, and you can still Option-Click which is crutical, so it’s live-with-able, but barely. “Profiles” in Dia are more like how other browsers do it. When you switch profiles, it’s kinda like you’re in a new isolated browser. If you’re logged into CodePen in one profile and then switch to another, you’re no longer logged in. I would think some people find this an improvement of Dia over Arc, as Arc didn’t have a profiles feature. If they love it, that’s cool, I just never used profiles and don’t like them. I preferred how spaces were just groupings of tabs in Arc. The Problem in Dia: Dia only has little dots for the Profiles where Arc has icons/emojis for Spaces. I’m not always on a computer with a touch pad, so I preferred the larger click area in Arc. It’s worth mentioning because Arc forced this, and Dia just makes it optional. To me, it’s required now. And literally all the major browsers offer this now, which to me proves how rad it is. Dia does side tabs just fine. There are little things I prefer in Dia, like I didn’t need the Easels and Boosts and all that, so the removal of those things is fine with me. It might offer to open up a web search for your thing. It might offer to switch to an already-open tab that you may or may not realize you already had open. It might offer a recently-visited page it can re-open for you. It might be a command.

0 views