Latest Posts (20 found)

Premium: AI Is Getting Way Too Expensive

A great deal of the discussion of the so-called benefits or problems with AI comes down to the theoretical jobs that are (or are not) lost as a result of things LLMs can (or cannot do), or the equally theoretical productivity benefits that’ll come from using LLMs in place of (or in conjunction with) humans. Anthropic’s Economic Index and OpenAI’s Economic Research Exchange are marketing operations that exist to propagate the (wrongheaded) belief that LLMs are either leading or will soon lead to massive economic or productivity shifts, even though little or no actual evidence exists to show that this is the case, other than the occasional story about LLMs make people worse or slower at their jobs or single lines in studies that are used (incorrectly) to prove that “ AI is making it harder to find a job for young people .” In fact, Anthropic’s Head of Economics recently said there was “no material increase in the unemployment rate to date.” These conversations materially detract from the actual harms or effects of AI, and exist only to make you scared that AI will take your job. They do not have any vested interest in expressing the actual economic effects of AI, which are, at this point, a simmering cauldron of different speculative bets on whether or not LLMs — a definitively niche technology — will create or become general-purpose software ( per Roger MacNamee ) that scales into the next Google Search, iPhone, or Microsoft 365. As I’ve argued again and again, the AI industry’s revenues are, outside of Anthropic and OpenAI, incredibly small. Even in Exponential View’s deliberately-pro-industry analysis , there’s only around $110 billion in trailing twelve-month revenues across the entire industry, including OpenAI and Anthropic’s cloud spend.   For those counting at home, that’s $12 billion less than the $122 billion OpenAI raised in March , and a full $145 billion less than all AI startups raised combined in the first quarter of 2026 .  Anthropic and OpenAI want you to talk about the theoretical so that you don’t focus on the tangible — their hundreds of billions of dollars’ worth of commitments, said commitments effects on the remaining performance obligations of hyperscalers and chip manufacturers, and the sheer scale of venture capital’s investment in AI, which ( as I’ve argued in the past ) largely allows for massive on-paper gains with little or no hope of liquidity. To put this bluntly, I believe the entire conversation around AI’s theoretical relationship to jobs to be masturbatory and a conscious attempt to avoid having a messy conversation about the scale of the actions taken based on the flimsily-founded promises of AI labs and hyperscalers.  Today’s piece will dig into the true scale of the money needed to make AI make sense, by which I mean how much OpenAI and Anthropic will need to meet their commitments, how much money hyperscalers will need to pay off their investments, venture capital’s true exposure to the AI bubble, and what will have to go right for the bubble not to be, well, a bubble. I’ll also make the case that the longer the bubble continues to inflate, the harder the basic economic puzzle of AI becomes to solve, as creating and deploying infrastructure becomes vastly more expensive — meaning that in order to achieve profitability, hyperscalers and neoclouds need to charge significantly more for compute than before, and the only two real potential customers are ones that cannot pay for it.  This will be a more more-pointed newsletter than usual, focusing on hard numbers and harder truths.  Let’s get it on.

0 views

End of the contact form saga

I can’t take it anymore! If you wan’t to speak to me, send an email. My contact form is out of service indefinitely. This is actually in lieu of moving my professional services to a yet to be announce limited company. But I can’t let opportunity for a dramatic blog post go to waste. Also, I’ll probably skip the contact form on my company website. Is that a bad idea? I always got more spam via the form than the publicly visible address. My contact form has been through a lot. Previous entries in the saga: I quite enjoyed the week in September when I opened a port to a self-hosted SMTP server I coded in 100 lines of TypeScript. The final iteration of my form included true end-to-end encryption. Through trial and error heuristics, I successfully eliminated all spam. (How many false positives I rejected remains unknown…) My privacy policy which was already simple is now entirely pointless. Are contact forms just outdated in general? Everyone seems to embed a Calendly widget these days. That’s not my style. I like the tiny bit of additional friction required to send an email. If someone can’t be bothered their message probably wasn’t serious. I’m not looking to maximise meaningless engagement. I look forward to moving business email to a separate domain. Biggest mistake I ever made was using for personal and business. Nothing worse than seeing an “urgent” request only to find out on Monday it didn’t matter. So long old contact form, it was fun! Thanks for reading! Follow me on Mastodon and Bluesky . Subscribe to my Blog and Notes or Combined feeds. SMTP on the edge Email: the final form I shut the emails out I let the emails in Progressive dehancement PGP encrypted contact form

0 views
Unsung Today

One and one thing only

I wanted to show you a year’s worth of messages from my barber’s software, because this is what software should be. = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/one-and-one-thing-only/1.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/one-and-one-thing-only/1.1600w.avif" type="image/avif"> No spam, no upsells, no growth hacks, no unwelcome cuteness or puzzling verbosity. It’s so curt and straight to the point it should be set in a 1970s Helvetica: The appointment’s tomorrow. Any questions?

0 views
Unsung Today

“Insights and ideas emerge from that interface.”

From Robin Sloan : Yet we should never forget that the product of work isn’t only the work — it’s also the worker. Doing the work changes you; the up-close expe­ri­ence trans­forms your capa­bil­i­ties and even your desires. Insights and ideas emerge from that interface, and I believe the agent mae­stros lose a lot — too much — when they take their big step back. But I suppose this is just my temperament, which I’ve written about before . I don’t merely want things done; I want to do them. #ai #craft

0 views

The Conductor Developer

TL;DR Why I think software development is starting to feel a little more like conducting an orchestra. There’s a shift happening in software development that I don’t think we’re talking about clearly enough. For the last couple of years we’ve framed AI as a productivity tool. How much faster can it write code? How many more features can we ship? How much cheaper can we build software? I think that’s the wrong question, but I understand why. The first thing AI became good at was writing code, so naturally that’s where we focused. As AI got better at coding, I expected the bottlenecks to move through the software delivery lifecycle: from coding to design and specification, architecture, then verification. And they have. We spent a lot of time at the most recent FOSE event discussing how we ensure good design, quality and resilience while agents increasingly write the code. That’s a topic for another ramble. A few months ago, though, I realised I was looking at the wrong bottleneck. I kept assuming it would simply move to the next phase of software delivery. I was wrong. AI didn’t change what great software looks like. It changed what’s scarce. Human attention is now the bottleneck. The next bottleneck isn’t design. It isn’t verification. It’s us. More specifically, it’s our attention. Developers have always protected long periods of uninterrupted focus because that’s where good software gets built. Pair programming. Quiet afternoons. Deep work. We optimized around flow because flow mattered. When we didn’t get that time, very little got done. But when I watch developers using AI today, I see something different. The best developers I know aren’t spending all day in flow anymore. They’re orchestrating agents. Great developers are starting to look less like programmers and more like conductors. I was watching Jacob Collier on YouTube recently because I’m hoping to see him in concert soon. Watching him conduct is fascinating. He’s not trying to play every instrument himself. He’s listening to the whole piece, hearing what doesn’t quite fit, bringing different voices in at the right moment, changing the energy, changing the tempo and shaping the performance as it unfolds. Increasingly, that’s what great software developers look like. A great conductor is first and foremost a great musician. They could play the instruments themselves. That’s not why they’re standing on the podium. Their value comes from understanding the whole score. The orchestra doesn’t need the conductor because the musicians aren’t talented enough. It needs the conductor because someone has to hold the whole system in their head. Increasingly, I think that’s what great software developers are doing. The AI agents are the musicians. The developer is the conductor. They’re deciding which agent should tackle which problem. They’re providing context. They’re evaluating what comes back. They’re spotting subtle mistakes. They’re deciding what deserves another iteration and what is ready to move on. I was talking to an engineer recently who told me they regularly have eight AI agents running in parallel. I’ve heard similar numbers from others. Ten. Twelve. Beyond that, they become the bottleneck. That number stuck with me because it sounded remarkably familiar. It sounded like my job. As CTO, I rarely produce the work myself anymore. Instead, I have lots of streams of work progressing at once. A strategy document comes back for feedback. A client opportunity needs a decision. Someone wants guidance on a technical trade-off. Another team needs context before they can move. None of it arrives neatly packaged. It comes as conversations, emails, documents, chat messages and half-formed ideas. My job is to decide where my attention belongs, make sense of incomplete information, provide context and help other people make progress. When I first became CTO, I thought I needed to get better at managing my time. I was wrong. What I really needed to learn was how to manage my energy. The challenge wasn’t the hours. It was the constant context switching. The endless stream of decisions. The feeling that nothing was ever completely finished. An executive coach taught me some things I’ve never forgotten. Protect your attention. Manage your energy. Reduce unnecessary decisions. Create systems that help your brain, not just your calendar. Lately I’ve been wondering whether developers are about to need exactly the same capabilities. A few weeks ago I shared this thought with our Chief People and Leadership Officer. His response surprised me. “I knew something fundamental was changing,” he said. “I just didn’t know how to help. Now I do.” That conversation stuck with me because we’ve spent decades helping executives succeed in this kind of environment. We coach them to make decisions with incomplete information, manage cognitive load, prioritize relentlessly and protect their energy. Yet we’re still preparing developers for a world of individual execution. We’re redesigning the tools, but we haven’t started redesigning the job. I don’t think software developers are becoming managers. I don’t think AI is replacing engineering. I think engineering expertise is simply being applied in a different place, and much more often, because execution has become so much faster. (I suspect software developers are simply the first knowledge workers to experience it, but I’ll save that thought for another rambling.) The question I’m most interested in now is this: How do we redesign engineering careers when human attention becomes the scarce resource? When I became an executive, learning to manage my own energy was one of the hardest things I’ve ever done. Even today, if I stop paying attention to it, I pay the price. I have a feeling software development is about to demand those same capabilities from many more people. And I don’t think we’ve quite realised how profound that change is.

0 views

Left-Handed People Can’t Use a Fountain Pen

Apparently left-handed people can't use a fountain pen as the left hand smudges the ink as they write. I call bullshit on that. I'm left-handed, and I use a fountain pen just fine - even with an "overhand" grip. I read this post from Neal Stephenson a few days ago (thanks to Sal for sharing it), where Neal - a professional, left-handed writer who writes his books with a pen and paper - shares some of his experience and advice on writing with a fountain pen. I've been back using fountain pens for a few months now and it's going great, but as Neal says - it's all about the paper. I use a decent quality notepad for writing my notes, with either a £30 Lamy Safari fountain pen, or a ~£100 Kaweco AL Sport. Both work great and I'm yet to find myself with pen on my hand. Here's an example of me writing in my notebook, using both my Lamy and Kaweco. As you can see, I have an overhand grip where my hand swipes across the ink as I write, but the ink is dry by the time my hand gets there. Both these pens have a medium nib too. So if you're left-handed and think you can't use a fountain pen, you can! You just need some decent paper. Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views

Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining

The GPT-2-style models that I've been training work really well, and I've even managed to train some that perform better than the original OpenAI small model in terms of cross entropy loss on a test set. But as I wrote previously , there's a mystery: why do they perform worse on my instruction fine-tuning evaluation? I had various theories about why that might be, and to me, the most plausible-seeming of them was the amount of data they were trained with. As best I can find out, OpenAI's models were, by modern standards, trained on much more data than they should have been, while I'd used the theoretically optimal amount of training data. To put it in other words, OpenAI's models were overtrained. If I deliberately overtrained my own models, could I match their performance? This post is a write-up of what happened, but so as not to bury the lede -- it didn't seem to help much, if at all. Let's see why. Let's start by getting a nice crisp definition of overtraining. It's important not to confuse overtraining with overfitting. Overfitting is where you train a model so that instead of learning a general rule about the data it's seeing, it learns something very specific to the training data -- for example, this: ...rather than this: Overfitting is pretty much always a bad thing. Overtraining, by contrast, is more of a judgement call. For LLMs, it's generally used as a shorthand for "training for more than the Chinchilla-optimal number of tokens". The Chinchilla paper makes a very specific case: if you train a model for roughly 20 times as many tokens as it has parameters, then you'll have as good a model as you can get for that budget in terms of compute. They were arguing against contemporaneous experiments where people were doing things like doubling the number of parameters but training on the same amount of data. If you overtrain, it means that you're training on more than 20 tokens per parameter. The Chinchilla argument is that instead of doing that, you should scale up the number of parameters and the number of tokens equally to keep the 20x ratio. Because the amount of compute used scales pretty much linearly with both tokens and parameters, you'll spend the same amount and you'll get a better model that way. Based on that heuristic, if you double your compute budget, then you should scale up your parameter count by 2 and your training token count by the same amount, and by doing that you'll get a better result than you would if you'd naively just doubled the parameters or the tokens, for the same amount of compute time spent. But overtraining is not always a bad thing. Keeping things Chinchilla-optimal means that you have to keep scaling up the model as you scale up the compute budget, and often you can't do that -- for example, let's imagine you're training a model that's meant to run on mobile phones. You have a hard limit on the number of parameters: what will fit in the target devices' RAM. And importantly, in general you will still get a better model by overtraining -- just not as much better as you would have done if you had been able to scale up the model as well as the training tokens. Now, as always, we don't know enough about the original GPT-2 training runs to be sure as to whether they were overtrained, and if so, by how much. But one thing that we do know is that GPT-2 was trained in 2019, three years before the Chinchilla paper came out, so they definitely didn't use it as a heuristic! One thing that they do say in the GPT-2 paper is that their dataset, WebText, is "a total of 40 GB of text". Assuming 4 bytes per token (a good rule of thumb for the GPT-2 tokeniser), that's 10B tokens. If the small model, which had 124M parameters, was trained on all of those, then it definitely was overtrained; 124M times 20 is about 2.5B. 1 On top of that, there's the question of epochs. In my training runs so far, I've been training through a Chinchilla-optimal 3.2B unique tokens. But being able to easily get your hands on that much training data is a relatively new thing, as you can see from the fact that the GPT-2 authors decided to document how they got theirs in the paper. So back then, pre-Chinchilla, people tended to train for multiple epochs so that they could make better use of limited data. Now, when I first started training my models, I tried to dig up some details on the GPT-2 training run beyond what was there in the paper. I found this report , and using data there, I calculated that it looked like OpenAI had trained for about 42 epochs over WebText. The report's author came up with 60 epochs as an equivalent-sized training run for their own dataset. 2 Again, those numbers are shaky; we don't have the real data. But I think it's not crazy to say that the OpenAI models were probably trained on more data than mine, and probably over more than one epoch. They were, in the Chinchilla sense, overtrained. That makes perfect sense for the time, given that it was three years before the Chinchilla paper! But that opened up a couple of interesting experiments that I could try. Obviously I didn't want to train a model over 10B tokens (not least because I'd need to download a larger sample of FineWeb). And I certainly didn't want to do 42 epochs of training, given that one epoch over 3.2B tokens took almost two days, even with my relatively fast JAX code . But it seemed plausible that I'd be able to get at least some signal if I trained longer. I decided to use , my new dedicated training box to train two new models: I expected each training run to take a bit less than four days on . Once they were done, I would be able to evaluate the models, both with the test loss eval and the IFT one. In an unusual-for-me fit of scientific good practice, I decided to write down what I expected to see up-front. I felt that: It was time to find out! I kicked off the extended train, over 6.4B unique tokens, using my JAX code because it's somewhat faster than the PyTorch version. It crashed after about 40 hours, with this error: There were no obvious issues with the machine -- the CPU temperature had been hovering at around 60°C, and GPU at around 70°C, as had been typical in training runs on in the past. There was nothing in or that looked suspicious. On Nvidia's documentation site I found a mention of that error message, but the page in question was all about PyTorch; I couldn't find anything relevant to JAX. For now, I decided to chalk it up to some bug somewhere in the training stack, and only dig in if it happened again. So I kicked the training run off again from the latest checkpoint to see what happened, and 38 hours later, it completed: Note that the numbers there -- tokens seen, time taken, and so on -- are for the portion of the training run after the restart. My checkpoint recovery code doesn't carry those over. The final train loss is just the loss over the period from the penultimate checkpoint to the end of the run -- there must have been some "easy" data there :-) The training loss chart looked like this: As you can see, the latest checkpoint wasn't the "best" one, so I pulled both latest and best down to , my workstation, for investigation. Firstly, I ran them both through my JAX smoke test, which asks them to complete "Every effort moves you" with 20 more tokens (using greedy sampling): Very inspirational. But coherent, and that's what matters. Now, the bulk of my evals use PyTorch rather than JAX -- I've been sticking to that for consistency -- so I converted the model Safetensors files so that they had the right structure to work with that: ...and did the equivalent smoke test on that side (which uses non-greedy decoding with a temperature of 1): I think that the Unicode junk at the end of the second is probably the first half of a two-token apostrophe or something along those lines. Anyway, those looked good. Next, it was time to run the first eval: how would the models perform on the held-back test set? So, the "latest" checkpoint was better than the "best" one. This is a problem with the way I define "best" in my training script. The original version of that script used a validation set to evaluate the model before every checkpoint; "best" at that point meant that the model was the best performer on that validation set. Later on, I decided to play a little loose with my training runs and pull out the validation -- this seemed reasonably safe because I was doing single-epoch training, so what I saw as the primary benefit of regular validation during training -- detecting overfitting by looking for rising validation loss -- was not so important. It's hard for a model to overfit on a single epoch training run. However, doing that meant that "best" had to be changed to mean "best-performing on the training data". And that's actually a really bad metric, because the training data is different for each measurement, so they're not really comparable. So, for this model, I decided that I'd discard the "best" checkpoint, and use the "latest" one. But all of that aside, those numbers were pretty impressive! Not only did it beat my best Chinchilla-optimal model to date, which had got 3.418784 loss on the same eval, and the original OpenAI small model with a loss of 3.499677, but it was actually getting quite close to the OpenAI medium's loss of 3.231442. 3 So now it was time for the two-epoch training run. Adding support for handling multiple epochs to the code was very simple , so having done that, I kicked it off. Just over three days later: This time the numbers are for the full training run -- there were no odd CUDA issues, and it ran straight through. Again, the final train loss is just the loss over the period from the penultimate checkpoint to the end of the run. For runs over the same dataset, it can actually be a useful quick-and-dirty way to compare results before running the full evals, but in this case it's obviously not comparable with the last run's result there, because we're talking about loss on different data. Anyway, once again, "best" and "latest" were different checkpoints, as you can see from the loss chart: ...so I pulled them both to , did the first smoke test: ...converted them to PyTorch: Did the PyTorch smoke test: So that all looked good (despite the Unicode junk), and it was time to work out the loss: That was pleasingly in line with my predictions: Both models would get better results on the test loss eval than my existing ones (90% probability). The one trained on more tokens would be better than the one trained for two epochs on the same tokens (70%). ...though of course the difference between the 3.326482 that this model got and the 3.324953 that the one-long-epoch one got was tiny and probably in the noise. I decided to count it as a win, anyway :-) So now it was time for the big test: how would they do at instruction-following? There are two phases to getting numbers for this test: firstly, I run the script that does the fine-tuning, and then generates the model's completions for the test set using the fine-tuned model. These are written to a JSON file for later use by the LLM-as-a-judge script. That script is the one I fixed in my previous post . I ran it for the extended train (one long epoch) model, taking care to make sure that the config it was passed matched the model's original training run in not having dropout enabled: Next, I ran it for the two-epoch model: With that done, it was time to run the LLM as a judge script . That gave us results, which I've put into the table below. For each model, I've shown the number of epochs of training the model needed before its validation loss started rising. For the pre-existing models, I've shown the score that they got in my baseline evaluation at the end of my last post , and their ranking in that eval. Then for all models, there's the score from this new run, and the new ranking. The two new models are in bold. A note on the numbers: in my previous post, I mentioned that different LLM judge runs differ due to randomness in how "strict" the judge model is. You can see that showing up here -- most of the old/new IFT scores are pretty close, within a point or so, but they do differ, some rising and some falling. This is with exactly the same set of responses for each model going into the judging program -- I used the same JSON files for this table as I did for the previous one for the models that were in there -- the only difference was that the new JSON files for the new models were added. As you can see, the new models scored better than the "JAX, with MHA bias, no dropout" one, which is the most similar: it's the same model config, just trained for the Chinchilla-optimal number of tokens. The one long epoch model got a score that was 1.22 higher than that one, and the two-epoch one scored 0.92 higher. However, my normal rule of thumb for comparing models in these evals is that differences of less than a point or two are probably in the noise. I did a second run of the LLM judge script -- it takes 20 minutes to run and costs a couple of dollars each time, so I don't like to run it all that often -- and this time around (just looking at the JAX numbers) things were a bit closer: You can also see that the two new models have swapped places. So I think that the principled approach here is to say that the improvement is almost certainly in the noise. Perhaps if I did a very large number of runs of the judge I'd get something more solid -- but what I'm hoping for is something a little more unambiguous; some change that really moves the needle in an obvious way, clearly outside the noise. That means that my second pre-registered prediction: Both models would score better than my existing models in the IFT test (70%) but would still be worse than GPT-2 small (90%). ...was half wrong. I got the "worse than GPT-2 small" bit right, at least. But while the models did look a bit better than the most similar pre-existing model, (a) the difference was too small for me to be confident in it, and (b) they were still worse than "JAX, no MHA bias, no dropout" and "Cloud FineWeb, 8x A100 40 GiB". What can we take away from this? The hypothesis that I was trying to test was whether it was simply overtraining that made the OpenAI weights better at this IFT evaluation than mine. Frustratingly, I can't say that the hypothesis was false. Perhaps there is an improvement gained by overtraining -- and the apparent gains which, on this experiment, appeared to be in the noise, would have been consolidated if I'd trained for even longer -- perhaps the 42 epochs on 10B tokens I suspect that the original weights were trained on? But equally, perhaps there is no benefit, and further training would have left the models exactly where they were in the ranking. Intuitively, you'd think that further training of a model would make it better at answering questions, at least up until the point that its parameters were "saturated" and could not absorb new information without forgetting something else. After all, a model that's never seen "Jane Austen wrote 'Pride and Prejudice'" will never be able to successfully answer when it's asked who the book's author was. But where that saturation point might be -- and indeed how much training you'd need to do to get there -- is not obvious. It's an annoying place to finish this experiment, but I guess at least an inconclusive result is better than never having run it at all. And it was at least good to see the test loss improvement that I expected. But given the opportunity cost of tying up in four-day training runs, I think I'll look into other possibilities next. As I was running this experiment, something came up -- and I'll post about that soon. Luckily, this time it won't involve training more base models... It's also worth noting that by the same, um, token, the extra-large model, with 1,542M parameters, was undertrained because the Chinchilla-optimal number of tokens would have been about 31B. Though that said, see later regarding epochs.  ↩ You might wonder whether training over the same tokens repeatedly over multiple epochs "counts" for Chinchilla purposes. Is training on 1.6B tokens over two epochs the same as training on 3.2B tokens over one? " Scaling Data-Constrained Language Models " poked into that in 2023, and from the abstract, came to the conclusion that you could do up to four epochs over the same data without losing much value, but after that returns diminished. If that holds for GPT-2, and they really did train for 42 epochs, maybe they wasted a lot of time? I'll need to read that paper in full at some point.  ↩ There is a table of results later on in this post where you'll be able to compare models easily.  ↩ Firstly, I'd train one on 6.4B tokens from my FineWeb dataset -- the original 3.2B that I had been training on to date, and then on whatever 3.2B came next. Secondly, I'd train another one on the same 3.2B tokens as usual, but I'd do two epochs. Both models would get better results on the test loss eval than my existing ones (90% probability). The one trained on more tokens would be better than the one trained for two epochs on the same tokens (70%). Both models would score better than my existing models in the IFT test (70%) but would still be worse than GPT-2 small (90%). It's also worth noting that by the same, um, token, the extra-large model, with 1,542M parameters, was undertrained because the Chinchilla-optimal number of tokens would have been about 31B. Though that said, see later regarding epochs.  ↩ You might wonder whether training over the same tokens repeatedly over multiple epochs "counts" for Chinchilla purposes. Is training on 1.6B tokens over two epochs the same as training on 3.2B tokens over one? " Scaling Data-Constrained Language Models " poked into that in 2023, and from the abstract, came to the conclusion that you could do up to four epochs over the same data without losing much value, but after that returns diminished. If that holds for GPT-2, and they really did train for 42 epochs, maybe they wasted a lot of time? I'll need to read that paper in full at some point.  ↩ There is a table of results later on in this post where you'll be able to compare models easily.  ↩

0 views

Nix finally has a source-bootstrapped OpenJDK

One of the earliest requested issues we had opened on GuixPkgs was to add more packages, specifically OpenJDK. “Specifically I would like to see openjdk translated. openjdk is not bootstrapped from source code in nixpkgs” [ issue#3 ] Ever since I learned about the stage0 bootstrap chain and how Guix announced full-source bootstrap for all packages in 2023, I was in awe. They provided a package graph of more than 22,000 nodes rooted in a 357-byte program, including the JDK. 1 GuixPkgs can now build OpenJDK 25 🎉 Does it really work? Let’s take it for a spin. lives in the output, since Guix splits the package. is written in Java. is C++, but the class library it needs is Java, and the compiler that compiles the class library is Java, and it runs on a JVM that needs a class library… and so on. This is a bootstrapping problem. Every distribution resolves this the same way in practice: download a JDK and use it to build your JDK. Debian documents the pain , and the Bootstrappable project has a whole page on it. They are fun reads, I highly recommend them. What makes JDK special is that there is no actively maintained JDK that can be built from source without a JDK. The authors had to go back quite a few years to find one, and bring it back to life. Nixpkgs does the same. is built by : a 135 MB prebuilt-tarball. Nixpkgs makes it pretty easy to audit in the meta.sourceProvenance of the package. What does the source provenance of GuixPkgs’ look like? It’s a little whacky but the overall build chain in Guix is the following: Nineteen complete JDK builds, from a C++ program. We can use to emit the entire closure as a dot file to visualize the differences. I restyled the nodes and edges: dots instead of labelled boxes, and the red to highlight the derivations that exist only because of the JDK bootstrap. The upper panel’s entire Java story is those two red dots and the one edge between them: → . The lower panel’s is the 33-node constellation. Why is the Nix graph still so big if I claimed it was from a binary distribution? It turns out that a JDK closure is mostly not Java. It is a large C++ program that wants X11, cups, fontconfig, freetype, alsa and zlib, sitting on a C toolchain that both sides bootstrap from source. That sub-graph is the same in both and it swamps everything. I was curious to zoom-in and see the source-bootstrap portion of the derivation graph. Since that shared sub-graph is drowning the signal, I chose to filter it out. Which derivations exist in this closure only because of how the JDK is bootstrapped? 🤔 That is a reachability query. We delete the Java-provenance nodes from the graph, see what is still reachable from the root, and whatever is left are the derivations that exist only because of the JDK bootstrap. This lines up with our intuition. Nixpkgs only needs two derivations to bootstrap the JDK, while GuixPkgs needs 876 . Note Interestingly, the bootstrap build for JDK has in its closure: IcedTea 8 wants GTK 2, which wants its own Mesa, which wants 😱 What started off as a fun art project , started during TacoSprint 2026 , has now found relevance to those interested in bootstraping and reproducible builds. Nixpkgs has made a lot of progress on this front as well, however it is still not fully bootstrapped and relies on plenty of prebuilt binaries.  ↩ jikes 1.22 (C++) compiles GNU Classpath 0.93 , a free reimplementation of the Java class library. Classpath is enough to bring up JamVM 1.5.1 , a JVM written in C. Now something can run Java. Now we can finally use Java tools: Ant and ecj , the Eclipse Compiler for Java. Then it doubles back and rebuilds all of that with the newer versions to get more modern JDK features. That stack finally compiles IcedTea 2.6.13 (OpenJDK 7) → IcedTea 3.19.0 (OpenJDK 8) → OpenJDK 9 . After which it is one rung at a time: 9 builds 10, 10 builds 11, …, 24 builds 25. Nixpkgs has made a lot of progress on this front as well, however it is still not fully bootstrapped and relies on plenty of prebuilt binaries.  ↩

0 views
Giles's blog Yesterday

Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix

I'm digging into why my GPT-2 style models score worse on an instruction-following eval than OpenAI's original weights; I gave the details in this post . While I was writing up the results of my first experiment into possible causes, I ran the post past ChatGPT -- I always use an "editorial board" of AIs to check my posts for flow, style, and any technical errors (though all writing is always mine). It took a look at the eval code that I was running, and highlighted a bug. Luckily, it doesn't change the important results -- OpenAI's models continue to be better than mine at instruction-following. But it was enough to change the baseline numbers, re-ordering how well my own models did. So I fixed it and regenerated the baseline so that future experiments are based on solid ground. The eval takes a model, and trains it over multiple epochs on a split of a subset of the Alpaca instruction-following dataset. At the end of each epoch, it evaluates the resulting model against a held-back validation split; if the eval loss starts rising, it bails out. Finally, it runs a test split of the dataset through the resulting model, and saves the result. Once I've run it for a bunch of different models, I use an LLM-as-a-judge script to get GPT 5.5 to score results -- for each question-answering result, it sees all of the responses for all of the models in the same prompt , shuffled in order each time, to try to make it judge models against each other as consistently as possible. Now, the idea was that the generation of the test split answers would use the model from the epoch prior to the rising-loss one. So I had code like this: Each time the validation loss went down, I wanted to store the model's parameters in so that they could be restored later in that last line. If you look closely, the error is pretty obvious. does not return a copy of the model's parameters -- it produces a dictionary containing references to the parameters inside the model. So although I was trying to stash away the parameters in each time loss went down, so that we had a copy of the ones that were live in the epoch prior to the rising-loss one, what I was actually doing was just pointlessly saving a reference to the "live" params. The call to was essentially a no-op. The solution was simple enough: That was enough to make the code do what it was meant to do. While I was there, I also noticed that the evaluation code was only using the first five batches in my eval dataset. This code was originally adapted from an eval in " Build a Large Language Model (from Scratch) ", where the eval was run much more frequently and had to be super-fast. Because my own code ran it more rarely, it made more sense to use all of it. Given that I was going to completely re-run the script to generate the test responses, I figured that I might as well fix that at the same time. I re-ran the fixed script on all of the models I'm comparing, and ran that past the GPT 5.5 judge; here's what I got: Let's dig into those numbers. Let's look into the number of training epochs first; I've highlighted the models for which it changed. For the "Cloud FineWeb, 8x A100 40 GiB" model, I think that was a result of the change to the number of validation samples. For the two JAX models, it's a bit more of a mystery. There may be some part of the same "more eval data" variation there, but I have a suspicion that there's something more, related to dropout. My methodology to date has been that the IFT training runs should use the same dropout setting as the original base model training run (I'll come back to this in a later post). However, due to differences between the JAX training script and the evaluation code (which is PyTorch), matching the dropout rate when evaluating those models is a bit fiddly and error-prone. I am 100% sure that I got it right this time around -- I've checked and double-checked the specific commands I ran. But I think -- from what I remember doing -- that I might have messed it up the previous time. That's not certain, though -- just a suspicion based on half-memories of some commands I ran several weeks ago, which never wound up in my (too many terminals open at once). For the IFT scores, remember that they're not strongly comparable between runs. Let's imagine that the LLM judge is given the following answers to the question "Who was the author of Pride and Prejudice ?" In some cases it might treat the first two as being 100/100, in others it might give the second 95/100 for being too wordy. Likewise, in some cases it might rank the last two as 0/100 for being wrong, in others it might give the "Sarah Palin" one 5/100 for at least being the name of a person rather than complete nonsense. Now, we always ask the LLM to judge all of the models' answers for a given question in one go, so at least we can be sure that it will be consistent for a given prompt about a given question. But what we can't keep consistent is which way it leans between different runs, or different questions within the same run. Sometimes it might be feeling "generous" and give the Sarah Palin answer a bit of grace, other times it might be harsher. So there's a significant amount of noise there; my rule of thumb is that a variation of a point or two is within that noise, so OpenAI medium going from 41.62 to 42.41 is pretty much meaningless, and likewise "JAX, with MHA bias, no dropout"'s move from 19.25 to 18.12. So what's important is the relative ranking -- which is first, which is second, and so on. Naturally, you have to allow for the fact that -- for example -- if one model goes from position 11 to position 3, like "JAX, no MHA bias, no dropout" did, then the previous number 3 will have to become number 4, 4 will become 5, and so on. You can see that happening in the results table. Obviously, as other models rise and fall, that has further knock-on effects. Anyway, with all of those caveats, the good news is that my original mystery remains. The OpenAI models were still doing noticeably better than my own ones. GPT-2 medium continued to lead the pack (unsurprisingly, given that it's a bigger model), and GPT-2 small was still in second place. If that had changed, it would have made a rather disappointing end to this series: "mystery explained, it was a bug in the eval :-(" But now let's look at my models. Firstly, it looked like the "no MHA bias" models might have benefited from the extra training -- or from having their dropout settings corrected. They rose from positions 11 and 15 to 3 and 8 respectively -- a huge swing for "JAX, no MHA bias, no dropout", and a solid improvement for "JAX, no MHA bias, with dropout". Most of the other changes in relative rankings can be explained by those two models having been promoted, but there are some other changes. In particular, three models dropped significantly in score (and, as a result, ranking): My suspicion -- a very weakly-held hypothesis, but an interesting one -- is that those models had previously been benefiting from the bug. Remember that the signal we're using to stop training is that the validation loss starts rising. We're using that as a proxy for overfitting, which in turn we're using as a proxy for "this model has had as much training on this data as it needs for the eval". But there's no guarantee that the connection is there. Perhaps the models weren't overfitting and if we'd waited for another epoch or two, the validation loss might have started falling again. Or maybe some amount of overfitting would be beneficial for this eval? There's probably a near-infinite amount of digging in that I could potentially do here. But I think it's best to stop. The bugfix was important because it meant that the eval was now doing what I thought it was doing. Importantly, it doesn't change the puzzling fact that my models were worse at this eval than OpenAI's, which is what I'm trying to untangle. And it means that I can now lean more confidently on the baseline numbers. So now it's time to actually start changing things to see if I can close the gap! Here's a link to the next post in this series: does overtraining help? . "Jane Austen" "The author of 'Pride and Prejudice' was Jane Austen" "The author of 'Pride and Prejudice' was Sarah Palin" "The author of 'Pride and Prejudice' was 'Pride and Prejudice'" "Cloud FineWeb, 8x B200 160 GiB" "Local FineWeb train"

0 views
Chris Coyier Yesterday

CodePen 2.0

Noting perhaps my largest personal career accomplishment, which is launching CodePen 2.0 . Far more work, believe it or not, than the entire creation of the original CodePen. This isn’t the place to describe every detail of what we did and why we did it. If you’re interested, perhaps our Why 2.0? podcast or the What’s New? page. Instead, a couple of stories from the first week of launch. I was working on a demo with someone I’ve never met before. It started on their (classic) Pen. They needed to import some other JavaScript, so they used 3 Pens and pulled in the JavaScript from the other two into the main demo. They also needed an npm package. I forked the Pen and invited them as a co-editor, so we could both work on it together anytime. I moved the JavaScript into files on the main Pen, as that’s much easier to work with. The npm package is in the file for easy version management. We both cleaned it up to our liking. The Keyframers (David and Shaw) got back together and did a live stream on launch day. They also used the invite feature and live collaboration . They worked together for hours, and while there was a bug or two, it was nothing super major, and it went great. One of my favorite bits was that they shared the Live View of the Pen, so as they were working on it, we could play with the demo ourselves. As I was working on the emails we were going to send out about the launch, I was building them in the special language for crafting them: MJML . I went ahead and added MJML as a block to CodePen so I could just build them right in CodePen. Works great , even for weird stuff . Many more Blocks to come. I friggin love how I can make little websites and deploy them right through the Pen Editor. Like the one for our slideVars library or codepen.school . It just makes me wanna build a ton of little weird websites.

0 views
Kev Quirk Yesterday

Toot.community is shutting down

by Jorijn Schrijvershof Sadly, toot.community is shutting down in October. This post talks about why Jorijn has made that decision. Read post ➡ So much of this post resonated with me and aligned closely with my reasons to step away from Fosstodon . Being a fediverse server admin is a thankless job and as Jorijn says: I find myself feeling more angry or discouraged after spending time as an admin, and that’s not what I want from social media or a volunteer project. Over a year on since I stepped away from the Fosstodon team, I'm much happier now, but the experience has tainted my experience of social media. These days I probably visit Fosstodon maybe once per week. Every post I write is written here, on my site, and syndicated over there. I feel this is much better for me personally, as it keeps me away from the latest drama. When I do go visit the timeline, it's usually fun because it's fleeting. I don't have time to get involved in the latest drama. I go, check notifications, quickly peruse the timeline, and leave. It's a much healthier way of engaging, I think. More broadly, I think the issue with admin and mod burnout on the fediverse is what will limit it in the long run. It's almost entirely propped up by volunteers, who don't want to be involved with the drama, or be attacked for making a decision you don't agree with. And as a result, servers like toot.community fall by the wayside. I don't know what the solution is here - I actually don't think there is one. It's a fundamental limitation of the way in which the fediverse is architected and as a result, another great server is gone. 😔 Thanks for reading this post via RSS. RSS is ace, and so are you. ❤️ You can reply to this post by email , or leave a comment .

0 views
Unsung Yesterday

“Next up? You know it. You hate it. It’s the invisible wall.”

Tangentially related to the previous post , this 8-minute video from CrowBranch is a fun tier list of various ways games try to keep you in bounds: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/next-up-you-know-it-you-hate-it-its-the-invisible-wall/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/next-up-you-know-it-you-hate-it-its-the-invisible-wall/yt1-play.1600w.avif" type="image/avif"> The limits here are I believe unrelated to precision, and rather to wanting to keep the player within a controlled, designed environment. (We once covered a story of this backfiring horribly , and also talked about skyboxes and streaming .) The video also made me want to see people – especially designers with less experience slash less baggage than I have – doing a tier list of e.g. common interface elements. Or maybe I should make such a video… #errors #games #youtube

0 views
Unsung Yesterday

“This is a sphere getting exponentially angrier as time passes.”

When you learn computer science, at some point you encounter floating-point numbers in all their peculiar glory. Floating-point numbers allow you to store values that are extremely huge or extremely tiny, but that comes with a strange price. The number is still just a number with a regular limited precision, but is accompanied by a floating point – another number that just says how many zeroes to put after or in front. In effect, when your number is millions, you get a precision of thousands. When your number if billions, you can only count in millions, and so on. This makes intuitive sense, in a same way a billionaire doesn’t care about pocket change. (Go to a floating point visualizer , and you will see you can input 9000004 and 9000005 and 9000006 with no problem, but add one more zero and 90000004 will be rounded down to 90000000, and 90000005 up to 90000008. Bigger numbers get even less precise.) This has a strange effect when it comes to videogames or graphic software: The further you move something away from the origin of your universe – where X and Y are 0 – the more you have to worry about its internal precision. This is not a problem in real world. You can send a tiny, incredibly meticulous watch screw far beyond the solar system and back, and it will do fine. But now jump to the digital world, assume Earth is 0:0 – apologies to Copernicus – and now the further away you go, the more the floating point moves to the left, and at some point the precision of the number on the other side runs out; a value that can express billions of kilometers is too coarse to even consider millimeters. This is a fascinating problem which is visualized nicely in this short video . This is what happens when you move an object with its constant internal precision further and further away: = 2x) and (width >= 700px)" srcset="https://unsung.aresluna.org/_media/this-is-a-sphere-getting-exponentially-angrier-as-time-passes/yt1-play.2096w.avif" type="image/avif"> = 3x) or (width >= 700px)" srcset="https://unsung.aresluna.org/_media/this-is-a-sphere-getting-exponentially-angrier-as-time-passes/yt1-play.1600w.avif" type="image/avif"> There are ways around this – you can increase the precision, introduce a “floating offset,” have two precision centers around two floating points, or do a bunch of other things – but each one will cost you. So, sometimes, the best solution is the simplest one: do not allow the numbers to get too big. And this is why Adobe Illustrator, Figma, and I bet at least a few other tools that promise infinite canvas, actually stop you if you stray too far away from the center. Precision is like atmosphere, the canvas says, thinning out the further you go. We cannot promise your structural integrity will survive going past a certain point, so we won’t allow you to do that. Here’s Figma, where I at some point the canvas stops following your scroll commands: The first invisible X value is 131,072 and it’s not a surprise it’s one of these numbers that will look very, very familiar. It seems like a huge enough value, but if you consider a slide in a slide deck is 1920×1080, it’s not as infinite as it might initially seem. And this is also why – in part – Figma Slides manually wraps you after 20 slides: #graphics #youtube

0 views
Martin Fowler Yesterday

The Economic Benefit of Refactoring

Giles Edwards-Alexander does an experiment to see if decomposing a large function helps reduce token costs, suggesting that is may now be possible to measure the economic benefit of refactoring

0 views
Farid Zakaria Yesterday

Guix by Nix

I have been working more on GuixPkgs in preparation for a talk at nix.vegas for DEFCON34 . At the end of my previous GuixPkgs post I left a teaser: We can then build a NixOS machine where every package is the Guix equivalent 😱. Well, adeci 1 took the bait and we went even further than that. 😈 Say hello to Guix by Nix : a bootable VM where the kernel is Guix’s Linux-libre, the userland is translated Guix packages, and PID 1 is GNU Shepherd . No systemd. No D-Bus. No NixOS activation. Not even a or binary in the guest. It’s Guile all the way down and Nix built all of it. Log in with / and poke around: As a reminder, even though these binaries live in , they are not Nixpkgs packages. They were translated from Guix derivations using guix-transfer and built by the . That was built from Guix’s package definition, source bootstrap and all , but it lives in , because built it. If you want to try out any of these packages on your own machine, you can use GuixPkgs . Tip Guix offers all packages built from source where Nix may offer it as a prebuilt binary. You can use GuixPkgs to get a source-built bootstrapped version of OpenJDK for example and all the whacky steps to get there. How does this actually work? The project is a three-stage pipeline: Warning AI was leveraged to write the initrd and activation scripts. That seems to trigger people lately, so consider yourself warned. Every program that can be executed, every ELF file, every script interpreter, all traces back to a translated Guix derivation. The only things Nix authored are text files: the init scripts, the Shepherd config, . This is Guix by Nix . “Every executable byte comes from Guix” is exactly the kind of claim that’s easy to say and easy to fudge. 🤥 A booted demo is great but that can’t prove it. A VM where secretly came from Nixpkgs boots identically . The flake ships an audit-check that classifies every store path in the shipped closure by provenance, unpacks the compressed initrd 2 , and inspects every executable payload. The audit derivation fails if anything is unclassified, any ELF file or script interpreter doesn’t trace to a translated Guix output, any reference survived translation, or anything systemd-shaped appears anywhere. The report includes a lot more information such as the exact Guix channel commit everything was translated from. There’s also some NixOS VM tests for good-measure. 🕵 Note The kernel is Guix’s , translated and built by like everything else. Nix wraps it in a thin adapter so NixOS’s VM tooling accepts it by augmenting with some additional metadata only; the is byte-for-byte Guix’s. One obvious next step would be a normal NixOS machine, where every package that exists in GuixPkgs shadows its Nixpkgs equivalent on . GuixPkgs offers such an overlay, but it needs quite a lot of CPU to build all of it…. A more realistic use case is to mix and match GuixPkgs and Nixpkgs packages in a single system. Get the best of both. I heard there were people in the Nix community who still want a non-systemd system, à la sixos . 🫠 For now, this project remains a minimal VM demo. If you make it PID 1 on real hardware, please send photos. 🙇 It has been extremely fun and rewarding to work with Alex on this project. We nix-pilled him at PlanetNix 2 years ago and since then he has been pushing the boundaries of what Nix can do and is currently employed at Shopify working on Nix.  ↩ Nix’s reference scanner can’t see inside archives and sneaky paths hide there.  ↩ guix-transfer is the tool that translates Guix derivation graphs into Nix derivations. GuixPkgs is the flake with all Guix packages, all built from the 357-byte seed, although there is a Cachix cache provided. Guix by Nix assembles ~42 of those packages, a subset of the overall set, into a useable machine. The initrd is a custom shell script, interpreted by Guix Bash, that loads eight modules, mounts the root disk and the 9p store, and s. Activation (accounts, , the setuid copy) is another custom Nix-generated script run by Guix Bash. PID 1 is from Guix, with a small Scheme config that starts , , , and a serial . works like you’d expect. is compiled from Guix’s own search-path specifications, so and friends point where Guix intended. It has been extremely fun and rewarding to work with Alex on this project. We nix-pilled him at PlanetNix 2 years ago and since then he has been pushing the boundaries of what Nix can do and is currently employed at Shopify working on Nix.  ↩ Nix’s reference scanner can’t see inside archives and sneaky paths hide there.  ↩

0 views
Maurycy Yesterday

You have been mislead about lightbulbs

There's a story that goes something like this: In 1925, lightbulb manufactures secretly colluded to standardized lifespans at 1,000 hours. They would test each other's products to ensure compliance. This is true. At the time, many bulbs lasted longer than one thousand hours. This is true. Therefore, this was done so that people would always need to buy more light bulbs. This is wrong , but it's the type of wrong that cites its sources and hides in the part you'd never think to fact check: The assumption that a longer lasting lightbulb is a good product. In truth, increasing the lifespan of a bulb makes it worse in every other way... but people think they want a long lasting lightbulb: the purpose of standardizing was (primarily) to avoid a race to the bottom. A old-school lightbulb is a rather simple device : A thin tungsten wire (~20 μm) sealed inside a glass envelope to protect it from air. When current is applied, the wire gets white hot and starts glowing. The most important parameter of a lightbulb is how hot that wire gets: this controls the peak emission wavelength (color) and brightness of the lamp. Room temperature objects do emit light (this is how thermal cameras work), but it's at the ~10 μm range instead of the 400 nm - 700 nm light that we can see. In order to put the emission peak in the visible spectrum, the filament would need to run at ~5700 °C ... that is, the temperature of the sun . No metal can survive these conditions: Tungsten melts at "only" 3422 °C. Since it has the highest melting point of any metal, tungsten is the obvious choice for filaments. However, the metal is quite brittle and drawing it into a wire isn't easy. The first commercialized lamps used carbon filaments that were made by charring plant fibers. However, the carbon would evaporate at fairly modest temperatures ~2000 °C. Tantalum filaments were briefly produced during the 1900s, because the metal was easier to draw into a wire than tungsten. These were the first lightbulbs that could actually be left on at night, although they were quickly replaced with tungsten manufacturing improved. There ware also some experiments using zirconium dioxide ceramics, which become conductive when heated. These allowed lamps to operate in air (obviating the need for a vacuum pump and glass seals ), but were limited by its melting point of 2,700 °C. Since any filament must run below its melting point , the peak emission is always in the infrared. This means that only the extreme high-energy edge of the spectrum is useful for illumination, so a small increase in temperature will make a lamp orders of magnitude more efficient. Also, since this increases the average energy of the atoms, the lamp is able produce shorter wavelengths: resulting in a whiter and less depressing glow. The snag is that when a metal is close to its melting point, the atoms are barely holding together: A hot tungsten filament slowly falls apart as the metal crystals slide past each other. Additionally, atoms can evaporate from the surface until there's no wire left. The rate of both of these processes increases with temperature, so there's a fundamental trade off between color/efficiency and lifespan. The lightbulb everyone always cites in the story is hanging in a California fire department. It's been running nearly continuously for over 120 years and racked up over a million hours of operation. Impressive right? What almost no one talks about is that it's hardly even glowing! Photo taken by Wikipedia user Rjaerial Despite nominally being a 60 W lamp, it draws only 4 watts... and is a lot dimmer than you'd expect from a 4 W lamp due to its poor efficiency. There isn't any documentation, but in all likelihood, the bulb was made wrong and ended up having a very high filament resistance. That's why it was sold for as a night light, because it wasn't usable for anything else. Early bulbs (like that one) were handmade , and quite expensive. Because of this, there were universally optimized for long lives. This resulted in light isn't anywhere near white, and a an efficiency that was a tiny fraction of a modern incandescent lamp. (which are also terrible by any objective standards) Once the production process was automated, new bulbs cost pennies, so it made sense to optimize them to work well... because less efficient bulbs cost more money to operate: Going off modern day prices, electricity costs around 0.10 [$/kW*h], so a 60 W lamp will consume 6$ of electricity over a 1,000 hour lifespan. Considering that such a lamp only costs around 3$, installing one that lasts longer but uses more power would be silly. Case in point , despite the cartel only lasting for 14 years, modern (non-halogen, incandescent) bulbs still last for between 500 to 2,500 hours, a range that includes the cartel's 1,000 hour standard. Instead of "making bulbs last longer", manufacturers spent huge amounts of time and money developing entirely new technology: fluorescent and LED lamps. Because these don't use a wire on the very edge of melting, they can be made to work well and last a long time. Of course, specialized lamps have different requirements: In photography, a truly white light is desirable, which leads to specialized "photoflood" bulbs that only last for a few hours. In the other direction, many indicator lamps are designed for 100,000 hours because they are difficult to replace. Ok, but what's with the testing ? If long lived light bulbs are worse products, why would they need a cartel to enforce a the thousand hour limit? Well, it's because people think long lasting bulbs are a good product: Lifespan is something everyone can understand, and has a direct effect on when you will have to go back to the store: if you saw two 60 W bulbs in a store, one claiming to last 400 hours and the other 2,000 hours, you'd probably get the longer lasting one without thinking about it. It's not that efficiency is hard to understand, but most people aren't doing homework before buying lightbulbs... and it doesn't help that bulb packaging uses input power as a proxy for brightness, so the idea that two bulbs both labeled as "40 W" would have a different brightness is rather confusing. As a result, competition was forcing lightbulb makers to produce worse products. To be clear , I'm not defending the Phoebus cartel: they absolutely engaged in price fixing and other anti-consumer practices, and it's difficult to imagine that profit wasn't a factor when deciding the 1,000 hour standard... but by nature, tungsten lamps are consumable items. I guess the the moral here is reality rarely fits into nice stories. Even something so obvious like "products designed to break are bad" often isn't — every manufactured object is the result of hundreds of overlapping compromises, most of which are invisible to the end user. Also , to preempt the orange site, I'm not saying that planned obsolescence doesn't exist. There are plenty of actual cases of products being made hard to repair so they can sell you another one. ... but lightbulbs aren't a good example. https://www.mouser.com/datasheet/3/299/1/T_1_Wire_Terminal.pdf : Indicator lamp datasheet featuring a 100,000 hour rating. https://www.1000bulbs.com/product/67291/STAG-PH213I.html : A photography lamp that lasts 3 hours. (store page) https://www.youtube.com/watch?v=zb7Bs98KmnY : An excellent youtube video on this topic. https://doi.org/10.1063/1.1657874 : Lab tests of tungsten wire evaporation https://doi.org/10.1016/s0016-0032(25)91062-9 : Brightness and color of light as a function of temperature.

0 views
DYNOMIGHT Yesterday

So you want to use plants to reduce indoor CO₂

Humans make carbon dioxide. Carbon dioxide is bad for cognition. But plants turn carbon dioxide back into oxygen. And plants are the one true home decoration strategy. So maybe if you get a lot of plants, you can you can keep carbon dioxide in check and keep your brain working? It’s theoretically possible. It’s probably just barely possible in practice. But it won’t be easy. People produce ~1 kilogram of carbon dioxide per day. That’s around 5.7 × 10²³ molecules or 0.948 moles per hour. (You may remember from high school that a mole is a gigantic number made up to avoid having factors of 10²³ everywhere.) Let’s keep it simple and call it one mole per hour. Meanwhile, plants turn carbon dioxide into oxygen through photosynthesis, i.e. the chemical reaction of (6 water molecules) + (6 carbon dioxide molecules) + (energy) → (1 glucose molecule) + (6 oxygen molecules). The minimum energy physically needed to convert 1 mole of carbon dioxide into glucose and oxygen via this reaction is ~477 kilojoules. So we’ve already got a lower bound. Say you have magical plants that somehow channel all incoming energy into photosynthesis with perfect efficiency. They’ll need ~477 kilojoules per hour, which converts to a continuous usage of 132.5 watts. 1 That’s a bit more than what’s used by two incandescent light bulbs, which isn’t too bad. But you don’t have magical plants. Real plants do photosynthesis through a physical process with two steps, each of which involves four electrons absorbing a photon. That means you need eight photons per carbon dioxide molecule. If you want to tune your lights for maximum efficiency, you should give each photon exactly the minimum energy necessary to excite an electron, which happens to be ~1.8 eV. That corresponds to pure red light with a wavelength of 680 nm, and a continuous usage of 386 watts. 2 No physical system using chloroplasts can neutralize your CO₂ using less than that. Somewhat high, but still manageable. But your houseplants won’t be able to grab every single photon that hits them and direct it towards photosynthesis. In practice, ~30% of photons will reflect off the plant, or go through it, or hit some part of the plant other than the chloroplasts. That brings us to 551 watts. 3 And there’s another issue. After plants make glucose, what happens to it? Some is used to grow more plant, which permanently sequesters carbon from the environment. But lots is also burned by the plant for the general business of staying alive, releasing the carbon back into the air. The exact amount burned in this way varies based on species and conditions, but around 40% loss reasonable, 4 bringing us to 918 watts. 5 That doesn’t sound that bad. But have you considered what it would be like to live in the same room with 918 watts of pure red light? In terms of radiant power, that’s the same as produced by ~765 incandescent lightbulbs. 6 Modern LED grow bulbs are ~50% efficient, meaning you’ll actually need to spend ~1836 watts. If you’re imagining plants that you can actually see, adjust that upwards again for all the light lost to the room. And if you want to use normal light frequencies instead of living Red Life, then your LED bulbs will be less efficient at creating light and your plants will be less efficient at capturing it. Realistically, we’re talking about something like 5,000-10,000 watts, most of which is lost to the room as heat. Imagine five space heaters blasting you on high all the time. But maybe you’re OK living in a tanning booth. Or maybe you’ll keep your plants in a perfectly reflective chamber. Or maybe your house has a glass ceiling and infinite free sunlight and free climate control. That’s cool. But have you forgotten about your old friend, photosynthetic photon flux density? Plants can’t absorb infinite amounts of light. Chloroplasts take time to “reset” before they can absorb more photons. Your pet fern can only absorb ~52 watts of energy per square meter of leaf surface area. 7 So no matter how much light you can produce , if you want to neutralize the carbon dioxide you make, you will need at least 918 / 52 = 17.6 square meters of fern leaf. Picture a 4.2 meter square wall, packed solid with ferns. If there are any gaps, stems, soil, or wall showing, it needs to be even larger. That’s the absolute minimum. But maybe that still sounds OK? Fine. But consider one last barrier: Plants obey the laws of physics [citation needed]. If they remove carbon from the air, they must put that carbon somewhere. The only place it can go other than back into the air is into the plant itself. The 1 kg of carbon dioxide you produce each day corresponds to 273 grams of elemental carbon. The only way for a plant to hide that is by making more plant. But dry plant matter is only ~50% carbon, and for each gram of dry plant matter, plants have 5-10 grams of water (varying a lot by species). So in order to sequester all the carbon you make, each day you will need to grow around (1 kg carbon dioxide) × (0.273 kg elemental carbon / kg carbon dioxide) × (2 kg dry plant / kg elemental carbon) × (8.5 kg actual plant / kg dry plant) = 4.6 kg actual plant. Your garden must grow that much, every day. That’s 140 kg per month. You must prune and discard all that outside, or your garden is not actually sequestering anything. In conclusion: Behold the power of arithmetic : (1 mole CO₂ / hour) × (477 kJ / mole CO₂) = 132.5 watts.  ↩ Again using the power of units: (1 mole CO₂ / hour) × (8 photons / CO₂ molecule) × (1.8 eV / photon) = 385.94 watts So chloroplasts are at most ~34% (132.5 / 385.94) efficient at channeling the energy in light into photosynthesis.  ↩ I find this 30% number amazingly low. (Well done, evolution.) And perhaps it should be somewhat lower. For one thing, the 30% figure comes from sunlight filtered to the 400-700 nm range. If you’ve got pure 680 nm light, absorption should be somewhat higher. Also, if photons are absorbed by some part of the plant other than the chloroplasts, they become heat and the energy is gone. But if they’re reflected or go through the plant, then they might go on to hit some other plant (provided you have a lot of plants around). If you really have pure 680 nm light and you have very densely packed plants, maybe you could drop this to 10-20%.  ↩ Wikipedia quotes a 35-45% loss just for respiration in the leaf itself. But then this paper shows numbers ranging from 30% to 56% depending on the species and growth rate.  ↩ I’ve estimated an overall efficiency of 132.5 watts / 918 watts ≈ 14.4%. If you go to Wikipedia, it estimates that ideal leaf efficiency with sunlight is only around 5.4%. That’s because sunlight contains a wide band of wavelengths and my calculation assumed an ideal 680 nm source. Around 47% falls outside the 400-700 nm range, and inside that range, around 24% is lost due to higher-energy photons with energy that gets wasted as heat. If you account for that, my estimate becomes 14.4% × (1-0.47) × (1-0.24) = 5.8%, which is close enough for government work.  ↩ A traditional “60 watt” incandescent lightbulb is rated based on the power input . But only around 2% of that energy is actually converted to light. So 918 watts of pure red light isn’t what you get from 918 / 60 = 15.3 lightbulbs. It’s what you get from 918 / 60 / .02 = 765 lightbulbs. That said, your eyes aren’t very sensitive to 680 nm light, so the perceived lux wouldn’t be nearly so bad.  ↩ The saturation point of plants is usually given in units of 300 μmol/m²/s. That the number of photons (in micromoles) that can be absorbed, per square meter of leaf, per second. A typical value for a shade-tolerant houseplant would be ~300 μmol/m²/s. If we assume again that the light is 680 nm so that each photon carries 1.8 eV of energy, then ~300 μmol of photons carries 51.92 joules. That’s 51.92 joules of energy per square meter of leaf surface, i.e. 52 watts.  ↩ Build an industrial indoor farm. Wait two weeks. Weigh it again. Divide the increase in weight by your own body mass. That’s the fraction of your CO₂ that you’re removing from the environment. Open a window. Behold the power of arithmetic : (1 mole CO₂ / hour) × (477 kJ / mole CO₂) = 132.5 watts.  ↩ Again using the power of units: (1 mole CO₂ / hour) × (8 photons / CO₂ molecule) × (1.8 eV / photon) = 385.94 watts So chloroplasts are at most ~34% (132.5 / 385.94) efficient at channeling the energy in light into photosynthesis.  ↩ I find this 30% number amazingly low. (Well done, evolution.) And perhaps it should be somewhat lower. For one thing, the 30% figure comes from sunlight filtered to the 400-700 nm range. If you’ve got pure 680 nm light, absorption should be somewhat higher. Also, if photons are absorbed by some part of the plant other than the chloroplasts, they become heat and the energy is gone. But if they’re reflected or go through the plant, then they might go on to hit some other plant (provided you have a lot of plants around). If you really have pure 680 nm light and you have very densely packed plants, maybe you could drop this to 10-20%.  ↩ Wikipedia quotes a 35-45% loss just for respiration in the leaf itself. But then this paper shows numbers ranging from 30% to 56% depending on the species and growth rate.  ↩ I’ve estimated an overall efficiency of 132.5 watts / 918 watts ≈ 14.4%. If you go to Wikipedia, it estimates that ideal leaf efficiency with sunlight is only around 5.4%. That’s because sunlight contains a wide band of wavelengths and my calculation assumed an ideal 680 nm source. Around 47% falls outside the 400-700 nm range, and inside that range, around 24% is lost due to higher-energy photons with energy that gets wasted as heat. If you account for that, my estimate becomes 14.4% × (1-0.47) × (1-0.24) = 5.8%, which is close enough for government work.  ↩ A traditional “60 watt” incandescent lightbulb is rated based on the power input . But only around 2% of that energy is actually converted to light. So 918 watts of pure red light isn’t what you get from 918 / 60 = 15.3 lightbulbs. It’s what you get from 918 / 60 / .02 = 765 lightbulbs. That said, your eyes aren’t very sensitive to 680 nm light, so the perceived lux wouldn’t be nearly so bad.  ↩ The saturation point of plants is usually given in units of 300 μmol/m²/s. That the number of photons (in micromoles) that can be absorbed, per square meter of leaf, per second. A typical value for a shade-tolerant houseplant would be ~300 μmol/m²/s. If we assume again that the light is 680 nm so that each photon carries 1.8 eV of energy, then ~300 μmol of photons carries 51.92 joules. That’s 51.92 joules of energy per square meter of leaf surface, i.e. 52 watts.  ↩

0 views
Theia Yesterday

Coming of a New Sun

Six weeks after the US and China hammered out the latest round of bilateral compute agreements in a midnight deal, reporter Jenny Gesteson visits a south Texas 'dark factory' to see how the machines there - and the minds that operate them - are learning to run themselves.

0 views
Grumpy Gamer Yesterday

Thimbleweed Park 2

I have good news and bad news. First the good news. The good news is that we just started production on Thimbleweed Park 2, due out in early 2028. We will self-publish with the help of a private investor. Mark Ferrari, Gary Winnick, David Fox, Octavi Navarro, Robert Megone, and others from the original team will be back! Wishlist on Steam! I’ll be starting up a Thimbleweed Park 2 dev blog like we did for Thimbleweed Park and posting regularly to keep everyone up to date. Now the bad news There is no bad news, it’s all good news. P.S. Thimbleweed Park 1 is on sale on Steam, Switch, iOS and Google. But I suspect all my readers already own it. P.P.S Steam doesn’t show it yet, but there will be Mac, Windows and Linux. P.P.P.S There will be a GOG version.

0 views
Jim Nielsen 2 days ago

The AI Aesthetic

Every zeitgeist comes with new design idioms unique to its challenges. Many of them disappear as fads change, but others bake themselves into deeper parts of existing software interaction paradigms. For example, there’s the hamburger menu (≡) which saw a proliferation during the rise of mobile due to the constraints around screen size. It has since spread to many other parts of software interaction design and will likely remain prevalent for a long time as a terse way of indicating “more menu-type content here”. As another example, before AI what were the connotations of the sparkle emoji ✨? Personally, I don’t know, but now it means AI. (AI = sparkles and rainbow colors — it’s funny when you think about it. They should’ve just thrown unicorns in there for the trifecta. AI = sparkles, rainbows, and unicorns ✨🌈🦄. Apt.) Some patterns are very specific to the interactions inherent to the nature of AI as a technology. For example: streaming text. This is a pattern made for and refined by chat interfaces, so it may not have tons of utility for reuse across other software interaction paradigms. Then there are other patterns that’ve been refined by AI interfaces and are starting to spread to other places in software. For example, the “shimmering text” which in AI land implies a kind of “thinking” but is being repurposed to indicate any kind of asynchronous task (thinking, fetching, computing, etc.). Then there are other influences my subconscious is picking up on. For example, a lot of AI apps use tiny icons. These are most obvious (to me) in desktop Electron apps because they clash with the system-level grain of applications . Take a look at this screenshot, where you have desktop AI apps on the left (Claude, Codex, Cursor) and macOS apps from Apple on the right (Finder, Photos, Mail). You can see how the AI apps all have much smaller, thinner icons than their native counterparts. Are tiny icons our collective future in interfacing with computers? (Personally, I hope not.) There are other aesthetics my brain associates with AI, like beige/cream colors, orange accents, and serif typefaces as well as whack-a-mole UI controls (you know, the ones where you click the toggle and the entire UI repaints and you have to move your mouse somewhere else in the UI to click the toggle again? The non-determinism of AI’s grain has seeped into its UI/X). It all makes me wonder what other aesthetics are being born out of this AI moment and how many will spread, take seed, and become part of common software interaction paradigms for years or decades to come? Reply via: Email · Mastodon · Bluesky

0 views