223 points | 2d ago | Discuss on Hacker News | Back to Radar
People say that agentic development is great because you can churn out so much so fast. But that doesn't mean that any of it will be truly good and reliable.
The things that are truly insightful and solid end up being used exponentially more, which makes the linear cost of extra development time (asymptotically) insignificant in the cost/benefit equation.
An agent can synthesize existing solutions but (because I see this failure mode at work constantly) it can't synthesize an architecture to resolve the sorts of tensions that the person prompting it doesn't yet understand (not that that is stopping anyone). You can't prompt it to build something if the operational primitives required to solve the problem haven't been mapped.
"Build a tool based on Scribe and Web Scrapbook" in 2003 would've made a fragile PHP wrapper because that's what the existing landscape looked like.
"But the math proofs," people will say. A lot of those seem to be spam-solving things with a huge swath of existing lemmas, and a some of these are being debunked and retracted.
Just today I was quizzing ChatGPT about a basic grammar question for a language that has huge training data but for which the grammar was not well documented. It kept giving me confidently wrong answers until I drilled and drilled it and then finally it found/gave back an explanation that perfectly fit a pattern given in one particular grammar, citing that as a source. It doesn't appear to have been able to figure out the inner structure on it's own. It appears only able to pattern match and put things together from what humans have already discovered and written.
Claiming we "solved" statistical interpolation long ago just means curve-fitting and basic regressions on structured data. Transformers are a truly impressive achievement, scaling all of this to unstructured high-dimensional text topologies, but it's fundamentally the same math operations on statistical proximity.
Like how do we explain hallucinations here? Tokens that are hallucinated are semantically "close" in that vector space but they're completely false in reality. If LLMs operated in a true semantic space they wouldn't hallucinate CLI flags that don't exist.
I propose that concepts of "concurrency" and "contention" are themselves vector in latent space. All concepts are. Recall that we're talking about a 10^4 - 10^5 dimensional space. You can fit in pretty much any conceivable association as some direction in there.
And try to zoom in on any concept you know. If you do, it should quickly become apparent that there's never any concept you can give a closed definition for. We can only define concepts, and we can only learn them, through generalizing from examples. Which is conceptually (pun not intended) regression - finding a vector along which examples live.
> My second warning remark is that I shall refuse to discuss the academic enterprise in financial terms. The first reason is that the habit of trying to understand, explain, or justify in financial terms is unhealthy: it creates the ethics of the best-seller society in which saleability is confused with quality. The other day we had to discuss the professional quality of one of our colleagues, in whose favour it was then mentioned that one of his Ph.D.s had earned lots and lots of money in the computer business, and few people seemed to notice how ridiculous a recommendation this was. We also know that the financial success of a product can be totally independent of its quality (as everyone who remembers for instance the commercially successful IBM360 should know). The second reason for my refusal is that the value of money is a very fuzzy notion, so fuzzy in fact, that efforts to understand in financial terms always lead to greater confusion. [Remember this, for it is quite likely that this afternoon will give you the opportunity to observe the phenomenon. Note that money need not be mentioned explicitly for the nonsense to emerge, a reference to "the taxpayer" can do the job. The role of "the taxpayer" then invariably leads to the conclusion that of State Universities at least the undergraduate curriculum has to be second- or third-rate.] The final reason for my refusal is that the habit appeals to the quantitative mind and I come from a culture in which the primarily quantitative mind does not evoke admiration. [A major reason that we considered Roman Catholics to belong to a lower class was precisely their quantitative bent: they always counted, number of faithful, number of days in purgatory, you name it.....]
[0] https://www.cs.utexas.edu/~EWD/transcriptions/EWD11xx/EWD117...
It was well thought through. I mean, the 360 architecture is still with us because it avoided the traps that killed the PDP-11, VAX and the 68k -- as much as I loved the PDP-11. It's true when it came out that nobody had any idea what a general-purpose operating system looked like and it took a decade for them to productize VM so that you could run as many operating systems as you needed simultaneously.
The quality and speed tradeoff doesn't matter if other incentives aren't aligned. In today's product-driven world, where usage stats are compiled in real-time, features are added and then pop-ups, nudges, etc are added to software to get product usage up. These features may or may not add quality, relevance, etc. But the "success" is measured in usage, forced or not. This is essentially Microsoft today (especially with copilot), but even after every major apple OS update, I get "what's new" pop ups for every Apple app (notes, reminders, mail, etc) on every platform (iPad/iPhone/Mac) that I own. I opened the email app to check my email, don't get in my way!
/change my mind.
speed kills quality. it’s literally impossible to make anything good fast.
we know this, and it still applies to software. while we may be able to make things faster, they will never become good (or great) without an incredible amount of care, patience, and joy from its maker.
there are no shortcuts to quality. it will always take a lot of time to make anything good.
Your comment made me think two things:
• Sometimes constraints make things better, and ‘speed’ can occasionally make you prioritise the things that actually matter so you deliver stuff that counts
• Sometimes, thinking longer about something doesn’t get you closer to the correct answer. You can rearrange and refactor and rewrite and redesign, but you won’t always get something objectively better than what you originally came up with. It’s still your thoughts, your brain and your ways of working that shape the output (and they haven’t changed).
Investing more time to make something better should be a conscious decision. Perhaps it’s one people decide against for the wrong reasons.
But sometimes things need something other than time and effort spent to make them better.
Ideas for songs. Good ideas for songs can come in an instant. And in the hands of an experienced songwriter, you can take that idea to a polished product very quickly. But that's because they spend years developing taste.
> fantastic ideas for a simple product came to somebody in an instant
What's the saying around here? Ideas are cheap. It's execution that matters? And the execution part often comes about through trial and error.
> But sometimes things need something other than time and effort spent to make them better.
I don't think anyone would disagree with this.
Over the two decades of writing software professionally, I learned to appreciate the perspective of software companies: most software is crap, and without pressure to ship, smart programmers will forever keep polishing the turd, way beyond the point the software stops being relevant or useful.
The best, high quality, slow-developed software isn't the best because programmers knew when to stop. It's because the scope and timeline were bounded up front.
Is Zotero good software? Yeah, it seems pretty good. I myself don’t use it, despite trying, because it didn’t solve my problems. It is free, has always been free, and its users are primarily academics who do not face market economics for their work output.
Speed is not the enemy. A good engineer knows that time is a precious resource, just like any other, and should be valued appropriately. And engineers are not academics.
This post is about Zotero. https://www.zotero.org/
If you are not an academic, you might not know Zotero.
It is such a pleasure to use. Every app should be like this.
I read everything in it, including books I'm going through right now from https://teachyourselfcs.com/
It also does an amazing job of taking snapshots of posts. I use it all the time to grab posts from HackerNews so I can mark them up.
And it automagically syncs everywhere across devices and lets me store way too many files on the web like the ADD packrat I am, without breaking a sweat.
In short, this software just works, and it works well. So when somebody behind Zotero talks about how to develop software, I listen.
And it's a fun post with some history. You should save this post to Zotero, and then read it.
If Zotero ever enshittifies or just disappears would I still be able to access my data and transfer it to something else?
Taking a guess here, but Evernote? :)
It was a classic example of "Microsoft made something surprisingly good, made people think it was crap because they stuffed five OneNote icons on your taskbar and three up your nose, and then they wrecked it", must have come from the same mind that thought buying Activision was a good idea.
I have many issues with Zotero (which has been dumbing down at the expense of power users for some time) – but easy private sync is definitely not one of them.
From Zotero's site:
Data syncing is free and unlimited, and it can be used without file syncing.
https://www.zotero.org/support/syncLooking at my own Zotero account, I can confirm that all metadata storage is free and not counted (I am using none of my 300Mb storage in my account according to zotero.org, despite syncing metadata for several thousand bibliography items). The only storage Zotero seems to count is attachment storage, and that's all synced via my WebDAV.
What am I missing here?
Maybe Zotero does more than Wallabag, but I am happy with it so far.
It suggests a certain... simplicity that I chase in all of my developments. The amount of EFFORT to get software from "works but..." is interesting. You can get something that works 99% of the time, but that never quite fits the "just works" level of quality that we love to see.
Really, there's an immediate feeling of joy I get when I interact with software that "just works" knowing how much effort it takes to get something to that point.
Like an application to book a haircut at a salon or log into the WiFi at a hotel that does something "on rails" and should work without any learning or training is one thing. A complex application for professional creative work is something else entirely.
Then there are the things that are just weird, like life got me looking at mobile games lately and I was really shocked to see how many of them have a horrific onboarding experience. The best games in the "clothes collecting" genre, for instance, like Love Nikki have a gentle onboarding process that teach you the game mechanics and get you into the story and maybe get you hooked. When I was developing consumer-facing social and entertainment apps (before Facebook!) I had my own belief that "easy onboarding is good for users and good for profits" and worked with founders who got it and pushed me to outdo myself. Most games in that genre have horrific UI and don't try to explain the mechanics because they could care less if users understand the currencies used in the system because they only care for the worst "whales" who click on any red dot and fold green at the slightest difficulty.
These days when it comes to "slow software", I think five years is an eternity from a platform perspective so if you start something you may need to rebuild it. Like I find it hard to justify developing on anything other than the web platform (like "write once and it runs on platforms you never thought of like VR headsets", "for christ's sake did you realize every Windows software vendor had at least one InstallShield engineer back in the 1990s and IT was pulling their hair out updating all their desktops?", "it took three weeks for the App Store to reject your update? what did you think would happen?") but the app I am working on right now depends on React components that haven't been updated since 2021 and the bill is coming due.
P.s. As a side note, I started making small dress up games that I draw myself to scratch my own itch. I am not a programmer so they are absolutely horrendous and I only show them to friends. Making videogames is not as easy as making paper clothes for paper dolls like I did as a kid, but it is so fun and I feel 30 years younger somehow. :D
(On the other hand, while I believe thinking about "gaming" the ranking takes away from the fun of the game, Dress to Impress is massively popular, so maybe this frustration is what makes it popular. IDK if kids play it because there are no competitors in the genre or because some competitive toxicity is addictive.)
In short, pairwise rankings are real and comparable between people whereas different people might use a scale differently. It's a difficult problem to deal with "i like the bagel shop a little bit more than the rice bowl restaurant but the party member who has celiac's opinion matters more than mine" See
https://brocku.ca/MeadProject/Thurstone/Thurstone_1927f.html
now if you have N options you don't have to have everyone evaluate N(N-1) pairs because preferences are correlated and should be transitive. There are many algorithms, like Elo and the methods in that paper that can be used to run tournaments and compute scores. The game Covet uses pairwise ranking and I think gets good scores.
To get back to social decision theory, it's probably not good for players to be in a team for purposes of winning and losing. Like I think sharing clothes with your friends is a really fun mechanic (might make you want to really buy or work for something!) but if people can help or hurt their friends you get into the whole can of worms of coalitions in N-person games.
I just quickly tried Covet and I can imagine how the act of comparing and picking one of the two options can feel fun. I can see myself just voting on outfits in the game without creating my own when I need some really simple "mindless" fun. Not a productive effect but also not addictive enough AFAIK to be a big problem.
I think teams can work as a concept but you need to design them with a great deal of care so that mechanics are pro-social and fascilitate friendships -- there's a nice article about it: https://www.gamedeveloper.com/design/game-design-patterns-fo...
- Calibre Ebook reader - Browser bookmarks for article sites - Reference manuals/docs - Biology papers of interest
It's baffling actually that Firefox doesn't let you annotate bookmarks, it seems such an obvious feature.
I might be dense - but I added this story as a URL to my library on Zotero on Android - and I don't see any way to read it "on" Zotero? I can apparently create a citation, however.
AI doesn't preclude this. In fact, it can help accelerate parts of it.
He's describing the typical big project lifecycle:
- Examine the landscape
- User research (how they use existing software, what their frustrations are, etc)
- Brainstorming
- Early ideas and prototypes
- Refinement, user feedback
- Solidify the vision and high level process design
- Choose technologies
- Design & architecture
- Plan out phases
- Build phases, then test them with users
LLMs are great at research, and great at prototypes. Once you have your design, they're good at coding as well. They're also good at distilling user feedback.
It can also slow it down as people get distracted experimenting with features they can build quickly.
Most V2 rewrites fall under a similar category.
It's like all the food being replaced by MacDonalds food that is served from a replicator (from Star Trek). Sure, we can survive on it, sort of, but if all we ever know is this food and we're too busy to cook (or no one knows how to manufacture a stove) then we'll never know anything else. Maybe someone will learn how to tweak the recipe a little but our minds will be conditioned to context-switch so much that we'll never even think of learning to cook great meals.
That's what LLMs do. They're shit.
Whenever a new technology makes what used to be hard easy, all the non-experts pile in and produce shit for awhile, until they're finally forced to admit that software engineering is hard (or die trying).
The same will happen with LLMs in a couple of years.
There seems to be this pervasive false dichotomy about software currently: that it's either lovingly hand-tooled and crafted and perfect; or that it's a vibe-coded one-shot sloppy-slop-slop-fest of spaghetti code.
There is a middle ground, where rigorous software project management and careful architectural guidance and oversight continues, but agentic coding takes away much of the coding drudgery. It takes time, and it's not even remotely one-shot, but it's still way faster than coding everything by hand. From what I can tell, this quiet approach is where most agentic coding effort is going currently -- it's not flashy, it's not sloppy, and mostly you won't even notice.
https://assets.buttondown.email/images/93382906-4996-445c-81...
is just of some passionate academics working on a project with no real economic or social media incentives driving it.
> “… we could not have accelerated Zotero’s conception, because we did not know exactly what we wanted, and so could not have written coherent prompts for an LLM.”
A surprising proportion of software products, maybe even businesses today, are solutions in search of a problem.
Sometimes that’s okay, but only sometimes. And being a solution in search of a problem requires you to get everything /else/ pretty much perfect if you want to succeed.
The fact Zotero paid attention to what people wanted, and gave it to them, and were market oriented, is demonstrably a big part of their success.
It is MUCH easier to make something people want, than to make them want something you made.
Rapid prototyping is an amazing opportunity afforded to us by AI. But some people use that potential to spend even longer on a more developed prototype that they are too attached to to get feedback on!
I see the opposite. I see thousands and thousands of quickly produced apps that have been put out there, but have 0 users that could shape it. Not so much because of their quality, but rather distribution in a very crowded and shouty marketplace.
So as a software developer now might be the best time to throw system design on it's head and develop without a direction but make sure everything is as good as one can forge it.
Surprising functionality and stuff not found elsewhere then more or less simply emerges from that.
Granted - that has a certain freedom and no pressure to make money as a prerequisite but so do the 5 years noted here.
At the same time, there are broadly three motions in the loop: - thinking/discussing (what should it do) - building (Make it do that) - using (Seeing whether that is actually what it SHOULD do)
And then of course you repeat until you run out of time, money, patience, volition, etc. For some software there is a terminal state - it truly does exactly what it should. For most you're always reaching for it.
LLMs definitely accelerate 2 and potentially can help accelerate 1 and 3. As such cycle time can be reduced. It still may take 100 turns to get what you want, but I'd be surprised if the clock time for the 100 turns would be unchanged using AI.
People that do think of using LLMs to accelerate 2 usually don’t think enough about 1. And from what I’ve seen, they quickly get tired of 3.
Most quality software I’ve seen usually starts from a small subset for the cycle and then incrementally add to it. LLM projects usually rush the building part and do too big of a job and it’s become cumbersome to design (sunken cost) and evaluate (too many variables). If you want to build shelters, you start with a small hut, not with a cathedral.
What on earth makes you think that you can't adopt this exact approach with agentic coding??
This is like saying that you shouldn't drive a car around suburban streets, because the maximum speed it can do is in excess of 250 kph. You absolutely can (and arguably should) take agentic coding slow and steady, and build software the way you always did. Like using a car to get to the shops at 50 kph, you'll still get there much faster than if you walk.
It can be done, but it’s not where the hype and the practice is. It’s all about number of commits, number of PR, churns in LoC, which has no bearing on software quality and usefulness.
Something like caddy[0] is just 2700+ commits over 7 years. That’s like 32 commits a month in average. Even if you apply an exponential decay (going from greenfield to mature project), the latter years would have been way peaceful. I know AI can help in some cases, but is it such a pressing need. That would been like taking a car to go somewhere less than 10m away.
Most (mature) open source projects progress a few changes at a time. Most of the time is about ensuring that nothing breaks due to the change.
How many cathedral does the world needs? And even if you wanted to build a new one, it’s no longer the middle age. No one is inventing UNIX all over again. Lots of hurdles have been solved. The difficulty of software is not technical or implementation related, it’s mostly about defining the problem.
The difference between an academic and an engineer is that the engineer understands that time spent in research and design does not magically justify itself. Quality and time-to-market are both important and to be an engineer is to figure out how to deliver along the optimal frontier.
> That winter, with Roy as the principal investigator and Josh and me as co-directors, we applied for a grant from the Institute of Museum and Library Services.
Per https://en.wikipedia.org/wiki/Institute_of_Museum_and_Librar... :
> The Institute of Museum and Library Services (IMLS) is an independent agency of the United States federal government established in 1996. It is the main source of federal support for libraries and museums within the United States, having the mission to "advance, support, and empower America's museums, libraries, and related organizations through grantmaking, research, and policy development"
You can guess the next part, I’m sure:
> On March 14, 2025, President Trump issued an executive order virtually eliminating IMLS that directed that "the non-statutory components and functions ... shall be eliminated to the maximum extent consistent with applicable law, and such entities shall reduce the performance of their statutory functions and associated personnel to the minimum presence and function required by law", along with minimizing several other agencies. The entire 70 person staff was put on leave on March 31, 2025.
> On May 1, 2025, a lawsuit brought by the American Library Association and the American Federation of State, County and Municipal Employees resulted in the U.S. District Court for the District of Columbia granting a limited temporary restraining order to block any further actions to dissolve IMLS.
I don’t think this is true. I’m working on medium to large features in companies that have around 1K engineers. These features typically involve 50% of work within your domain plus 50% of work in dependant domains. You cannot get anything doing without previous alignment with such domains. There are discussions, tradeoffs, design docs, approval committees, etc. AI can (and does) help in every step, but it’s not a magic wand that can solve the whole thing with a well crafted prompt. Ithe hands of inexperienced people, AI slows down things (e.g., ai-generated robotic and lengthy slack messages that go nowhere, PRs that implement what a jira ticket says… but jira tickets without alignment are worthless, etc)
Comments are loaded live from Hacker News and are not stored by Mid or Real.
ghoshbishakh 8h ago on HN
adamddev1 8h ago on HN
huijzer 8h ago on HN
josephg 8h ago on HN
LLMs are excellent at making prototypes though. And prototypes can be an excellent way to stop yourself from implementing the wrong features in the first place.
lelanthran 6h ago on HN
Well, yes. I mean, it's in the name Generative Pre-trained Transformer: they're text generators!
To delete text using a text generator, you have to emit the original thing taking care to omit the deleted stuff during emission. It's more work.
iamnothere 4h ago on HN
tripleee 2h ago on HN
We've had that since 1971
rrr_oh_man 1h ago on HN
What are you talking about?