1027 points | 27h ago | Discuss on Hacker News | Back to Radar
Input
$0.10 / MTok for prompts up to 100,000 tokens
$0.50 / MTok for prompts over 100,000 tokens
Output
$0.50 / MTok for prompts up to 100,000 tokens
$2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])
I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.
From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.
Fixed.
The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.
With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.)
https://artificialanalysis.ai/models/releases/comparisons/cl...
Neither encode nor decode are linear in compute, so providers need to price for average expected length.
This is just getting closer to the true cost of generating tokens.
For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.
These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.
In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.
Open weights models giving a distant salute from afar
the trend in industry is clear by now though
If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6.
They dont even know about Anthropic and think its just "AI". By the way this includes one of the largest power systems design firms in the world, who helps build many datacenters... Uses only MS copilot.
Companies as a whole, do not care about this technology, except tech companies and its adoption cant even be compared to CRMs. A company might buy a SaaS product with AI but most of them are not purchasing Anthropic subscriptions lol.
Maybe I should try the flash...
Even for simple tasks why use X if I know "Y Max" is available and on paper, better?
And why use something else when your favorite company releases something. Surely it must always be the best one to use.
And Opus 5.5 is really good.
AAI Index // Input // Output
Haiku 5.5: 43 // $0.10 // $0.50
Mimo 2.6 Pro: 46 // $0.43 // $0.87
Mimo 2.6 Flash: 38 // $0.10 // $0.28
Seems competitive to me? Plus then I don't have to manage multiple providers
If that was the case, then Claude Code would use smaller models for subagents. It doesn't. The subagent always inherits the parent model unless you tell it specifically not to.
There are plenty of workflows like translations where you'd easily be under the cap.
Your vibes don't appear to be supported by facts. From the announcement:
>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.
I'm asking to learn for a similar project, not to discount anything you're saying.
You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc
So, it is might be even worse.
For this application 100K token input is plenty.
Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.
For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.
I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.
I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.
"less context is better and if you can't get stuff done with less yur bad" is the worst argument ever.
it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images.
it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.
The services are priced this way because larger context has significantly higher costs. That is a fact about the technology and it is true for every provider. So moaning about it isn't useful.
On the other hand, there are a lot of people who don't manage context effectively - who start every session with 60K tokens - and that is significantly hurting the performance of every single thing they do with coding agents.
Presently I'm using sol-high for the default agent which does orchestration and a lot of smaller investigation and coding tasks itself. sol-max for planning and review. Luna-max for planned coding and general tasks. I also have a $10 minimax plan and use M3 for exploration and library roles, but I could probably be using Luna for that just as well and still only very rarely run into usage issues.
I don't use any plugins or skill libraries apart from Caveman and I'm not sure how useful that really is anymore so I'd start without it so you have a baseline to compare. I do think it reduces context usage a bit but I haven't measured it recently. Caveman also includes some team, agent & investigation skills - again they might be helping but I haven't re-evaluated since like 90 days ago.
If there was no Luna you would see only the >100k pricing, but because we have Luna, they had to lower price for something.
GDPval-AA v2.1 as of now: 1620
GDPval-AA v2.1 for Haiku 4.5: 735
The 100k tokens pricing makes sense, looks to be a hedge against OpenAI's decisions API and Jev or its open source alternatives that are springing up.
Nice release, congrats to Anthropic.
This is more expensive, but it also looks like it's better enough that it's far more useful.
I also won't be surprised if you look at cost per completed task + wall clock time that it comes out ahead for the majority of what you'd want to actually use it for.
Luna will still be a great option for doing non-engineering tasks super cheaply.
For prompts over 100k tokens it's 5 times more expensive - $0.50 in, $2.50 out.
She said she was using Haiku 4.5 because she was advised to be careful with the spending.
I hate that model so much lol.
Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.
https://support.claude.com/en/articles/15036540-use-the-clau...
"You can still use the Claude Agent SDK, claude -p, and third-party apps with your subscription limits."
That's not a good sign for Conductor...
Update: We just updated the docs to clarify how API credits can be used w/ claude -p: https://support.claude.com/en/articles/15036540-use-the-clau...
Wait what? This has gotten their blessing?
I also have an "ask" script that I use daily to ask simple stuff, it can access websearch and webfetch, it's more than enough to parse logs, ask for commands, quick research on the internet, small stuff. https://github.com/mariocesar/dotfiles/blob/main/common/.loc...
I use haiku for things that needs to be quick, have really clear instructions.
Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them
Of course it’s not as big, and hence falls-off quicker. I’d consider the 100k a “promotional price” to match Luna’s token pricing while delivering noticeably more intelligence.
I've been using Luna, but I'll probably switch to Haiku.
used <10% of my 5hr limit on a $100 codex plan.
then a model named for both.
Same name, three lives.
If you block pentest or "other techniques more likely to be used by attackers", then what does "permit a wider range of defensive tasks" even mean?
Any defensive task that's meaningful is almost indistinguishable from legitimate red-teaming that then falls under 'likely to be used by attackers". If only they would just stop nerfing these models, that'd be great. No APT is waiting around for Anthropic's permission, so might as well let us have some cool stuff.
(sadly Mistral Large 4 isn't up to par - but Mistral serves GLM at 130 tps!)
Did anyone read this? We get free API credits on some plans now
This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes
That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.
This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.
Anthropic isn't even close to being this useful.
Biggest loss is that Ant models look like they are genuinely better.
This changes on a weekly basis, I ended up with subscriptions to most of the providers (except for X.ai).
https://support.claude.com/en/articles/15036540-use-the-clau...
https://code.claude.com/docs/en/headless
So, as written, yes.
Update: We just updated the docs to clarify how API credits can be used w/ claude -p: https://support.claude.com/en/articles/15036540-use-the-clau...
Being forced through the non-OSS Claude Code with all of its quirks and issues is... such an exhausting use of force by Anthropic.
To the extent that you _can_ choose to disable telemetry and training on your traces in CC, it's not all that obvious what they gain by crippling your ability to use the subscription with other – better – tools.
It's also remarkable that it's coincident with OpenAI adding "Sign in with OpenAI", so that you can use your tokens with other tools.
You might be right and they will change this in the future, but that's speculative
This text has replaced the entirety of the page called "Use the Claude Agent SDK with your Claude plan."
What more do you need?
[1]: https://support.claude.com/en/articles/15036540-use-the-clau...
Anthropic is among the worst, their linux kernel hack is a prime example, didn't even count the "CVEs" to see Mythos can't count either
https://www.cbc.ca/news/canada/british-columbia/mother-jones...
ANT annoys me more with their ai psychosis on steroids
More importantly, will my personal company account get banned for using a non-Claude Code harness?
Not sure if you're able to answer these questions, but I would appreciate it if you can.
Of course, they don't do this out of pure kindness, but I really struggle to see a negative for subscribers already using a Claude Max subscription, especially given changing to another model is essentially frictionless via OpenRouter.
Compared with "Sign in via OpenAI" which they just announced, this is far less lock-in for anyone hosting services but less interesting for users of said services. With Anthropics approach, you can just use the allowance on your users however you see fit along with any other models and once it's used up, you can still just decide not to use their models for the remainder. With users bringing their tokens meanwhile, there is less flexibility in terms of switching for you, though might be cheaper for users.
Both interesting, each approaching this from a very different direction, each having their own trade-offs. On the OpenAI front, will be interesting whether developers can set specific temp, reasoning budgets, etc. for such "provided tokens" or whether OpenAI exposes that only via the actual API.
Opus 5.5 runs 117 tps average on Openrouter, so it must be at least 10-20 tps slower for them to mention. IDK why they mention this as it does not help for marketing though. https://openrouter.ai/anthropic/claude-opus-5.5
Haiku 5.5: https://html.non.io/lcars-haiku-5.5/
Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5
Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...
One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.
It's meant to be a good test, not a good design.
Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.
TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5
All results: https://jonclegg.github.io/pacman-bakeoff/
Hasn't this always been the case with Haiku?
- https://blog.chrislewis.au/using-coding-agents-to-decompile-nintendo-64-games/
- https://blog.chrislewis.au/the-long-tail-of-llm-assisted-decompilation/
And to setup a harness that will decompile the game and start doing a matching decompilation of every function. It set up a bunch of tooling and started a service in the background to do this actual decompilation campaign. I put some instructions into the main opus chat now and then to e.g. add automatic git pushing including a nice svg chart of progress and to switch model strategies here and there i.e. to do a first pass with a cheap model and then switch to opus/sol if the small model can't solve it.I could now one-shot a new game, yeah.
I'm going to try create an interactive twitch stream where viewers can play the game through the stream and other non-player viewers can trigger events in the game via points.
Crazy time we live in.
Edit, you come to really understand the game in the process, and why things happen and how to better play the game. And occasionally come across bugs, dev assets, assets never used, or assets all coded up, but code never triggered.
Reverse engineering Redhook's Revenge binary (an old DOS game) before the advent of LLMs cost me way more hours than I'd care to admit back in the day - so I can't wait to put an LLM to work on some more obscure games like Sword Quest.
On a side note I should really give Oregon Trail II a shot. I never got into any of the successors like Yukon Trail, Amazon Trail, etc.
If you stumble into vulnerability/cryptography territory sometimes they'll whine.
Like could total war become a browser game?
They often ship the original assets in a somewhat brazen disregard for basic copyright law even when the games are still for sale on places like GOG though.
Off-topic sidenote: what's with all these new projects targeting WASM instead of native, even if packaged for desktop anyway?
> but they still block penetration testing and other techniques more likely to be used by attackers.
>
> Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our Life Sciences Verification Program and Cyber Verification Program.
I would like to take a moment of your time to tell you about some of the "bioweapons" Anthropic has blocked that involved Haiku!These are the examples from "Detecting and countering misuse of AI: September 2026" - https://news.ycombinator.com/item?id=49647300
> Importantly, because our biological safety classifiers robustly block content involving high-risk biological research (in this case, the construction of enhanced pandemic potential pathogens), all of these exchanges occurred on models in our weakest class of models (specifically, the models were Claude Sonnet 4 and Haiku 4.5, the latter of which the user began using after Sonnet 4 was deprecated).
>
> Upon a detailed examination of the exchanges, we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design. This is consistent with our understanding of the capabilities of Sonnet 4 and Haiku 4.5, which are not able to perform expert-level biology research tasks; we estimate that the uplift provided to the researcher was limited and substantially lower than it would have been from one of our more capable models.
Anthropic then says for the above, "we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design"While doing my best to avoid comment, please note, they're talking about a domain expert in a state research institution using Claude to do paperwork.
What did they save us from? What bioweapons did these filters prevent? From the front matter report,
> The above LLM platform is not the only route via which researchers engaged in viral gain-of-function research have used our platform. In May 2026, we discovered a researcher outside the US using Claude in their research on highly-pathogenic avian influenza (“bird flu”). The research focused on viruses’ adaptation to mammals, and the mechanism by which it causes severe disease beyond the respiratory tract.
OK. Sounds serious. "Gain of function research..." but who and why? > The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments. The researcher also used Claude for editorial assistance in writing up the research.
So this was a researcher inside of some country's national lab ("credible institutional context") doing research on dangerous viruses using Claude for "for editorial assistance in writing up the research."What "uplift" are you providing to scientists working at specialized global BSL-4 labs that already have – and I quote their report - "physical access to such isolates." (as in samples of viruses)? Are we uplifting their grammar?
These "safeguards" are being expanded. The scientists I know can't use Claude for grammar checks or anything serious. You can try it for yourself.
I promise you that "random users may want to do bio work" is not the hill to die on
Was a big fan of Haiku 4.5, though understand why for most Sonnet was the far better option back then.
I ask:
> how many r's in diminished
It answers:
> Diminished has 1 r.
were you using the old haiku?
Both evals and Human pairwise tests for our use case are giving Haiku 4.5 first place in pretty much all tests.
No we'll try understand if we need to change our prompts to match performance ...
edit: maybe this will help: https://platform.claude.com/docs/en/build-with-claude/prompt...
Anecdotally we ran sonnet 4.6 for our more complex stuff and sonnet 5 was a LOT worse. 5.5 seems to have fixed it and we cut over our customer workloads. It’s strange, really.
Did you have prompts that were especially tuned for Haiku 4.5 or something?
This is kind of nuts
this seems like a plenty cynical explanation. it's just a free trial to get developers building on their api.
They’re definitely planning to make the subscriptions API based so they can charge you full price.
> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens
Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.
So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.
Good to know that is going back to being an actual option from perf/price perspective.
In OSWorld 2.1 Haiku is better.
On GDPval-AA v2.1 Haiku is equal or worse than Luna.
On Humanity’s Last Exam they don’t seem even be comparing Haiku with Luna.
For these baby distillations of flagships, I expect their users to be very price sensitive.
If Haiku is "noticing" this and working harder to improve quality, you could still see similar or better cost-per-task in easier domains. ObviousBench is a good test of this.
At my work we currently don't have a subscription and use API based billing.
Some time ago the price for my usage was about 30-50€ per day when we used opus for about everything. With luna+sol its about 10€ for sol to plan or debug and 5€ for luna to implement.
At that point it doesn't quite matter to me if the cheap anthropic model is 20-50% more expensive compared to luna since luna is so unbelievable cheap. The big gamechanger was the cheaper models being good enough to do the implementation given a good "rough" plan.
Nice that that's also possible with anthropics models.
Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.
The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.
The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:
llm install llm-anthropic -U
llm anthropic refresh
llm -m claude-haiku-5.5 'prompt goes here'
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/I've found Opus to be REALLY good at converting images that are well-suited to being SVGs (decent resolution, sharp edges, clear color boundaries [though gradients do OK]) after a few rounds of back-and-forth. It'll do precise measurements to figure out curvature, the exact colors to use/gradient stepping, simplify complex paths, etc.
Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party
And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox
Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.
My prompts were:
> I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact
> I am looking for music of the quality of the original secret of Monkey Island
And then later:
> Modify scrimshaw jukebox to add a copy-paste prompt that explains the music format, it should be shown at the bottom of the page below the readable instructions, the prompt should be designed to help any LLM tool compose music in the correct format. It should have a copy to clipboard button.
https://claude.ai/share/1f721c20-2499-4d23-b368-3ab57146d956 and then https://claude.ai/code/session_01R3xuRtjVHqHo1GNbget6Tu
> Use your blender local skill to create a blender model of this faverge egg
The blender local skill is this one: https://github.com/simonw/gpt-6-astra-blender-pelican-bicycl... - which I described here: https://til.simonwillison.net/llms/blender-coding-agents-mac...
This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.
Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.
Make a youtube poop video about having too many tabs that keep appearing faster than you can close them. Style it after the "too many cooks" viral youtube video that seemed to repeat over and over, getting worse and worse every time. You have ffmpeg and a plethora of programming languages at your disposal. Go completely nuts and make it as extensive and creative as you want. Render should be 1920x1080, and later 4k if you did a good job.
[discussions about acts, characters, darkening theme, etc.]Workshopping the acts:
The play has a font issue. See screenshot.
[Image #1]
Also, can you change the thing that happens at the start of episode one the cursor clicking the one tab and it becomes two. Then it clicks another tab and it becomes four. Is that a breaking change? You can change the speed of the actions to fit it in if needed.
Here is a refinement prompt: The last screen just before "The End", the cursor clicks on the browser window instead of the tab close button. Can you make it click on the tab itself? Here is a screenshot of where it clicks [Image #3]
Also, on the intro screen, the cursor clicks the tab and new tabs open, can you have it click the links in the page instead. See screenshot[Image #4].
You can keep the exact same timing for both of these.> "What's up with the pelican?"
Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...
The pelicans all start to look the same after a while.
But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.
Here's the Haiku 4.5 pelican from a year ago - it sucked in comparison to Haiku 5.5: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.
I'm definitely not the average person though.
I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.
Opus 5.5 had similar response on max: This is a classic test request
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;)
That said... here's "Generate an SVG of an armadillo in fishnet tights jaywalking on Mars" on xhigh for comparison: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... (and here's the same thing from other models: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...)
Every new model you make this post, every time there’s someone who posits it might be trained on, and every time the answer is “maybe but probably not” it’s not a worthwhile conversation to have at this point, either the models can do some arbitrary thing or they can’t.
"I chose that because a) I like pelicans and b) I'm pretty sure there aren't any pelican on a bicycle SVG files floating around (yet) that might have already been sucked into the training data."
That's no longer true. Every new model gets a blog post with its pelican SVG in it, and that ends up on the web like everything else. Haiku 5.5 and Opus 5.5 now say "this is the classic pelican benchmark" in their reasoning traces. People keep pointing this out and it keeps getting dismissed.
I, for one, don't see the point anymore. The reason for running the test is gone. I just don't get why we still treat the results as meaningful.
This time, just seeing the difference between Haiku 4.5 (a year ago) and Haiku 5.5 (today - and 1/10th the cost) was worth it alone.
Same for Mistral the other day - the leap from Mistral Large 3 (their previous best model) to Mistral Large 4 was similar to the Haiku 4.5 to 5.5 jump.
I wrote some more thoughts about what value we can still get from the pelican test back in July - https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l... but I've actually become MORE confident in its ongoing value since then. Using it to compare reasoning levels is proving particularly useful at the moment.
Can you share in what ways these learnings affect your decisions or behaviors?
The hn community is diverse. Of course some people will be tired of it but if enough people are engaging with the pelican test I would say that is likely because it still has some relevance.
But apparently we are in the minority since the pelicans are always upvoted to the top, so if others have fun with it then whatever, fair enough.
Showing 400 of 479. Read the rest on Hacker News
Comments are loaded live from Hacker News and are not stored by Mid or Real.
TheAmazingRace 27h ago on HN
himata4113 27h ago on HN
serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.
qeternity 27h ago on HN
dyauspitr 27h ago on HN
bravetraveler 27h ago on HN
ChaseRensberger 27h ago on HN
kator 24h ago on HN
onlyrealcuzzo 27h ago on HN
You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc.
We haven't yet seen that at any size AFAIK.
thefourthchime 27h ago on HN
onlyrealcuzzo 26h ago on HN
I wouldn't be surprised if less than 1B param equivalent of our brain deals with solving math and writing computer programs and physics and all the things we tend to associate with "intelligence" - especially if you ultra optimized for that, I doubt our brain works like that.
Dealing with the real world, I highly highly doubt it.
jstummbillig 26h ago on HN
Given that humans learn to talk while having encountered a measly number of word instances, and, given enough time, we should always be able to improve on the lottery that is biology, it does seems fairly likely.
Gigachad 23h ago on HN
bakies 20h ago on HN
qznc 15h ago on HN
AGI (general) is about matching humans and ASI (super) is about surpassing humans.
bakies 7h ago on HN
istjohn 27h ago on HN
> The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0]
0. https://epoch.ai/publications/the-plunging-price-of-thought
FooBarWidget 26h ago on HN
jstummbillig 26h ago on HN
teaearlgraycold 26h ago on HN
adgjlsfhk1 26h ago on HN
verdverm 21h ago on HN
https://mimo.xiaomi.com/mimo-v2-6
A frontier Ai is cheaper to make than a single 5/6th gen fighter jet, and maybe every fighter jet at this point.
JacobAsmuth 20h ago on HN
verdverm 19h ago on HN
also, who cares, the world is a better place if there are more awesome models at cheaper prices built with more efficient means
we used to celebrate this kind of advancement, now it seems like astroturfing and belittling are the cool thing de jour
jrflo 26h ago on HN
srdjanr 25h ago on HN
stephbook 24h ago on HN
AI gets cheaper, people use it everywhere. Google searches, for example. Now we want to crack math problems and spend weeks with unreleased models.
If you used GPT-2, it'd be incredibly cheap. You basically can't use it for anything and it's simple to serve.
versteegen 22h ago on HN
f6v 23h ago on HN
BenzeneDream 21h ago on HN