Discussion intelligence

Where the AI market is talking

A source-linked Hacker News monitor for the tracked company graph. These are attributable public discussions—not sentiment, endorsements, or unverified company facts.

Attributed mentions

2181

Across all collected Hacker News results

Companies represented

11

Exact-name monitor matches only

Collection source

Hacker News

Latest new record seen

Source-backed records

Latest market discussions

HN v1 · Reddit and X require approved API adapters

Showing 861880 of 2181 matching discussions

Anthropic
comment

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

Grok is quite interesting. I run comparisons almost daily on tasks and Grok is its own beast, in a good way. It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same thing with diminishing quality. It makes the lineup pointless. Ditto for Anthropic. Gemini-3.6-Flash and 3.1 Pro genuinely behave differently. Opus 5 and Fable are.. cousins. I find that when I want to test a complex creative challenge, having 4 "families" to choose from makes the experience interesting since they will excel in different areas. Grok might implement unique lighting, Opus, elegant primitives, Sol, accurate snowfall in one pass, Gemini, silky movement. Combined, you can pick and choose best. For what its worth, Grok always feels "messy" but finishes. Grok 4.6…

Anthropic
comment

DeepSeek V4 Pro 0813

> An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms This is a different thing. Yes, giving multiple example is called "few-shot prompting". But one-shot vs few-shot benchmarking is different. In this context "one-shot" means "pass at 1 effort" as opposed to "multi-shot". In the literature this is called "pass@k". Anthropic has a good explanation here: https://www.anthropic.com/engineering/demystifying-evals-for... (search for "pass@k"). In this discussion we are discussing pass@1 (single shot) vs pass@(k>1) (multi shot). > The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session …

OpenAI
comment

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

Do you really think OpenAI's leadership are as stridently, overtly political as Musk? OK.

OpenAI
comment

Grok 4.6

> The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. Couldn't be further from the truth. The models can be tested and statistically evaluated. I ran a massive Fable max code review on my lone lisp codebase. Now that I have switched to OpenAI, I decided to run an equivalent review using Sol max and compare them. I'm keeping all data so I can thoroughly evaluate their performance in multiple areas such as correctness, rigor, performance, security, maintainability, consistency, among others. Fable pass is 100% done and I'm around 70% done with the Sol pass. Preliminary results are already becoming clear: Sol is capable of reproducing around 70% to 90% of Fable's performance. Haven't tested open weight models…

OpenAI
comment

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

I've never used Grok, but I'm very dissatisfied with the writing style of frontier models from OpenAI and Anthropic. I only use them for coding now. ChatGPT is very long-winded, sometimes producing multiple bullet point lists for a simple answer. Claude is full of mannerisms: 'not merely x, but y', 'Here's where it gets interesting', 'the real question is', etc.

Anthropic
comment

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

I've never used Grok, but I'm very dissatisfied with the writing style of frontier models from OpenAI and Anthropic. I only use them for coding now. ChatGPT is very long-winded, sometimes producing multiple bullet point lists for a simple answer. Claude is full of mannerisms: 'not merely x, but y', 'Here's where it gets interesting', 'the real question is', etc.

OpenAI
comment

Qwen3.8-2.4T

There’s an alternate universe in which OpenAI stays open, licenses according to revenue, Chinese models don’t gain traction in the US because domestic models take all the capacity…whatever, $1T IPO beats the right answer ever time

OpenAI
comment

Hax – a minimalist, terminal-native coding agent written in C

From the README.md: If "fancy new AI tech in an old-school minimalist package" sounds like your vibe, you might like this. What is the question about: A: Why start a 'new' software project? (instead of old, none, multiple, ...) B: Why in 'C'? (instead of Mojo, Java, D, ...) C: Why 'today'? (instead of Yesterday, Tomorrow, never, ...) For me the beauty of this project is: Someone tried to figure out for themselves what an end-to-end process would look like for what they wanted to do. What happens to the input? How are files written? What is sent to the LLMs? And they used the tool they know best. No `import openai`.

OpenAI
comment

Grok 4.6

Does anyone know how the grok allowances compare to OpenAI / Anthropic for the monthly plans? I heard they're not generous, which means I never really bother testing Grok.

Anthropic
comment

Grok 4.6

Does anyone know how the grok allowances compare to OpenAI / Anthropic for the monthly plans? I heard they're not generous, which means I never really bother testing Grok.

Replicate
comment

Grok 4.6

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

Anthropic
comment

Grok 4.6

The Opus 5 release was a perfect example of how useless these benchmarks are for a head to head model comparison. Anthropic published a post showing Opus 5 beating Fable in almost every eval but then added a disclaimer that it was still a tier below Fable in intelligence (and thus pricing). So then what did all the numbers represent exactly?

Cohere
comment

DeepSeek V4 Pro 0813

Opus 5 doesn't really even speak coherent English. I'm not sure what's going on, but it can't explain anything. It still does an excellent job with code and writing tests and code review and creating and completing a plan, and it seems to be able to understand English instructions, but it sure as hell can't explain what it did or how to use the code it wrote. That was true before they announced the watermarking, I'd already started to back off of using Opus as much because I like to understand what the model is doing and have it write documentation I can use to reproduce its results, but maybe watermarking was already in there unannounced.

Cohere
comment

Show HN: Woxi - Open-source Mathematica / Wolfram Language reimplementation

I love the idea of Sage, and think an open-source state of the art CAS is essential in this day and age, but in trying to master it I got sick of Python. I truly, honestly loathe the way Python handles for symbolic computation, and digging down I traced my disdain all the way to the fundamental object model of Python; so there is no way some surface-level modification will work for me. On the one hand, it allows the sloppy "integration" of various systems (like you mention). On the other hand, it is not a real integration, just a patchwork of the worst kind. You never know what kind of interface you will face next, there is no coherence among tools, everything is its own world and you need to translate manually between them. Assuming you know the interfaces beforehand. If not... good luck.

OpenAI
comment

OpenAI’s head of ethics leaves less than a year after joining

Isn’t ethics incompatible with the nature of a commercial enterprise? Now, if OpenAI was a non-profit… as it was intended originally.

OpenAI
comment

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

grok offers a team subscription which respects user privacy and does not send your prompts out to a team for moderation. that alone makes it the closed source subscription i would choose. claude and openai are spying on you. as it stands i don't have it because the reasoning is encrypted, so i feel that it still is not working for me, it's two faced.

Anthropic
comment

Grok 4.6

Yes that's been obvious since the beginning. That's why you should always monitor your agents closely. Just like supervised self driving cars, you have to watch the road and do some hand holding. The tooling around isolation, logging, and real time security/anonomly detection for regular LLM laptop users is very immature right now. I expect that to change soon. The alternative is extremely locked down models which is what Anthropic seems to want to do.

OpenAI
comment

Qwen3.8-2.4T

The obvious goal is to destabilize the western economy and prove that US tech is a worthless bubble - but I agree, OSS AI is great for everybody and what OpenAI was supposed to be

Methodology: HN Search returns recent public items matching a monitored company name. AIIStack stores a short normalized excerpt and the original link, deduplicates by company/provider/item ID, and creates an activity signal only after a threshold of newly observed records. Review the original discussion before making a decision.