Attributed mentions
2579
Across all collected Hacker News results
Discussion intelligence
A source-linked Hacker News monitor for the tracked company graph. These are attributable public discussions—not sentiment, endorsements, or unverified company facts.
Attributed mentions
2579
Across all collected Hacker News results
Companies represented
12
Exact-name monitor matches only
Collection source
Hacker News
Latest new record seen
Source-backed records
Showing 1661–1680 of 2579 matching discussions
The difficulty of evaluating coding agents is indeed a really big challenge. We built evals on our own codebase and shared some information about that to allow other companies to replicate. We found our own evals correlated loosely with public generic SWE benchmarks. In large user populations like at Databricks I think the ultimate answer will come from experimentation instead of offline evals. We are already doing this in small groups, exposing them to new candidate models and then measuring per-developer cost and perceived quality changes.
> My concern is what a misaligned model will do when they’re even more competent. I think the alignment talk is a red herring. It won't matter in the end, because there will be (if there aren't already) efforts to train offensive models without any guardrails whatsoever. And RL has another advantage: you can reward for whatever you need, and get different results. Right now they're training for general capabilities, but in the future I could see models trained for stealth intrusion and ensuring access, or for all out "milspec" penetrate, replicate and disable, or anything in between.
Coacervates, liposomes, acetogenic precursors, etc are all commonplace and could be considered proto-life. Some scientists have even argued that viruses are proto life. Prions (misfolded proteins that can self-replicate) and plasmids (non-chromosomal DNA strands that play a major role in horizontal gene transfer) are also candidates for proto-life that blur the line between living and non-living
From the article: In African tech hubs, developers are picking China’s cheap, freely available artificial intelligence models over more powerful U.S. ones. ... Chinese “open-source” models account for roughly half of total A.I. use on OpenRouter, a service with 400 A.I. models for users to choose from, up from less than 25 percent a year ago, according to a New York Times analysis of data from eight million customers. Chinese models are also 19 of the 25 most downloaded open-source systems on Hugging Face, another A.I. database.
> they can replicate or beat the price with rented GPUs They "can" is the caveat here. Rented GPUs are going up in pricing. I recently got an email that DigitalOcean pricing of GPUs were going up. So 1. They have to get a hold of them (availability is bad) 2. They have to maintain the pricing
I also think you under appreciate the value of toy software. We’ve spent 20 years talking about serious practices for making serious software but what if the next 20 years is about making personal toys that scratch our own itch? Perhaps LLMs fill this need quite well and they aren’t suppose to replicate what already exists.
The nice thing about natural language is that nesting semantic layers is free and arbitrary, and far more tractable than in a formal grammar. Every natural language is like coherentist ω-order logic. Effectively, I don't have to write the 10k requirements. I only need to provide a sufficient metatheory that can be extrapolatable to those 10k requirements, and that can include embedded theory I did not write myself but am familiar with enough to invoke, as well as refinement criteria ranging from the fuzzy to the explicit with priority weighting parameters to describe the shape in which I want the search space pruned. This isn't anything new or particularly interesting. It's the entire basis upon which ILP demonstrated generality. A metatheory to synthesize 10 trillion rules isn't even scr…
Isn't this the opposite of what everyone is saying should happen? That is, lead with open models -- or at least "openness" and don't leave the capabilities in the hands of an elite few? Did they learn nothing from the Hugging Face incident, where HF wasn't even able to use the models to defend itself from OAI's attack?
Their conclusion is also interesting. They don't see this as an alignment failure. They just think that their internal security measures in the training/evaluation environments were insufficient, and that this accidental (unintentional on the human side) attack on Hugging Face is a warning shot for intentional attacks by bad actors, which will occur very soon. For defense, they say models should be able to not just autonomously fix security vulnerabilities but also to then deploy them to production, without any human approval in the loop, otherwise the offense will be favored compared to the defense. (I guess the latter won't be popular among organizations, though they might warm up to it.) But the more interesting thing is, as I said, that at least in this talk, they don't even ment…
The recent Hugging Face incident did not seem like FUD to me
My brain went from "wow this is cool" at the first OpenAI x Hugging Face report. Now I feel I feel like this is intentional, in an attempt to force regulation that would squeeze out open-weight models.
The sweet spot is representative democracy with a tendency of socialism, to mildly distribute welath and assure social coherence. A voting system which results in a division of force, representing the voters will. The current US administration calls such things comunism.
It’s the lack of coherent leadership at the top, the lawyers, junk bond, equity bankers, oil-men, and the military have had their chance, it’s time for the engineers, science, medical, builders, local bankers and internal infrastructure people to run things. And when I say run things I mean, stop accepting their word as the truth.
Location: New York, NY Remote: Yes Willing to relocate: Not specified Technologies: TypeScript, React, Next.js, Node.js, NestJS, Express, Go (Golang), Python, FastAPI, Django, Java, Spring Boot, GraphQL, PostgreSQL, MongoDB, Redis, Kafka, AWS, GCP, Azure, Kubernetes, Docker, Terraform, LLMs, LangChain, LangGraph, RAG, Stable Diffusion, ComfyUI, TensorFlow, PyTorch Résumé/CV: https://www.linkedin.com/in/joshua-johnson-77271a116/ Email: Joshua.johnson.tech@hotmail.com About: Highly accomplished Senior / Founding Full-Stack Engineer with 10+ years of experience across the full software development lifecycle, specializing in designing, building, and scaling cloud-native, high-performance applications. Core expertise includes TypeScript, React, Next.js, Node.…
> There is no single effect of AI on work quality. Depends on use and effort. Quoting the article... "Nearly half of senior leaders say they would trust a chatbot’s judgment over their own, on questions they were hired to answer." If they're using it like an oracle, then that is worrying, but how worrying - practically speaking - would depend on the baseline prior to the use of AI and how that compares with what the AI is spitting out. The first big hazard of this AI worship is, of course, the atrophy of reasoning skills. I don't think there's an especially distinct skill needed here. A rational person is able to evaluate claims made by others responsibly based on a combination of proxy signals like an earned reputation of trustworthiness, integrity, and authority, the logical coherenc…
I’m not endorsing the current administration’s policy, nor do I think it would succeed, but I find this comment really perplexing since this action is PRECISELY in line with the 2025 National Security Strategy. [1] Also, for further insights on some of the unstated policy, I recommend reading the project 2025 document as well. [2] Perhaps by “coherent policy” you meant “sound” or “effective”? But as it stands it certainly is coherent. [1] https://www.whitehouse.gov/wp-content/uploads/2025/12/2025-N... [2] https://www.documentcloud.org/documents/24088042-project-202...
Likely we don't need to replicate. We just need to get a sample of life with completely independent origin to our kind of life here on Earth, to study. This way we would be able to separate what are universal properties for any life from historical contingencies specific to us. To have samples of a single kind of life complicates things.
unfortunately the field of abiogenesis, just like extraterrestrial life, is a lost cause. The reality we will never be able to replicate creation of life nor find life on other planets (except via contamination). it's why we've never gotten any closer. and i'm not even religious
Bad analogy. In GH you have to press the buttons at the exact time that is prescribed and you're scored on that. With AI there is no such prescribed goal. The world is your oyster. You can use it in boring ways, but you can use it in immensely creative ways too. You can chain together AI agents, image generators, video generators, 3D asset generators, computer vision models, text embeddings etc. etc. and prompt all the orchestration into existence with your ideas, taste and direction. It's the opposite of a constrained game. Boring people use it in boring ways, interesting people use it in interesting ways. The rest is old man yells at cloud, like how people criticized "laptop music" or photography.
I feel two ways about it. I've written some algorithms and architectures for motion control that were good, after many years of chewing on the problem through iterations. I don't think I'll replicate that effort or those results, now that I use LLM's so much. On the other hand, I did that in pursuit of a goal, and LLM's take me further, faster, toward that goal.
Methodology: HN Search returns recent public items matching a monitored company name. AIIStack stores a short normalized excerpt and the original link, deduplicates by company/provider/item ID, and creates an activity signal only after a threshold of newly observed records. Review the original discussion before making a decision.