Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It has to be satire.

What if they spent an hour?

It happened to me as well but in a different direction: i.e adds non library code in a shared library. Another issue I have with is that it is chasing too much edge cases/security issues(I.e chasing ghosts).

However this makes it also a strong model because it fixes/solves problems that both Opus and Fable are incapable. In reviews it catches bugs that both Fable and Opus are missing to spot. To me the “best of both worlds” is to research the problem with got sol, create a plan with Fable and dual review it with both Fable and GPT-SOL and implement it with GPT-SOL. You can see in the code reviews how many times both Fable and opus are sloppy and superficial while GPT-SOL just does its due diligence …


we can just fork android

That's an "overthinking" model.

The appeal has been completed and he's been acquitted. Note the past tense in how the CBC wrote this part:

"The Wisconsin police identified Mr. Klayme as the offender using the incorrect username. His conviction rested on the internet username being attributed to him,” the appeal panel wrote.

The Appeal Court decision said no evidence was found linking Klayme to the girl, the Kik conversations, or the intimate images.

In addition to an acquittal entered on the charges, the sentence and ancillary orders against Klayme were vacated.


Genuine question to AI users upset by this: if you trust AI to generate your code, why don't you trust it to help you "deploy https://forgejo.org/ on cheap VPS make no mistakes"?

Every single time any story has the words "Google developers" in it, they're behaving like arrogant, disconnected, anti-consumer jerks. From constant GCP breakage, to not acknowledging obvious Android bugs, to intentionally braking Chrome. It's a stark contrast to behind the scenes interactions with them.

Wish more web looked like this really. Has distinct character.

solving ARC-AGI and being useful turned out to be two different problems

The article overlooks security implications from spyware, which is a huge problem not only for financial applications, but personal safety, too [0].

Per FTC, a stalkerware will: geo locate, read call list & record calls, read notifications, texts, & possibly emails, access gallery, camera, & files, and monitor network activity. [1]

You could do all of those with "on-device adb" (in some cases, with just the appropriate permissions), without root access. Stalkerware & financial fraud enabled merely due to the scale & reach of Android (half of humanity uses it!) and lack of basic security literacy warrants such protective measures, as (Thaler & Sunstein would like to remind us) defaults matter.

With conspiracies abound, we must not lose sight of tech safety and related issues, which almost exclusively affect the most vulnerable & the most disadvantaged.

[0] https://www.techsafety.org/spyware-and-stalkerware-phone-sur...

[1] https://consumer.ftc.gov/articles/stalkerware-what-know


Maybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.”

You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvious show stopping bug yourself.

The technology is not a brand new one that fixed everything wrong with the old one, no, but not sure I would’ve noticed your comment if it had been such a bland observation. I genuinely assume good faith here… will say am tempted to assume the standards of someone posting such a thing might be impossibly high. Glad to be having a fun conversation instead of getting your grades on my work product or something :)


If you’re trying to quantify fun, you’ve already well and truly missed the point.

Especially the (slightly frightening) animations they show!

> This is a CVE by any definition and you'd be screaming your head off if any other OS would allow this kind of permission bypass.

Or perhaps they wouldn't? CVEs aren't holy scripture, and "security" isn't the most important consideration in computing. This theatre has gone too far IMO.


> Please don't use HN primarily for promotion. It's ok to post your own stuff part of the time, but the primary use of the site should be for curiosity.

https://news.ycombinator.com/newsguidelines.html


In France, where we have relatively cheap and clean energy, they plan to build GW DCs with GWs setups of diesel backup generators. When these will run, the pollution will be staggering.

Yeah I tried to have a conversation about security yesterday with Claude and it immediately stopped me - I was taken aback, I didn’t expect it at all.

This is highly problematic.


Exactly. I'm a nerd too and we nerds can real hypocrites because we spend too much time alone. I will happily turn a blind eye to world suffering if it means I get to work less but get paid the same.

AI CEOs aren't different. At some point we'll have to admit that we're sacrificing others for our own benefit.


That's a valid point. Maybe better to have no ads at all.

This is very risky, water damages are very expensive, please do not do this in an apartment but at your own rural house!

Imagine if I hired someone to do a cad drawing, and they spent X days/weeks of time building a pipeline to visualise the part instead of asking how can they view the part? We’d call that a complete and utter waste of money and time.

That's a valid point. Maybe better to have no ads at all

This is an error by US (Wisconsin) police misidentifying a man in (Halifax, Nova Scotia) Canada. Why do you want to hear from a UK lawyer?

It sounds like the appeal hasn't actually happened/completed yet anyway, that they're trying to raise public interest/outrage to help the case.


Convicting an person who is not guilty, that's how good they are. /s

You really should play the 25 games before stating that it's "simple". The benchmark doesn't just track "completion", it also tracks the number of steps, and the score is based on the median steps took by human players. So in order to get 99% it would mean that the model solved every level of every game in less steps than the median humans. Which, having played the games and having setup harnesses for local models, I find hard to believe.

Also the models have to figure out what "end" means. And each game involves some kind of "gotchas" thrown in the harder levels. Some games are only solved by about 2/10 people trying them.

The 99% result most likely has some leakage somewhere, either in the preparation of the environments, or from session to session.

Seriously, play some of the games. They're fun.


While this is certainly impressive, I believe it is not a good answer from a user's perspective. Instead of having Claude automatically spend tokens in an unrelated task, I would have preferred it to first tell me: "I can't view the drawing. Would you like me to build a computer vision pipeline to do it"?

So you could say that the goose was not valued, huh?

Is surely super agressive to outright downvote given evidence says otherwise but that’s fine I guess. Let’s drop it, apparently we’re not getting to mutually satisfying conclusion.

What better way to spend token. I’m amazed that people talk about this as it’s a good thing!

Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: