It happened to me as well but in a different direction: i.e adds non library code in a shared library. Another issue I have with is that it is chasing too much edge cases/security issues(I.e chasing ghosts).
However this makes it also a strong model because it fixes/solves problems that both Opus and Fable are incapable. In reviews it catches bugs that both Fable and Opus are missing to spot. To me the “best of both worlds” is to research the problem with got sol, create a plan with Fable and dual review it with both Fable and GPT-SOL and implement it with GPT-SOL. You can see in the code reviews how many times both Fable and opus are sloppy and superficial while GPT-SOL just does its due diligence …
The appeal has been completed and he's been acquitted. Note the past tense in how the CBC wrote this part:
"The Wisconsin police identified Mr. Klayme as the offender using the incorrect username. His conviction rested on the internet username being attributed to him,” the appeal panel wrote.
The Appeal Court decision said no evidence was found linking Klayme to the girl, the Kik conversations, or the intimate images.
In addition to an acquittal entered on the charges, the sentence and ancillary orders against Klayme were vacated.
Genuine question to AI users upset by this: if you trust AI to generate your code, why don't you trust it to help you "deploy https://forgejo.org/ on cheap VPS make no mistakes"?
Every single time any story has the words "Google developers" in it, they're behaving like arrogant, disconnected, anti-consumer jerks. From constant GCP breakage, to not acknowledging obvious Android bugs, to intentionally braking Chrome. It's a stark contrast to behind the scenes interactions with them.
The article overlooks security implications from spyware, which is a huge problem not only for financial applications, but personal safety, too [0].
Per FTC, a stalkerware will: geo locate, read call list & record calls, read notifications, texts, & possibly emails, access gallery, camera, & files, and monitor network activity. [1]
You could do all of those with "on-device adb" (in some cases, with just the appropriate permissions), without root access. Stalkerware & financial fraud enabled merely due to the scale & reach of Android (half of humanity uses it!) and lack of basic security literacy warrants such protective measures, as (Thaler & Sunstein would like to remind us) defaults matter.
With conspiracies abound, we must not lose sight of tech safety and related issues, which almost exclusively affect the most vulnerable & the most disadvantaged.
Maybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.”
You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvious show stopping bug yourself.
The technology is not a brand new one that fixed everything wrong with the old one, no, but not sure I would’ve noticed your comment if it had been such a bland observation. I genuinely assume good faith here… will say am tempted to assume the standards of someone posting such a thing might be impossibly high. Glad to be having a fun conversation instead of getting your grades on my work product or something :)
> This is a CVE by any definition and you'd be screaming your head off if any other OS would allow this kind of permission bypass.
Or perhaps they wouldn't? CVEs aren't holy scripture, and "security" isn't the most important consideration in computing. This theatre has gone too far IMO.
> Please don't use HN primarily for promotion. It's ok to post your own stuff part of the time, but the primary use of the site should be for curiosity.
In France, where we have relatively cheap and clean energy, they plan to build GW DCs with GWs setups of diesel backup generators. When these will run, the pollution will be staggering.
Exactly. I'm a nerd too and we nerds can real hypocrites because we spend too much time alone. I will happily turn a blind eye to world suffering if it means I get to work less but get paid the same.
AI CEOs aren't different. At some point we'll have to admit that we're sacrificing others for our own benefit.
Imagine if I hired someone to do a cad drawing, and they spent X days/weeks of time building a pipeline to visualise the part instead of asking how can they view the part? We’d call that a complete and utter waste of money and time.
You really should play the 25 games before stating that it's "simple". The benchmark doesn't just track "completion", it also tracks the number of steps, and the score is based on the median steps took by human players. So in order to get 99% it would mean that the model solved every level of every game in less steps than the median humans. Which, having played the games and having setup harnesses for local models, I find hard to believe.
Also the models have to figure out what "end" means. And each game involves some kind of "gotchas" thrown in the harder levels. Some games are only solved by about 2/10 people trying them.
The 99% result most likely has some leakage somewhere, either in the preparation of the environments, or from session to session.
While this is certainly impressive, I believe it is not a good answer from a user's perspective. Instead of having Claude automatically spend tokens in an unrelated task, I would have preferred it to first tell me: "I can't view the drawing. Would you like me to build a computer vision pipeline to do it"?
Is surely super agressive to outright downvote given evidence says otherwise but that’s fine I guess. Let’s drop it, apparently we’re not getting to mutually satisfying conclusion.