Yep, that happened to me too. The biggest joke was that LinkedIn forced me to create a password before I could delete the profile they'd created without my permission. This happened after I'd visited several LinkedIn links on my Pixel phone so they grabbed my Google email address and used that. Bear in mind also that I'd been retired for about six years at that point.
Yes, that is precisely what happened. Over and over again. Every time I had to delete it. It was something to do with the way they embedded Google login on the page.
Interesting to see how much more popular creating games has become as AI has become more powerful.
Right now I am working on a house price guessing game and I know I would not have be able to get anywhere with it a couple of years ago. It has still taken me a few weeks to get it where I wanted but I have had to intervene a lot with things the AI just wasn't good at.
> Even worse, it's not a fair comparison: they purposefully just used "adaptive" instead of "max" for Fable.
We agree models should be compared on a fair basis. Unfortunately, adaptive was the only publicly available number. Anthropic doesn't generally let us run their models for evals, so we rely on whatever Anthropic or third parties have published. In this case, the Agents' Last Exam leaderboard has Fable Adaptive, but not Fable Max.
Would have loved to publish a full curve for Fable if anyone makes the data available.
Although we do bias toward publishing evals where we're ahead, we have historically been unafraid to publish evals where we're behind (e.g., GDPval). The point is give people useful information to decide what's best, not to trick people.
Edit: Now I see there's a second entry with xhigh effort. Not sure if that was added or recently or we skipped it.
Probably because it's only 9 months old, developed by a Chinese developer (from what I can gather), and vibe coded. The developer has plenty of prior pre-AI experience though.
Why doesn't something like this exist for real estate? A popular open source AVM (automated valuation model) that helps home sellers get an idea of what their home will sell for.
Right now it seems AVMs are mainly seen as just a way to capture leads. Every estate agent will tell you they have some magic recipe that makes their valuation better than anyone else's.
I have had a bunch of ideas on how to approach this, but I really could do with a collaborator or two.
FWIW, I don't think the GP post is disparaging (at least as I read it right now.)
I think it is fair to list limitations from using a library that provides an abstraction; it can suggest why a tool isn't right for a person's use cases.
But it also sounds like this API handles those pretty well.
The issue is that it’s relatively low effort to make false and unverified claims. Defending and refuting it is a much higher effort task for the person doing the work to everyone else’s benefit.
RubyLLM dev literally had to take time to provide code samples and doc links.
No issue with listing legit limitations, but be a bro and fact check claims before wasting a volunteer’s time - and potentially leading other developers on a public board astray.
I completely agree that unverified claims create a heavy burden for maintainers. My only point was about the language used: 'disparaging' to me implies a bad-faith attack or a dismissive attitude, whereas this was just an honest technical mix-up that the poster immediately corrected.
I think part of the confusion with that word comes from things like corporate non-disparagement clauses. In those contracts, lawyers write the terms so broad that "disparagement" means saying anything negative, regardless of malice or intent.
Thanks for the discourse. I never meant it to be disparagement nor do I think it really was.
I checked and it turns out I remembered correctly that setting effort and some of its settings are not portable between providers.
There are some different settings that each provider uses and in order for it to be portable, you have to force some defaults on provider A when using a setting that is almost only supported in provider B.
In our implementation we decided to drop a certain setting when using OpenAI in one case and we decided we can just force some other setting when using Anthropic. But this 'solution', might not be what others expect.
When you build an open-source library you can go this opinionated route and force these settings, or you might go the config route and force people to explicitly handle per-provider differences. I will have a look at what I am able to do in terms of a contribution and then in the PR Carmine can decide what they like.
Somehow linkedin creates a new account (on linkedin? - so a brand new account?) without your agreement?
reply