Sure but the intention was to compete with Figma, and they are to a degree.
The success of their tomfoolery is really a separate issue from if they are engaging in said tomfoolery, which I think they are. I also think anyone expecting honestly from this is naive.
Bad people doing bad things are not stupid, in fact they're usually pretty smart. Nobody is gonna come out and say what's going on, because that's purely self-destructive. We will get leaks and lawsuits slowly, much like OpenAI.
They systematically violated copyright when they grabbed whole internet to train their models. Do you really believe that they will stop stealing because they signed some funny ToS? Especially when every bit of data they have and competition does not have is making their model better.
This. People repeat stuff they suspect might be happening like it's facts. Would I be surprised? Only a little. Does that mean it's definitely happening? of course not.
Blame copyright law. AI training is very comparable to compiling an inverted index, which is considered transformational even though you can recover the input documents from an inverted index.
To be maximally charitable, while creating large language models has been ruled to be transformative and thus fair use under current copyright law, Anthropic did separately violate copyright law when procuring copyrighted text from LibGen and Pirate Library Mirror. That being said, I agree with you, creating LLMs from copyrighted material is very clearly not a violation of copyright under the current legal regime, so long as you procure the text legally, be that through web scraping or by purchasing books directly. And it’s annoying that people try to muddy the water on this.
People downvote you like you're being paranoid, but we're literally discussing this in a thread that shows how little respect those companies have for any sort of trade secrets.
AI labs can hardly just throw random confidential data into the training and then hope it does not leak into the output of their model in an obvious way.
If that would be found it would destroy their main source of revenue, it could became a major national security or healthcare enforcement matter, and result in criminal investigations.
Some of the smartest people on the planet all in the same room, data at their fingertips… they randomly add it to the training set?
Labs at least must study prompts in an airgapped fashion. From there, consider how they could generate synthetic data to train another model. After, require trusted staff to do multiple levels of independent granular reviews of all fruits of the highest-value stolen inputs. (Or for model training data only, data never has to leave the airgap.)
Definitely risky, anyway. Surely some AI user has sent data, in confidential mode, with a unique shape they expect to be able to recognize if a later model recreated a facsimile even with heavy substitutions… but labs could bring risk of getting caught (over next few years) down quite low with extraordinarily ultraparanoid strategy. (But hopefully everybody is just behaving!)
They could run some sort of analysis to find high value input, such as proprietary technology, algorithms, or strategy.
Then they could group them together for one specific topic, and produce a report that analyzes if the information is plausible.
If so, they can have it send to staff for review, who could then create a test set that rewards the model for going into the direction of the proprietary solutions known to work.
I'm no expert, but at least something like that sounds plausible to me. I still very much doubt they are doing this.
It's actually simpler: your code is not the valuable part, it's the telemetry/metadata/conversation surrounding your session that's valuable. Every time you press escape, every time tell Claude/Codex that it's being an idiot, your back-and-forth conversations, etc. "when/why did we fail and how can we improve?" is what they want to know.
They can use LLMs to launder confidential customer sessions into trainable data. Then they can claim that they don't train on "your data" without it being technically incorrect.
Exactly. They can also feed you a stupider model to goad you into handing over more of this training guidance as well. The incentives are aligned for evil behavior. Open source really needs to win this race, or we need much tighter scrutiny on these AI companies.
That's not at all how they do it. They wash the data. The end result is they can steal your IP insights without it being explicitly tied to your business. All of the decisions you made to build your product? Those become the standard suggestions in the next model for people building the same sorts of products. All of that error correcting you did, which in a normal business would be considered hard-earned IP (like in the case of Apple's lawsuit - the "what not to do" is just as valuable as the "what to do"), is now free correction for the next model to produce the perfect result in a one-shot. Now the AI company can commoditize your labor and your industry, or compete against you if they so wish.
And yet it seems quite clear that they have been directly instructing Apple employees on how to steal data from Apple and bring it to OpenAI - explicitly stealing internal trade secrets from the largest company in the world, one known for its paranoia about leaking any company secrets.
I downvoted because the reply has nothing to do with the argument.
If I know for a fact that you're cheating on your wife, and someone else asks how I know that, then a third person chirping about your sketchy business dealings is entirely irrelevant to the question, no matter how much suspicion it might otherwise raise.
Your parent comment isn’t saying they doubt the assertion, they’re asking for details.
If you say you know for certain, it makes sense to ask how. It makes a big difference if the answer is “I used to work there”, or “I implemented those systems myself”, or “I heard my cousin’s second ex-wife say she heard it from her hairdresser”, or “aliens visited me in my dreams and told me”.
I don’t doubt these companies are lying through their teeth. We have plenty of proof of several cases where they did, to the point believing they are liars is a sensible default, but still I could not say I know for certain of every instance of their lies. Knowing how empowers you to do something about it and convince others.
> No one is going to admit to having an inside source on HackerNews.
Not only is that not true (people make throwaway accounts specifically to share insider info), no one has said this was insider information, there are plenty of other ways to know these details.
> Read between the lines.
That means nothing. There’s no information given, there’s nothing to read between.
How do you know that?