Is scraping social media legal? The clauses, quoted
Every article on this cites hiQ v LinkedIn as the case that made scraping public data legal. Almost none mention how it ended: hiQ paid $500,000, accepted a permanent injunction, and deleted everything it had built. Here are the actual clauses, photographed in the documents they live in.

Scraping a public social media profile is not a crime in the United States, and three separate courts have now said so in one form or another. It is also, on every major platform, a breach of the contract you agreed to. And if the data describes identifiable people it is a data protection question before it is either of those. All three are true at once, which is why every confident one line answer you have read on this subject is wrong.
Our articles are still written by humans!
Get human written articles in your Google feed.
The clauses, quoted whole
Start here rather than with the law, because this is the part that actually binds you and it is written in English you can read. Every quotation below was read on the platform's own page on 8 September 2026, and each is shown in a screenshot of the document so you can see what sits either side of it.
| Platform | What the terms actually say | Where |
|---|---|---|
| Meta, so Facebook and Instagram | "You may not access or collect data from our Products using automated means (without our prior permission) or attempt to access data that you do not have permission to access, regardless of whether such automated access or collection is undertaken while logged in to a Facebook account. We also reserve all of our rights against text and data mining." | Terms of Service, section 3.2, point 3 |
| Meta | "Except as provided in the Platform Terms, you may not sell, license or purchase any data obtained from us or our services, regardless of whether such data was obtained while logged in to a Facebook account." | Terms of Service, section 3.2, point 5 |
| "You can't attempt to create accounts or access or collect information in unauthorised ways. This includes creating accounts or accessing or collecting information in an automated way without our express permission, regardless of whether such automated access or collection is undertaken while logged-in to an Instagram account." | Terms of Use | |
| Meta | "[A]cceptance of these Terms alone does not constitute the required written permission to conduct Automated Data Collection; such permission must be obtained separately through Meta's formal authorization process." | Automated Data Collection Terms, effective 7 October 2024 |
| TikTok | You may not "use automated scripts to collect information from or otherwise interact with the Services" | Terms of Service, prohibited activities |
| TikTok | "Our Services are provided for private, non-commercial use" | Terms of Service, opening section |
| Do not "develop, support or use software, devices, scripts, robots or any other means or processes (such as crawlers, browser plugins and add-ons or any other technology) to scrape or copy the Services, including profiles and other data from the Services" | User Agreement, section 8.2 | |
| Do not "copy, use, display or distribute any information (including content) obtained from the Services, whether directly or through third parties (such as search tools or data aggregators or brokers), without the consent of the content owner" | User Agreement, section 8.2 |
Meta, and the three sentences that get cut off

Point 3 is the sentence you will find quoted everywhere, and it is almost always cut short. Three things live past the usual cut. The clause names logging out and prohibits collection anyway, which is a phrase drafted against a specific judgment we come to below. It then reserves Meta's rights against text and data mining, which is a European copyright concept sitting in an American contract and is aimed squarely at anyone building a training set. And point 5, two sentences later, is the one that describes the reader of this article rather than the vendor: you may not "sell, license or purchase any data obtained from us or our services". Buying a competitor dataset from a scraping company is named in Meta's terms as its own prohibited act, separately from collecting it.
Point 7 completes the picture by prohibiting anything that would "circumvent, bypass or override any technological measures that Meta uses to control or limit access". Read together, the three points cover collecting, buying and getting around the fence, which is most of what the tools compared in the best Instagram scrapers do for a living.
Instagram, where the same idea is said again

Instagram repeats the prohibition in its own Terms of Use with the same "regardless of whether" construction attached, and the paragraph directly beneath it mirrors Meta’s point 5: "You can't sell, licence or purchase any account or data obtained from us or our Service, regardless of whether such data was obtained while logged in to an Instagram account." The pairing is deliberate. Collecting and buying are prohibited in adjacent sentences on both documents, which is hard to read as an accident. Note also the British spelling of "unauthorised", which tells you this is the version served to a reader in Europe. The document is localised, so a compliance note should record which version you read as well as when.
The document underneath, which almost nobody opens

This is the document that turns Meta's parenthesis, "without our prior permission", from a loophole into a process, and it is worth reading because it closes the loophole in an unusual way. Clause 3 says in terms that accepting the document is not the permission: "such permission must be obtained separately through Meta's formal authorization process". Clause 4 then restricts what a permitted collector may do with the result to running a search engine, showing previews of Meta URLs, or whatever else Meta has authorised in writing. Everything else is prohibited, "including, but not limited to, transferring, selling, licensing or sublicensing Collected Data". Read plainly, that is the business model of most of the tools on the market. Clause 2 adds a term that runs until you certify to Meta that you have deleted the data, which is an obligation most people would not expect to have signed up to.
TikTok, where the non-commercial line does the work

TikTok's scraping clause is short and unambiguous, and it is not the sentence that decides most cases. Four bullets above it in the same list sits a broader one: you may not "use the Services, without our express written consent, for any commercial or unauthorized purpose". The document also states that "Our Services are provided for private, non-commercial use". If the grant itself is non-commercial, a competitor analysis for your agency is outside it before any script runs, and no amount of care about how you collect brings it back inside. TikTok is also the platform where the sanctioned alternative is hardest to reach, which we go through in the best TikTok scrapers, and where the tools that do work are measurably the least reliable. If your interest is really in what performs rather than in raw data, the four big studies of when to post on TikTok disagree with each other more than anyone quoting one of them admits.
LinkedIn, and the clause aimed at the customer

LinkedIn's first clause is the conventional one. The second is the interesting one and it mirrors Meta's point 5: you may not copy, use, display or distribute information obtained from the service "whether directly or through third parties (such as search tools or data aggregators or brokers)". That reaches the customer of a data vendor rather than the vendor, and it is the clause a buyer is least likely to have read. If LinkedIn is the platform you actually care about, what the feed rewards is a separate question and we cover it in the LinkedIn algorithm in 2026.
The case everybody cites, and how it actually ended
If you have read anything about scraping law you have read about hiQ Labs, Inc. v. LinkedIn Corp., No. 17-16783 (9th Cir.). It is the case that established, in the Ninth Circuit, that collecting publicly available data does not violate the Computer Fraud and Abuse Act, 18 U.S.C. § 1030. The reasoning rests on the Supreme Court's decision in Van Buren v. United States, 593 U.S. 374 (2021), which read "exceeds authorized access" as a gates up or down question: either the information is off limits to you or it is not, and using information you were allowed to reach for a purpose you were not is a different thing from breaking in. That is real law, it is correctly reported, and vendors have been putting it on their home pages ever since.
Then the case went back down for the claims that were left, and in December 2022 it ended. hiQ agreed to a consent judgment of $500,000. It stipulated that LinkedIn could establish liability for breach of the user agreement, for violating the Computer Fraud and Abuse Act after all on the basis of hiQ's use of fake accounts to reach password protected pages, for the California equivalent statute, and for trespass to chattels and misappropriation. It accepted a permanent injunction requiring it to stop scraping and to delete the source code, data and algorithms it had built. hiQ, as a company, was finished.
The company that won the landmark ruling that scraping public data is legal paid half a million dollars, promised never to do it again, and deleted the product.
Both halves are true and they do not contradict each other. The Ninth Circuit answered a narrow question about a criminal statute. The contract claim, which was never the glamorous part, is the one that produced the judgment. The CFAA came back at the end anyway, not for reading public pages but for the fake accounts, which is a different act wearing the same word. One more caveat that is almost always dropped: a stipulated judgment is an agreement between the parties, not a finding by the court, so it sets no precedent. It is not law. It is what happened to them.
The two cases that went the other way
Against that, one company has beaten two platforms in eighteen months, and the reasoning in both is more useful to you than hiQ is.
Meta Platforms, Inc. v. Bright Data Ltd., January 2024
In No. 3:23-cv-00077-EMC (N.D. Cal.), Judge Edward Chen held that Meta's terms did not prohibit Bright Data's scraping of publicly available data while logged off. The reasoning is the useful part. Meta had amended its terms in 2009 to remove the language binding every visitor rather than only its users, and the court read that as Meta knowing how to write terms that reach a logged out visitor and choosing not to. The court also held that once Bright Data closed its accounts Meta could not keep binding it, and that the survival clause purporting to forbid scraping in perpetuity after termination was unenforceable.
Note what that decision is not. It is not a ruling that scraping public data is lawful in general. It is a ruling about the specific words in one contract, and a contract can be rewritten. Meta rewrote it. The clause in the screenshot above now names logging out and prohibits collection anyway, and the Automated Data Collection Terms took effect on 7 October 2024, nine months after the judgment. Winning a contract case against the company that writes the contract buys you the time it takes them to redraft.
X Corp. v. Bright Data Ltd., May 2024
In No. C 23-03698 WHA (N.D. Cal.), Judge William Alsup dismissed X's claims on entirely different ground: federal copyright law preempted them. X does not own what its users post, and it was trying to use contract and state law to exercise a copyright owner's right to exclude while keeping the safe harbours of a platform that does not own the content. Alsup wrote that X Corp "wants it both ways: to keep its safe harbors yet exercise a copyright owner's right to exclude, wresting fees from those who wish to extract and copy X users' content", and warned that giving platforms complete control over public web data "risks the possible creation of information monopolies that would disserve the public interest".
That is a district court decision rather than an appellate one, so its reach is limited. But the direction across all three cases is consistent, and it is not the direction the platforms want: the further your collection sits from an account and a login, the weaker every claim against you becomes. Which is precisely why the clauses have been redrafted to stop asking.
Logged in or logged out, and why that answer expired
Every case above turns on it, so it is worth stating alone. A contract binds the person who agreed to it. If you never logged in, never made an account, and read a page the platform serves to anybody with a browser, the argument that you accepted terms is weak, and Meta lost on exactly that.
It is also the reassurance in almost every scraper round-up currently ranking, and on Meta's platforms it is out of date. Both Meta's Terms of Service and the Instagram Terms of Use now name logging out and prohibit collection anyway. Logging out is no longer an answer to the question, because the sentence now contains the question. Whether a clause drafted that way survives contact with a court is not something we can tell you, and neither can a vendor whose landing page cites a 2024 ruling as though nothing happened afterwards. The distinction still does real work on platforms that have not redrafted, and it never applied to data protection law at all, which does not ask whether you were logged in.
What has not expired is the practical version of the same distinction, which is a different risk from the legal one. Tools that operate through your own logged in account, which is how the cheaper unofficial APIs and every browser extension work, put your account inside the blast radius. Tools that fetch public pages from somebody else's infrastructure do not involve your account at all. Both are sold under the word automation and they are not the same product. We rank on exactly that basis in the best Instagram scrapers, where distance from your own account is treated as a feature, and the same split explains why several products in the best AI social media tools refuse to touch this category.
One thing is unambiguously worse than either: creating accounts to reach content that is not public. That is the fact pattern that brought the Computer Fraud and Abuse Act back into hiQ at the end. A gate you went around is a different legal object from a rule you ignored.
In Europe, the terms are the easy part
Everything above is United States contract and computer crime law. If the profiles you collect belong to people in the EU or the UK, none of it is the binding constraint, because the moment your spreadsheet holds a name, a handle, a photograph or a biography you are processing personal data and you are the controller of it.
In July 2026 the European Data Protection Board adopted Guidelines 03/2026 on web scraping in the context of generative AI. They are written for AI training rather than competitor research, so read them for the reasoning rather than as a rulebook covering your case. The reasoning is what transfers.
- Consent is effectively off the table. You collect indirectly, at scale, with no relationship to the people involved. The Board is explicit that publishing something publicly is not consent to its collection.
- Legitimate interest under Article 6(1)(f) is the route, and it is a test you have to pass. Three cumulative limbs: a real and precisely defined interest, a demonstration that scraping is necessary with no less intrusive way to get the same result, and a documented balancing against the rights of the people in the data.
- Technical signals now count against you. The Board treats robots.txt, ai.txt, CAPTCHAs and login walls as relevant to what people reasonably expect. Ignoring robots.txt was never illegal in itself. It is now evidence in the balancing test.
- Special category data needs a second lawful basis. Anything revealing health, beliefs, politics, union membership or sexual orientation needs an Article 9(2) exception on top of Article 6, and social media bios are full of it.
- Meta's text and data mining reservation points the same way. A rights reservation of that kind is the mechanism the EU copyright framework gives a rightsholder to opt out of mining, and it is now sitting in the terms of the largest platform on the list.
The practical consequence for a small business is smaller than that sounds. Collecting public post counts, captions and engagement figures for twenty competitor companies is a long way from collecting profiles of individuals. Building a list of named people with their photographs and biographies is the thing to stop and think about, and it is also the thing most likely to be described to you as lead generation.
What actually happens, as opposed to what could
Separate the case law from the realistic outcome, because the gap is enormous and reading three lawsuits in a row distorts it. Platforms sue data brokers selling scraped social data at industrial scale. They do not sue a plumber who pulled four hundred rows about local competitors.
| What you do | What realistically happens | How likely |
|---|---|---|
| Fetch public pages through a vendor's infrastructure, no login | Nothing, or the vendor absorbs an IP block and retries | Everyday |
| Run a tool through your own logged in account | Rate limiting, an action block, then a restricted or disabled account | Common enough to plan for |
| Buy a dataset from a scraping vendor | Nothing in practice, though Meta and LinkedIn both name the purchaser in their terms | Rare, and more exposed than buyers assume |
| Create accounts to reach non public content | Account termination, and the one pattern with real legal exposure | Avoid entirely |
| Resell or publish a dataset of scraped personal data | Cease and desist, and in the EU a data protection complaint | Rare, but this is where the lawsuits are |
| Publish scraped content as your own | Copyright claim, which is separate from all of the above | Depends entirely on what you republish |
The routes that are actually sanctioned
Every platform has a front door. On each one it is narrower than people expect, and knowing how narrow saves an afternoon.
- Instagram. The Platform APIs reach Business and Creator accounts that have authorised your app, plus limited details about other professional accounts through business discovery. There is no official call that returns an arbitrary public account's posts. Full detail in the best Instagram scrapers.
- TikTok. The Research API is genuinely excellent and requires an eligible academic or non profit affiliation, an approved proposal, and research on a non commercial basis. For a company that is a closed door rather than a queue, as set out in the best TikTok scrapers.
- YouTube. The Data API is open and free at a usable quota, and cannot reliably tell you which of the videos it returns is a Short. That single gap is why a whole tool category exists, covered in the best YouTube Shorts scrapers.
- Meta, if you have a real case. The parenthesis in Meta's clause is a process, run through the Automated Data Collection Terms. It is slow, most applicants are not what it is for, and clause 4 limits what you may do with the result even if you get in. But the clause is not absolute, and articles that quote it without the parenthesis make it look that way.
What we would actually do
If the question behind your search is whether to buy a scraper, the honest answer is usually that you want the data once. Twenty competitors, their posting frequency, their formats and roughly what earns engagement, collected in one afternoon from a per result tool with no subscription, is defensible research and costs a few dollars. No clause anywhere makes reading twenty public company profiles a problem, and the data protection question barely engages because you are not collecting people.
The version that gets expensive and gets noticed is the standing pipeline: every profile, every day, indefinitely, accumulating personal data nobody has a documented purpose for. That is the shape platforms sue, the shape the EDPB guidelines are written about, and in our experience the shape that quietly stops being read three weeks after somebody builds it.
And if the real reason you are researching competitors is that your own feed has gone quiet, the data is not the bottleneck. Knowing a rival posts four times a week does not put anything on your page on Thursday. That is a different problem with different tools: what an AI social media manager actually does is the honest version of the category, the free tiers are worth understanding before you pay for anything, and if you are weighing this against hiring somebody, what social media management actually costs is the comparison that matters. If the goal is being found at all, social media SEO for small business and getting your business to show up in ChatGPT are the more useful afternoons. And if you simply need something to post on Thursday, Instagram post ideas for small business will get you further than a spreadsheet of somebody else's captions.
Frequently asked questions
Is scraping social media legal?
Collecting publicly available data is not a crime under the US Computer Fraud and Abuse Act, which the Ninth Circuit held in hiQ v LinkedIn following the Supreme Court's reading of the statute in Van Buren. That is separate from the platform's terms of service, which every major platform breaches you under, and from data protection law, which applies the moment the data describes identifiable people. Legal, in this area, is three questions rather than one.
What does TikTok's terms of service say about scraping?
TikTok's Terms of Service prohibit you from using "automated scripts to collect information from or otherwise interact with the Services". The same document states that "Our Services are provided for private, non-commercial use", which rules out commercial research independently of the scraping clause.
Does Instagram allow scraping?
No. Meta's Terms of Service state that "You may not access or collect data from our Products using automated means (without our prior permission) or attempt to access data that you do not have permission to access, regardless of whether such automated access or collection is undertaken while logged in to a Facebook account", and the Instagram Terms of Use say the same thing in its own words. Meta's separate Automated Data Collection Terms, effective 7 October 2024, add that accepting them is not the permission, which "must be obtained separately through Meta's formal authorization process".
Is it legal to buy scraped social media data?
The platforms name it separately from collecting it. Meta's Terms of Service say you "may not sell, license or purchase any data obtained from us or our services", and LinkedIn's User Agreement prohibits using information obtained from the service "whether directly or through third parties (such as search tools or data aggregators or brokers)". Enforcement against a small buyer is rare, but a buyer who assumed the terms only bound the vendor has read only half the clause.
Did hiQ win against LinkedIn?
On the narrow question, yes: the Ninth Circuit held that scraping public data does not violate the Computer Fraud and Abuse Act. On the case as a whole, no. In December 2022 hiQ accepted a consent judgment of $500,000, stipulated liability for breach of contract and, on the basis of fake accounts reaching password protected pages, for the CFAA as well, and agreed to a permanent injunction to stop scraping and delete its source code, data and algorithms.
Is it legal to scrape data if I am not logged in?
It was the strongest argument available and on Meta's platforms it has been answered. In January 2024 the court in Meta v Bright Data held that Meta's terms did not prohibit scraping public data while logged off, partly because Meta had removed the language that once bound every visitor rather than only users. Meta then rewrote both its Terms of Service and the Instagram Terms of Use to prohibit collection "regardless of whether" it happens while logged in. The distinction still matters on platforms that have not redrafted, and it has never applied to GDPR, which does not ask whether you were logged in.
Does GDPR apply to scraping public profiles?
Yes, if the profiles belong to people in the EU or UK. Public availability is not consent. The EDPB's Guidelines 03/2026 treat consent as effectively unavailable for scraping and direct controllers to legitimate interest under Article 6(1)(f), which requires a precisely defined interest, evidence that no less intrusive method would work, and a documented balancing test. Robots.txt, CAPTCHAs and login walls now count as signals about what people reasonably expect.
Will scraping get my account banned?
It depends entirely on whether the tool uses your account. Fetching public pages through a vendor's own infrastructure does not touch your account. Anything that acts as you, which includes browser extensions and most of the cheaper unofficial APIs, puts your account in the blast radius, and rate limits and action blocks are the normal first consequence.
Can I get permission to scrape rather than risk it?
On Meta, in principle. The prohibition is written as "without our prior permission", and the Automated Data Collection Terms describe the process. It is slow, most commercial applicants are not what it is designed for, and clause 4 limits permitted uses to running a search engine, showing previews of Meta URLs, or whatever else Meta authorises in writing. On TikTok the sanctioned route is the Research API, which requires an academic or non profit affiliation and non commercial use, so for a company it is closed rather than slow.
How we researched this
- Every clause quoted on this page was read in a browser on the platform's own legal page on 8 September 2026, and each one is photographed here in the document it sits in. That is more effort than quoting is usually worth and on this subject it is the whole job, because two of the five documents, the Instagram Terms of Use and Meta's Automated Data Collection Terms, render their text through scripting. A plain fetch of either returns an empty shell, which is why most articles quote them at second hand and why several quote them wrong.
- We wrote a small script for this. It loads each document in a headless browser, searches the rendered page for the clause, scrolls it to the centre and photographs what is on screen, and it fails loudly when a clause is not found rather than shooting whatever happened to be in view. A silent miss would put a screenshot of the wrong paragraph under a quotation, which is worse than having no screenshot at all. If you want to check any of this, you do not need our script: open the link in the sources below and search the page for the sentence.
- Doing it that way changed the article. Our first draft quoted Meta's clause as ending at "data you do not have permission to access". The clause does not end there. It continues "regardless of whether such automated access or collection is undertaken while logged in to a Facebook account", and then adds a reservation of rights against text and data mining, and the paragraph two below it prohibits purchasing scraped data as well as selling it. Three material things sitting past the point where the quotation is usually cut off. That is the argument of this piece, and we made the mistake ourselves before we caught it.
- The court decisions are reported from published analyses by law firms with no stake in the scraping industry, with the case name, court, judge, docket number and date given for each so any of it can be checked. We did not read the full opinions and we are not lawyers, so what each case decided is stated narrowly and the limits are stated with it.
- Nothing here is legal advice and it cannot be. Liability turns on which platform, whether you were logged in, what data you took, whose personal data it was, which jurisdiction you and they sit in, and what you did with the result. Any article that answers yes or no without asking those questions is selling something. This one is a reading of documents, which is a different job from advice about your situation.
Sources
- 1Terms of Service, section 3.2
- 2Terms of Use
- 3Automated Data Collection Terms, effective 7 October 2024
- 4Terms of Service
- 5User Agreement, section 8.2
- 6Van Buren v. United States, 593 U.S. 374 (2021)
- 7hiQ Labs, Inc. v. LinkedIn Corp., No. 17-16783 (9th Cir. 2022)
- 8LinkedIn's data scraping battle with hiQ Labs ends with proposed judgment
- 9Major decision affects law of scraping and online data collection, Meta Platforms v. Bright Data
- 10District court adopts broad view of copyright preemption in data scraping case
- 11Guidelines 03/2026 on web scraping in the context of generative AI



