Google Patents Search Scraper
Search Google Patents and export one row per patent: number, title, assignee, inventor, dates, grant status and links. Filter by office and date. No API key.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
queries,maxItems,inventor(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.00185 per patent = $1.85 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Patent scraped | One patent returned. Queries that match nothing, and pages that fail, are never charged. | $0.00185 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-09-20, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
queries | Keywords, quoted phrases or a CPC classification code. Each term is searched separately, one row per patent. Quoted phrases work ("solid state battery"), and so does a classification code on its own (H01M10/0525). Up to 20 terms per run. | array |
maxItems | Total rows across every search in the run. The budget is split evenly between your searches, so four terms and 400 rows gives you 100 of each. Google hands over at most 1,000 results for any single search, whatever the match count says. Keep this low while you are testing — you pay per row. | integer |
inventor | An inventor's name, as it appears on the patent. Works on its own or alongside search terms. Leave empty to search everyone. | string |
assignee | The company or institution the patent is assigned to. Works on its own or alongside search terms. Leave empty to search everyone. | string |
countries | Restrict results to these offices. Two-letter codes: US, EP (European Patent Office), WO (WIPO/PCT), CN, JP, KR, DE, GB, FR, CA, AU, IN and about forty more. Leave empty for every office. | array |
status | Granted patents only, pending applications only, or both. | string |
patentType | Utility patents, design patents, or both. | string |
language | Restrict to patents filed in one language. Leave empty for all languages. | string |
dateType | Patents carry several dates and they can be years apart. Priority is when the idea was first claimed anywhere, filing is when this application was lodged, publication is when it became public. | string |
dateFrom | Earliest date to include, as 2020-01-01. A year on its own (2020) means 1 January of that year. Leave empty for no lower bound. | string |
dateTo | Latest date to include, as 2024-12-31. Leave empty for no upper bound. | string |
onlyLitigated | Return only patents Google has a court record for. Useful when you are looking at enforcement rather than coverage. | boolean |
sortBy | Relevance is Google's own ranking. Newest and oldest sort by date instead, which changes which 1,000 results you get on a broad search. | string |
searchUrls | Build the search on patents.google.com until the results look right, then paste the address bar here. The URL's own filters are used exactly as they are, and the fields above are ignored for it. Up to 20 per run. | array |
proxyUrls | Leave this empty. By default the run rotates a large pool of addresses that cost you nothing per gigabyte. Fill it in only if you specifically want the traffic to leave through proxy servers you already pay for, in the form http://user:pass@host:port. | array |
What you get
A structured dataset — each result includes fields like:
querypublicationNumbertitleassigneeinventorpriorityDatefilingDatepublicationDategrantDateisGrantedlegalStatuspatentUrlpdfUrlsnippetExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Where this tool sits
- Categories
- Developer & Research Tools
Google Patents Search Scraper
Type what you would type into Google Patents — a keyword, an inventor, a company, a classification code, a date range — and get the results back as a table, one row per patent. Publication number, title, assignee, inventor, priority and filing and publication and grant dates, whether it was granted, whether the family is still active, and links to the page and the PDF.
Read this part first: these are search results, not patent documents. You get what the results list shows plus the abstract snippet. There are no claims, no full description, no citation lists, no family tree and no downloaded PDF — just the link to each one. And Google hands over at most 1,000 results for any single search, however many it says it found. If either of those is a problem, this is the wrong tool and you should not buy it.
- The same search grammar the website has: inventor, assignee, patent office, grant status, utility vs design, language, priority/filing/publication date ranges, litigated-only, and sort order.
- Already built a search on patents.google.com? Paste the address bar in and it runs that, filters and all.
- Every row is checked against the filters you asked for before it is delivered. A row that does not match is dropped and not charged.
- No API key, no Google account, no browser, no quota to top up.
- Empty input returns one labelled sample row, free, so you can see the columns before spending anything.
Price
$1.85 per 1,000 patents, plus a $0.0005 start fee per run.
This is a flat rate on every plan, free or paid. There are no volume tiers, no minimum spend, no subscription and no add-on fees. What you read here is what you pay on day one and on day four hundred.
| Patents | Total cost |
|---|---|
| 100 | $0.1855 |
| 1,000 | $1.8505 |
| 10,000 | $18.5005 |
| 100,000 | $185.0005 |
A search that matches nothing costs the start fee and no more. So does a search Google refuses. You are billed for patent rows and for nothing else.
What is actually charged
- One
patent-scrapedevent per patent row written to the dataset. Nothing else is metered per row. - Free: the sample row an empty run returns, and every diagnostic row — a blocked target, a dead URL, a search that matched nothing. Those rows all carry
"charged": false. - Searches that match nothing — you get a diagnostic row saying so, not a bill.
- Duplicate patents already returned earlier in the same run.
- Rows that do not match the filters you asked for; they are dropped before they are charged.
- The notice telling you a search hit the 1,000-result wall.
- A run that finds nothing costs the start fee and nothing more.
- Rows never leave the dataset without a charge, and are never charged without a row. The billed event is a named one, so there is no price quietly attached to
apify-default-dataset-item— the trick that makes some scrapers bill you for their own error messages.
Input
{
"queries": [
"solid state battery",
"lithium metal anode"
],
"countries": [
"US",
"EP"
],
"status": "GRANT",
"dateType": "priority",
"dateFrom": "2020-01-01",
"maxItems": 200
}
| Field | What it does |
|---|---|
queries | Keywords, quoted phrases, or a CPC code on its own. "solid state battery" searches the phrase; H01M10/0525 searches the classification. Up to 20 terms per run, each searched separately. |
maxItems | Total rows across every search in the run, split evenly between them. Four terms and 400 rows gives you 100 of each. Default 100, ceiling 5,000. Keep it low while you are testing. |
inventor | An inventor name as it appears on the patent. Works on its own — leave queries empty and you get that person's patents. |
assignee | The company or institution that owns it. Also works on its own. |
countries | Two-letter office codes: US, EP for the European Patent Office, WO for PCT applications, CN, JP, KR, DE, GB and about forty-five more. Empty means every office. |
status | Granted patents, pending applications, or both. |
patentType | Utility patents, design patents, or both. |
language | The filing language. Sixteen are accepted, from English to Danish. |
dateType | Which date the range applies to. Priority, filing and publication can be years apart on the same patent, and picking the wrong one is the usual reason a search looks empty. |
dateFrom / dateTo | Written as 2020-01-01. A bare year means 1 January. Either can be left empty. |
onlyLitigated | Only patents Google holds a court record for. |
sortBy | Relevance, newest first, or oldest first. On a broad search this decides *which* 1,000 results you get, so it matters more than it looks. |
searchUrls | Paste a patents.google.com address. Its own filters are used exactly as they are, and the fields above are ignored for that one. |
proxyUrls | Leave empty. Fill it in only if you want traffic to leave through proxy servers you already pay for, as http://user:pass@host:port. |
Run it with empty input and you get one clearly labelled sample row, free, so you can see the output shape before you spend anything.
Output
One row per patent. A real row from a real run:
{
"ok": true,
"charged": true,
"recordType": "patent",
"query": "solid state battery",
"publicationNumber": "US11942620B2",
"title": "Solid state battery with uniformly distributed electrolyte, and methods of …",
"assignee": "GM Global Technology Operations LLC",
"inventor": "Yong Lu",
"priorityDate": "2020-12-07",
"filingDate": "2021-12-06",
"publicationDate": "2024-03-26",
"grantDate": "2024-03-26",
"isGranted": true,
"legalStatus": "ACTIVE",
"patentUrl": "https://patents.google.com/patent/US11942620B2/en",
"pdfUrl": "https://patentimages.storage.googleapis.com/f8/02/22/a9e3950cab3340/US11942620.pdf",
"snippet": "Accordingly, it would be desirable to develop high-performance solid-state battery materials and methods that improve contacts between the solid-state active particles and the solid-state electrolyte particles in the electrodes …",
"thumbnailUrl": "https://patentimages.storage.googleapis.com/90/b8/0e/ed6c106ab89ebf/US11942620-20240326-D00000.png",
"figureCount": 12,
"resultRank": 0,
"resultPage": 0,
"totalResultsEstimate": 3799,
"scrapedAt": "2026-09-20T08:29:14.306Z"
}
Field notes
publicationNumber— the full publication number including the kind code —US11942620B2. Stable, and the right thing to use as a primary key when you re-run.title— the title as Google shows it, with the search-term highlighting stripped out.assignee— the owner at publication. It is not updated when a patent is later sold, so treat it as "who filed it", not "who owns it today".inventor— the first named inventor only. The results list does not carry the rest.priorityDate— when the idea was first claimed anywhere in the family. Usually the earliest of the four.filingDate— when this particular application was lodged.publicationDate— when it became public.grantDate— when it was granted, or null if it has not been.isGranted— true when a grant date exists. The quick way to split grants from pending applications.legalStatus—ACTIVEif any member of the family is still in force, otherwiseNOT_ACTIVE, and null when Google publishes no status. It is a family-level summary, not a jurisdiction-by-jurisdiction answer, and it is not legal advice.patentUrl— the Google Patents page. Open it for the claims and the description.pdfUrl— the original document. Sometimes null — not every record has one published.snippet— an extract from the abstract, chosen by Google around your search term. It is not the whole abstract and it is not the whole patent.thumbnailUrl— the first drawing, when there is one.figureCount— how many drawings the results list mentions. Often 0 on foreign-language records.totalResultsEstimate— Google's own count of matches for the search, copied through untouched. It is an estimate and it wobbles: the same term with a country filter can report a *higher* number than without one. Use it for a sense of scale, never as a population count.
Every real row carries "charged": true. Sample rows carry "_sample": true and diagnostic rows carry "_diagnostic": true with an errorCode you can filter on, and neither is ever billed.
How it works
- It calls the same search endpoint the Google Patents website calls when you type in the box, and asks for 100 results a page. No page rendering, no headless browser, no login, nothing to install.
- Requests leave through a large pool of rotating addresses, and the run moves to a new one every few requests rather than waiting for the throttle to arrive.
- Each page is checked for the fields it should contain before anything is delivered. A response that is not an answer — a throttle page, an interstitial — is treated as a failed request and retried from a different address, not billed as data.
- Rows are then checked against your own filters. If you asked for US grants since 2020 and something arrives that is not one, it is dropped and you are not charged for it.
- Publication numbers already seen in the run are skipped, so a patent matching two of your terms is charged once.
What people use it for
- Prior-art sweeps before filing: run the phrase four ways, dedupe on publication number, and read the shortlist instead of clicking through result pages.
- Watching what another company's R&D is filing — set
assignee, sort newest first, schedule it weekly and diff on publication number. - Mapping a technology area by classification code, then counting filings per year from the dates.
- Finding who actually works in a field: pull a few hundred rows on the topic and sort by assignee or inventor.
- Due diligence on an acquisition: everything a company has filed, with grant status and whether the family is still active, in one table.
- Feeding a spreadsheet or a database. The rows are flat and typed, so they load without cleaning.
The 1,000-result wall, and how to get round it
Google Patents will tell you a search matches 106,876 things and then hand over the first 1,000. That is the endpoint's limit, not a setting, and no amount of paying changes it.
The way round it is to cut the search into pieces that are each under 1,000 and run them all:
- By year. Set
dateTypetopriorityand run 2020, then 2021, then 2022. This is usually the cleanest split, because filings spread out fairly evenly over time. - By office. Run
US, thenEP, thenCN. Useful when you only care about some of them anyway. - By assignee. If you already know the twenty companies that matter, twenty searches beat one.
- By status. Grants and applications, separately, is a free doubling.
When a search does hit the wall you get a free notice row saying so, with Google's match count on it, so you know to split rather than quietly receiving a truncated answer.
Reading the output
Every run writes three kinds of row and they are easy to tell apart:
- Real rows carry
"charged": trueand"recordType": "patent". One billed event each. - The sample row carries
"_sample": trueand"charged": false. There is exactly one, it only appears when you ran with nothing filled in, and it exists so you can see the columns before spending anything. - Diagnostic rows carry
"_diagnostic": true,"charged": falseand anerrorCodeyou can switch on:NO_RESULTSwhen a search matched nothing or hit the 1,000 wall,BAD_INPUTwhen Google rejected the search,NETWORKwhen it could not be reached,TIME_BUDGETwhen the run ran out of time first. Each carries a plain-Englisherrorand thequeryit belongs to.
If you only want the data, filter on charged == true. That count always equals the number of events you were billed for, so the dataset is its own invoice.
Limitations
- Search results only. No claims, no full description, no citation list, no family tree, no downloaded PDF — every row carries the link, and you open it for the rest.
- At most 1,000 results per search, whatever the match count says. Split the search by year, office or assignee to reach more; there is a section above on how.
- The snippet is an extract Google picks around your search term, not the full abstract.
- Only the first named inventor is in the results list, so that is all a row can carry.
assigneeis the owner at publication and is not updated when a patent changes hands later.legalStatusis a family-level summary —ACTIVEmeans in force somewhere, not everywhere. It is not a legal opinion and should not be used as one.- Titles and abstracts of foreign-language patents are Google's machine translations, and inventor and assignee names often stay in the original script.
totalResultsEstimateis Google's estimate and it moves. Adding a filter can make it go up. Do not report it as a population count.- Sorting by relevance, newest or oldest changes which 1,000 results you see on a broad search. There is no way to page past them in any order.
- Google throttles per address. A very large run slows down as it rotates; it does not fail, but it is not instant either.
- The hard ceilings are 5,000 rows and 20 searches per run. Split bigger jobs across runs.
- Design patents carry sparse metadata — often no assignee and no snippet — because that is what the source publishes.
Questions
Do I need a Google account or an API key?
No. Nothing to sign up for, nothing to install, no quota to top up. You pay Apify per row and that is the whole arrangement.
Can I search by CPC classification code?
Yes — put it in queries like any other term, for example H01M10/0525. It is not a separate field because Google Patents does not treat it as one; the code works as a search term and it works well.
I asked for 5,000 rows on one term and got 1,000. Why?
That is the endpoint's ceiling for a single search and it applies to everyone. You will also have a free notice row saying so. Split the search by year, office or assignee — there is a section above explaining how.
What happens if a search matches nothing?
One uncharged diagnostic row with errorCode: "NO_RESULTS", and the run carries on to your other searches. You are never billed for a search that returned nothing.
Can I paste a search I already built on the website?
Yes. Put the address bar into searchUrls. Its filters are used exactly as they are, which is the easiest way to run something complicated — build it where you can see the results, then paste it.
Why is the assignee sometimes in Chinese or Japanese?
Because that is how it was filed. Google translates titles and abstracts but leaves names in the original script, and inventing a transliteration here would be making data up.
Will the run fail if something goes wrong?
No. A refused, empty or broken search produces an uncharged diagnostic row explaining what happened and the run still finishes as succeeded. A failed run would still bill you the start fee, which would mean paying to be told something went wrong.
Can I run this on a schedule?
Yes. Nothing is held between runs, so the same input is safe to repeat. Diff on publicationNumber to see only what is new, and sort newest first so the new filings arrive at the top.
How do I get exactly the rows I paid for?
Filter the dataset on "charged": true. Sample and diagnostic rows are always false, and the number of charged rows always equals the number of billed events.