Pular para o conteúdo
← Back to Skalablog

Published article

Add Custom ScrapeBox Proxy Sources for Free — Part 2

Software EngineeringStripe

ScrapeBox proxy sources can be extended beyond the built-in list by opening the proxy harvester, choosing add and edit sources, and pasting in the URLs you want to harvest from. This article walks through the exact click path, how to find candidate pages with a Google search, and how to keep the sources that still return working proxies.

Part 2 of a series, after How to Add Proxy Sources in Scrapebox.

How to Add Custom ScrapeBox Proxy Sources

Custom ScrapeBox proxy sources are added through the proxy harvester's source editor: open Manage in the bottom-left area of the main interface, click Harvest Proxies, then click Add and Edit Sources. Each entry takes a source name, the page URL to scrape, and optional include or exclude filters. ScrapeBox is a Windows desktop SEO toolkit from SweetFunny; the same click path applies in the 2.x interface. Review the official ScrapeBox documentation for the version you have installed, because the proxy harvester module lives under different menu labels in some builds.

The pre-installed sources are a starting point, not a permanent list. Some produce far more working proxies than others, and the weak ones are worth swapping out one at a time rather than deleting the whole set. That way the harvester keeps running while you test replacements.

The practical steps are:

  1. Open ScrapeBox and click Manage in the bottom-left corner of the main window.
  2. Click Harvest Proxies to open the proxy harvester module.
  3. Click Add and Edit Sources.
  4. Enter a name for the source, the URL to scrape, and any contain or exclude text filters.
  5. Save the source, then run a harvest and test the results.

Filters matter more than they look. The 'extract URLs containing' and 'URLs must contain' fields let you keep only the parts of a page that are likely to hold proxy entries, while 'URLs must not contain' Stripe out navigation, comment, and ad links that would otherwise waste a harvest run.

Free proxy list pages can be found with a plain search query such as 'free proxy lists for scrapebox updated daily'. The results surface forums, aggregator pages, and directories that publish proxy addresses, and some of those pages publish thousands of candidate sources at once. Treat every result as a candidate page rather than a verified feed: you still have to harvest from it and test what comes back.

Forum threads are the most common hit, and they have a built-in drawback. A public thread that shares proxy lists is read by everyone else searching the same query, so the addresses on it are likely to be over-used and short-lived. They are still worth harvesting, but they should not be the only source you keep.

Aggregator pages can be more productive. A single directory page can list hundreds or thousands of individual proxy sites, and you can feed that whole list into the source editor instead of typing entries one by one. Because the list is a directory of sites and not a guarantee, expect a wide spread in yield across them.

One practical caution: a page that advertises 'verified' or 'updated daily' proxies is making a claim about its own freshness, and nothing in ScrapeBox validates that claim before harvesting. The only verification that matters is whether the harvested addresses pass your proxy test afterward.

How the Proxy Source Editor Filters URLs

The source editor has three fields beyond the name and URL, and they decide which links inside a page get scraped. 'Extract URLs containing' and 'URLs must contain' act as include rules, while 'URLs must not contain' acts as an exclude rule. Used together, they turn a messy HTML page into a narrower set of candidate URLs, which reduces the number of dead entries you have to test.

Include filters are most useful when a site lists proxies on sub-pages rather than on the landing page. If every proxy entry follows a predictable path, such as a dated subfolder or a numbered page, an include filter can walk those pages instead of pulling the whole domain.

Exclude filters do the cleaning work. Common exclusions are login, register, contact, privacy, and tag pages, plus any path segment that appears in a site's navigation on every page. Without exclusions, a harvest run spends much of its time fetching the same boilerplate links from every source in the list.

The filters operate on the URLs found in the page, not on the proxy strings themselves. If you need to separate working proxies from dead ones, that happens after the harvest, during testing, not inside the source editor.

ScrapeBox Proxy Sources Compared by Type

Not every source type behaves the same way. The table below compares what each kind of source typically gives you, based on how these pages are structured and shared publicly.

Source typeYield per pageFreshnessShared with others
Built-in sourcesVaries; some strong, some deadUnknownYes, shipped with the tool
Public forum threadModerateErraticYes, high
Aggregator directoryHigh page count, mixed qualityDepends on the list behind itYes, variable
Site that publishes its own daily listModerate, more consistentOften updated on a scheduleYes
Paid proxy providerPredictableMaintainedNo

The built-in sources are the ones the tool ships with, and their performance changes over time as the underlying pages change. Public forum threads and aggregator directories are free options with the sharing problem described above. A site that publishes its own list on a schedule tends to hold quality longer than a one-off thread, though the addresses are still public. Paid providers sit outside the free-harvest workflow and are the subject of a separate buying decision.

The pattern worth noticing is that free sources and long-lived proxies pull in opposite directions. The more widely a list is shared, the faster its addresses are consumed and blocked.

How Long Do Harvested Free Proxies Last?

Harvested free proxies are temporary by nature. They come online and go offline without notice, and the ones on heavily shared lists tend to stop working soonest because more people are routing traffic through the same address. The transcript puts it plainly: they will be online, then they will be off. That is why the source list needs periodic replacement rather than a one-time setup.

This is the argument for keeping a larger set of sources than you currently need. If a third of your sources stop producing, a list with twenty entries still has a working core, while a list with three entries may have nothing.

It also explains why testing has to follow harvesting rather than be skipped. A harvest run returns whatever the source page contained at that moment, including addresses that were already dead when the page was published or that expired between the page update and your run.

Budget the time accordingly. Free proxy harvesting is a repeating task, and the click path in this article is short enough to repeat whenever the supply drops.

A Testing Loop for Custom Proxy Sources

A workable routine is to harvest, test, and then prune the source list based on what actually returned working addresses. ScrapeBox includes a proxy tester for exactly this step, and the results tell you which sources earn their place in the list.

Run the loop like this:

  1. Harvest from the full source list and save the results.
  2. Test the harvested addresses in the proxy tester and note the pass rate.
  3. Identify the sources that contributed few or no working proxies.
  4. Replace those entries in Add and Edit Sources with new candidate pages.
  5. Repeat on a schedule, keeping the strongest sources in place.

Keep a written note of each source's rough yield. After a few cycles, the pattern becomes clear and you can stop testing marginal sources as often as the good ones.

If the free list never stabilizes, that is a signal rather than a failure. The next video in the series covers paid proxy options, which trade cost for addresses that are not being shared with every other person running the same search.

FAQ

  • How do I add a custom proxy source in ScrapeBox? Open Manage, click Harvest Proxies, then click Add and Edit Sources. Enter a source name, the URL to scrape, and any contain or exclude filters, then save. The new source joins the list the harvester uses on the next run.
  • Where can I find free proxy list pages to scrape? Search for phrases such as 'free proxy lists for scrapebox updated daily'. Forums and aggregator directories dominate the results, and one directory can list hundreds or thousands of candidate sites for the source editor.
  • Why do harvested free proxies stop working? Free proxies come online and go offline without notice, and widely shared lists burn through addresses fastest. Any source page can also go stale between its own update and your harvest run.
  • What do the contain and exclude filters do? They decide which links inside a source page get scraped. Include rules keep URLs matching a pattern, and exclude rules drop boilerplate such as login, contact, and tag pages before they waste a harvest.
  • Should I delete built-in sources that stopped working? Yes, replace them one at a time with sources that return live addresses. Swapping entries individually keeps the harvester usable while you evaluate replacements instead of clearing the whole list.
  • Is a 'verified daily' proxy page actually verified? That is the page's own claim. ScrapeBox does not check freshness before harvesting, so the only verification that counts is whether the addresses pass your own proxy test.
  • How often should I refresh my proxy sources? There is no fixed interval. Refresh when the tested pass rate drops noticeably, since that points to source pages that have gone stale or been over-shared.
  • Can I feed a whole list of site URLs into ScrapeBox at once? The source editor is built for individual entries, but an aggregator directory gives you the candidate URLs. Work through them in batches and keep the ones that produce working proxies.
  • Do I need paid proxies to use ScrapeBox? No. Free harvesting works, but it is a recurring task with an unstable supply. Paid providers are an alternative for when the free list cannot keep up with your scraping volume.

Source video