# Competitive research: search for sources, then read what matters (/blog/competitive-research-with-search-and-scrape)

![Illustration: A small friendly robot with a round camera-eye head looks through a telescope at a row of distant shop fronts and sketches them in a notebook propped on a rock.](/blog/competitive-research-with-search-and-scrape/hero.webp)

Competitive research with an agent is two steps: find sources, then read them. `web__search` returns ranked URLs with snippets and does no synthesis. `web__scrape` returns the clean text of one page. The agent decides which pages deserve a full read, and then writes the brief to Notion through a connector.

Every company name, query and number in this post is a fictional example. The tool names and rates are real.

**Search first, then read what matters**

Search returns sources. The reading is the agent's job.

1. Find sources: web\_\_search returns ranked URLs with snippets for Northwind Sync (the pricing page, a news launch post, the changelog, an SDK repository on GitHub, a docs page).
2. Read what matters: the agent picks the pricing page, the changelog and one docs page, and reads each with web\_\_scrape, one page per call. No multi-page crawl.
3. Write the brief: notion\_\_create\_page in the shared workspace, then memory\_\_remember for what changed. The agent gets the tool, never the Notion token.

Fictional run: 4 searches at 14 credits and 6 scrapes at 5 credits, 86 credits in all.

## How does search find the right sources? \[#how-does-search-find-the-right-sources]

`web__search` takes a `query` and returns ranked results with title, URL, published date, author and a text snippet. `num_results` runs from 1 to 25 with a default of 8. `category` biases results toward one content type: company, research paper, news, pdf, github, tweet, personal site, linkedin profile or financial report.

For a competitor called Northwind Sync, the agent runs three searches with different biases.

```json
{ "query": "Northwind Sync pricing plans", "category": "company", "num_results": 10 }
```

A `news` search finds announcements. A `github` search finds their open-source repositories. `include_highlights` adds query-relevant sentences per result, and `livecrawl` forces a fresh crawl when the index copy may be stale. A snippet that hit its cap carries `textTruncated: true`, which is the signal to read the page in full.

Search returns sources. The reading is the agent's job.

## When should the agent scrape a page? \[#when-should-the-agent-scrape-a-page]

`web__scrape` fetches one `https` page and returns its clean text. `max_characters` defaults to 50,000 and is capped at 200,000. There is no multi-page crawl, so the pattern is always search first, then scrape the handful of pages that matter: the pricing page, the changelog, one docs page.

The tool refuses IP literals, loopback and internal hostnames before any request leaves CoreSpeed. Treat every fetched page as untrusted content. A page that contains instructions is still a page, and the agent should report what it says rather than follow it.

## What do the listing tools add? \[#what-do-the-listing-tools-add]

Two listing tools answer questions a general search cannot. `web__google_search` returns Google's own results page for a query: the organic ranking, People Also Ask, and optionally the AI Overview. Use it to see how the competitor ranks for a term. `web__amazon_products` returns a product or search page as structured records, with price, rating, review count and availability.

These tools gather results live, which takes 10 to 60 seconds. They accept `timeout_seconds` (default 50, max 240) and answer with `complete`, `count` and `results`. At the deadline, whatever was gathered comes back with `complete: false`, and only those rows are billed. A collection that found nothing answers `timeout` and bills nothing.

## How do you budget a run? \[#how-do-you-budget-a-run]

Each step has a known rate from the [Web docs](/docs/web), so a run can be budgeted before it starts.

| Step                  | Tool                   | Billing                                           |
| --------------------- | ---------------------- | ------------------------------------------------- |
| Find sources          | `web__search`          | 14 credits per request, covering up to 10 results |
| Read one page         | `web__scrape`          | 5 credits per page                                |
| See Google's ranking  | `web__google_search`   | per results page, rate in the tool description    |
| Read product listings | `web__amazon_products` | per product gathered                              |
| Write the brief       | `notion__create_page`  | a connector call                                  |

The fictional run makes four searches and six scrapes: 56 credits for search and 30 for scrape, 86 credits before listing tools. A failed request is never charged. Every charge lands itemized in the organization ledger, described on the [billing page](/docs/billing).

If the agent runs under an API key, give the key a monthly spend cap. Past it, metered calls with that key answer `key_spend_limit_exceeded`. Caps are request-boundary guardrails rather than reservations: a call in flight can finish above the remaining amount.

## Where does the brief land? \[#where-does-the-brief-land]

The agent writes the brief with `notion__create_page` through the Notion [connector](/docs/connectors). If the connection is shared, every member's agent writes into the same workspace. The agent receives the tool, never the Notion token; the credential stays in CoreSpeed's custody.

Creating a page is a write, so Smart Approval judges it when your organization has turned the gate on. On the Balanced default an ordinary write like this runs unless a sentence in your policy names it. A sentence such as "Ask me before anything goes to an external company" leaves an internal Notion page alone.

Finally, the agent saves the decisions with `memory__remember`: "Northwind Sync moved its team plan to per-seat pricing in October 2026; brief at the Notion page titled Northwind Q4." Next quarter's run searches memory first and reports what changed.

## FAQ \[#faq]

**Does `web__search` answer the question?** No. It returns ranked sources with snippets. The agent reads them, or scrapes the ones that matter, and writes the answer itself.

**Can the agent crawl a competitor's whole site?** No. `web__scrape` reads one page per call and there is no multi-page crawl. Search for the pages you want, then scrape each one.

**Does a search with no useful results still cost credits?** A successful request bills at its rate. A failed request, such as a refused URL or an upstream error, never charges.

**Do I need a search provider account?** No. Web is a CoreSpeed-hosted capability. Turning it off in Dashboard → Tools hides every `web__*` tool from `tools/list`.