> ## Documentation Index
> Fetch the complete documentation index at: https://docs.shieldlabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Web scraping

> Run the standalone web-scraping app and compare its starter and final versions.

## What you will build

Unverified or automated searches receive no flight prices. This is an illustrative application policy, not a default rule that ShieldLabs applies to every customer.

## Run the starting application

You need Node.js 22 or later. The starter runs without ShieldLabs keys; the final version needs a registered HTTPS hostname and matching keys.

Clone the public [tutorial repository](https://github.com/ShieldLabs-ai/use-case-tutorials), then start this standalone application:

```sh theme={null}
git clone https://github.com/ShieldLabs-ai/use-case-tutorials.git
cd use-case-tutorials
git switch starter
cd web-scraping
npm ci --omit=dev
cp .env.example .env
npm run dev
```

Open [http://127.0.0.1:3000](http://127.0.0.1:3000). The starter has no ShieldLabs dependency or identification check. Its own SQLite state is separate from every other tutorial.

## Add the integration

Stop the starter server. Follow these changes, then run the complete implementation from the `final` branch.

<Steps>
  <Step title="1. Prepare the keys and HTTPS hostname">
    The `.env` copied from `starter` contains no key entries yet. Add the two lines below with this domain's real values:

    ```dotenv theme={null}
    SHIELDLABS_PUBLIC_KEY=your-public-key
    SHIELDLABS_API_KEY=sec_your_private_api_key
    ```

    Register the app's hostname in **Integration > Domains**. Put its matching Public Key and Private API Key in the private `web-scraping/.env` as `SHIELDLABS_PUBLIC_KEY` and `SHIELDLABS_API_KEY`. Serve this app through HTTPS on that registered hostname. The Private API Key stays on the server; a customer key does not automatically authorize localhost.
  </Step>

  <Step title="2. Identify the action in the browser">
    The `public/index.html` page loads the locally served JS SDK. In `public/index.js`, each action takes a fresh Request ID and sends it with the action:

    ```js theme={null}
    search.requestId = await identification.take();
    const data = await postJson('/api/flights', search);
    ```

    Only the Request ID goes to the server; the browser does not supply the risk result.
  </Step>

  <Step title="3. Verify the identification on the server">
    `server/shieldlabs.js` uses `@shieldlabs-ai/node` and the Private API Key to find that exact Request ID in History. The action route passes it to `server/flights.js`, where `verifyIdentification(requestId)` rejects missing, stale, replayed, automated or unusable checks before this scenario's rule is applied.
  </Step>

  <Step title="4. Apply the automated scraping rule">
    This excerpt from [`server/flights.js`](https://github.com/ShieldLabs-ai/use-case-tutorials/blob/final/web-scraping/server/flights.js) shows the decision's core:

    ```js theme={null}
    const check = await verifyIdentification(requestId);
    if (!check.ok) {
      return { success: false, flights: [], message: `No flights shown: ${check.message}` };
    }
    ```

    `server/flights.js` requires a verified identification before returning any generated flight schedule. Missing checks, automation, disabled JavaScript or a Dangerous result return no flights.
  </Step>

  <Step title="5. Run the completed application">
    From the repository root, compare the two versions and start the final app:

    ```sh theme={null}
    cd ..
    git diff starter origin/final -- web-scraping
    git switch final
    cd web-scraping
    npm ci --omit=dev
    npm start
    ```

    A fresh clone has the remote `origin/final` ref even before its local `final` branch exists. Open the completed app through the registered HTTPS hostname. The Node server listens on `127.0.0.1:3000` behind your reverse proxy.
  </Step>
</Steps>

## Follow the integration code

The completed source is in the [web-scraping application](https://github.com/ShieldLabs-ai/use-case-tutorials/tree/final/web-scraping). Read these files in order:

1. [Browser helper](https://github.com/ShieldLabs-ai/use-case-tutorials/blob/final/web-scraping/public/shieldlabs.js): loads the installed SDK, serializes checks and returns a request ID.
2. [Server verification](https://github.com/ShieldLabs-ai/use-case-tutorials/blob/final/web-scraping/server/shieldlabs.js): reads real History, checks the matching ID, rereads delayed results and rejects unusable or replayed checks.
3. [Scenario decision](https://github.com/ShieldLabs-ai/use-case-tutorials/blob/final/web-scraping/server/flights.js): applies this app's rule using the verified identification and stores its sample state.

The [finished application](https://github.com/ShieldLabs-ai/use-case-tutorials/tree/final/web-scraping) keeps its browser, server and SQLite state inside this folder. The tests use synthetic History responses; the normal server does not manufacture a successful identification.

## Try the completed application

1. Pick two different sample airports and a date in a normal browser. The app returns an illustrative flight schedule.
2. Run `npm test` in `web-scraping` to inspect the automated and missing-identification paths. Both return an empty flights array.
3. Directly calling the API without a fresh Request ID must not reveal the schedule.

Use **Reset demo DB** with `DEMO_ALLOW_RESET=1` to repeat the exercise. These controls are for a disposable demonstration, not production authorization.

## Check your result

In DevTools **Network**, inspect the /api/flights request and copy its `requestId`. Find the successful search Request ID in History. The local search row stores its verified Device ID and Request ID. The returned flights are deterministic teaching data, not live travel inventory. Run `npm run check` and `npm test` from `web-scraping` to exercise failure cases locally without making more identifications.

For a real run, obtain fresh request IDs in the browser and confirm their Device IDs and signals in History. If identification returns HTTP 429, stop and check your account limits before retrying. A cookie change does not prove that the Device ID changed: compare the History rows.

## Adapt it to your product

This example is a starting point, not a production-ready authorization system. Bind each action to an authenticated session where appropriate, authorize administration screens, persist state and atomically consume request IDs across all server instances. Review legitimate shared-device behavior before applying a device-based restriction. See [Acting on results](/guides/acting-on-risk-score).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.