DATA_HANDLING / IMPLEMENTATION_NOTES

Privacy Policy

How SpiderFoot.tools currently processes research queries, browser identifiers, account data, request metadata, and third-party service calls.

Last updated: September 1, 2026

This is a product-specific notice. It describes the application paths and configuration reviewed on September 7, 2026. It does not claim to audit production hosting logs, backups, provider dashboards, or settings that are not present in this repository.

Search terms can identify or relate to a person. Submit only a username, email address, or IP address that you are authorized to investigate, and do not enter passwords, authentication codes, private messages, payment data, or unrelated sensitive information. Usage rules are in the Terms of Service.

01 / SCOPE_AND_EVIDENCE

What this notice covers

This notice covers the public SpiderFoot.tools pages, username, email, and IP workflows, optional Google connection, browser- or account-linked search history, blog interactions, and the contact route.

Statements about storage and data flow are based on the current ThinkPHP controllers and services, browser JavaScript, routes, shared page template, feature flags, database columns, and local response headers. Provider-side retention and production infrastructure outside this codebase are identified as unknown rather than estimated.

02 / PRIVACY_AT_A_GLANCE

The current processing map

Activity Data involved Current destination Application storage
Visit a pageVisitor IP, request headers, path, browser/device signalsWeb infrastructure, Cloudflare, Google Analytics, loaded asset hostsFirst-party browser identifier; infrastructure and analytics retention is not set in this repository
Username scanUsername, scan categories, site definitions, returned account signalsSpiderFoot.tools origin, scan service endpoints, relevant public sites, and Serper where enabledUser record and research-history record
Email scanEmail address, one of 18 categories, returned account-context signalsSpiderFoot.tools origin, API scan service, and Serper where enabledUser record and research-history record
Target IP lookupSubmitted IPv4 or IPv6 address and returned network contextSpiderFoot.tools origin, then Cloudflare Radar server-side15-minute application cache; not added to current search history
Connect GoogleGoogle account ID, name, email, and profile image URLGoogle OAuth and the SpiderFoot.tools user tableStored in the user record; Google password is not received
View or like a blog postPost ID plus ordinary request metadataSpiderFoot.tools blog APIAggregate counters only in the blog controller; no per-reader blog action record

03 / RESEARCH_QUERY_FLOWS

Where username and email queries go

USERNAME_FLOW

  1. 01.The browser sends the username to create a first-party history task.
  2. 02.The browser sends the username and batches of locally bundled site definitions to workers.spiderfoot.tools.
  3. 03.It also sends the username and each enabled high-level category to apiv2.spiderfoot.tools.
  4. 04.Possible site names, profile URLs, category, confidence/status fields, and save timestamps are returned and saved to the task.

EMAIL_FLOW

  1. 01.The browser validates the format and sends the email address to create a first-party history task.
  2. 02.The browser sends the email address and one category at a time to apiv2.spiderfoot.tools across 18 configured categories.
  3. 03.Returned site, URL, category, status, reason, extra context, and the queried email can be saved to the task.
  4. 04.A registered, not-registered, unknown, or error response is an account-context signal, not proof of identity or mailbox control.
The scan-service runtime is not contained in this repository. The application can confirm what the browser sends to those endpoints, but it cannot use this codebase alone to establish their internal logs, downstream request details, or retention.

04 / IP_BOUNDARIES

Your visitor IP is not the same as a target IP

Visitor IP and connection headers

Cloudflare and the web server process the IP used to connect to the site. Where available, Cloudflare headers provide country, city, region, timezone, approximate coordinates, and related connection location fields. The visitor country also controls whether indexed public-web results are available.

Starting a username or email task stores the visitor IP in that history record. It also stores the site language, Cloudflare country value, and a shortened Accept-Language header in the associated user record.

IP entered for research

If the IP field is empty, the IP page returns available Cloudflare-aware request headers and does not send the visitor IP to Radar through the explicit lookup code path.

If a valid target IPv4 or IPv6 address is entered, the origin receives it and queries Cloudflare Radar server-side. The returned target data is cached for 900 seconds. The current IP workflow does not create a row in the username/email search-history table.

05 / URLS_COOKIES_STORAGE

Browser storage used by the current site

For a supported username parameter on the home page or email parameter on the email page, an early page script tries to place the value in sessionStorage under spiderfoot_pending_search. It then removes that parameter from the visible URL before the shared external page scripts load. The scanner consumes and removes the stored value.

If session storage is unavailable, the script leaves the URL fallback in place. Even when cleanup succeeds, it cannot undo the initial HTTP request, an earlier browser-history entry, network or extension access, or infrastructure logging that occurred before the script ran.

StoragePurposeConfirmed duration or behavior
osinttools_user_idFirst-party pseudonymous browser identifier used to locate the user record and historyConfigured for 2,592,000 seconds (30 days), refreshed on requests; current response uses SameSite=Lax
google_callback_referReturn location for the optional Google OAuth flow1 hour
spiderfoot_pending_searchTemporary username/email deep-link handoffRemoved when read or rejected; otherwise limited to the browser session
_ga and related Analytics stateGoogle Analytics client/session measurementControlled by Google Analytics and browser/provider settings; no duration override is present in this code
localStorageNo application use found in the audited search, account, or blog flowsNot applicable

06 / ACCOUNTS_AND_HISTORY

History exists before a Google account is connected

Every page request receives the first-party browser identifier described above. A database user is created when that browser starts its first valid username or email task. The user record stores the cookie identifier, site language, Cloudflare country value, Accept-Language value, and created/updated timestamps.

A history task stores the raw query, scan type, task identifier, completion state, normalized result records, simplified site names, visitor IP, created/updated timestamps, and—when those paths run—public-web results and AI status, errors, duration, search output, and summary output.

History is available to the current browser identifier; Google login is not required. Anyone with access to the same browser profile may therefore be able to view its username and email history. IP lookups are not displayed in this history interface.

If you connect Google, the OAuth request asks for openid, email, and profile. The returned Google subject identifier, name, email address, and profile-image URL are stored with the user record. The Google password is entered with Google, not SpiderFoot.tools, and the access token used to request profile data is not written to the user table by this code.

07 / PUBLIC_WEB_AND_AI

Separate enrichment paths, with separate data

Indexed public-web results

This feature is currently enabled except for countries blocked by the application configuration. The server sends the raw username or email query and the visitor country code to Serper. Returned result titles, links, source domains, and groupings are stored with the history task. Successful responses are also cached for seven days.

AI Search Results and AI Summary

Both panels are disabled in the current public-interface feature flags. The implemented backend uses Google Gemini. If enabled, AI Search sends the query, search type, and language context to Gemini with web-search retrieval. AI Summary sends the query, language context, and URLs extracted from saved scan and AI-search results. Generated output and task status can be stored in history, and successful AI cache entries use a seven-day application expiry.

The AI worker currently writes the provider response to an application log. AI output can be inaccurate; see the Editorial Policy for review expectations.

08 / ANALYTICS_AND_PAGE_ASSETS

What the shared page template loads

The shared layout loads Google Analytics property G-WMG0MQ892B on page load. Google describes its default implementation as collecting user counts, session statistics, approximate geolocation, browser/device information, a first-party _ga client identifier, and interaction data.

The site’s Analytics configuration supplies a page location and referrer with query strings removed. For supported username/email deep links, the earlier URL-cleanup script runs before the Analytics script. This reduces the chance of a query appearing in the configured Analytics page-view fields, but it is not a promise that no infrastructure, browser component, or provider ever sees the original URL.

No application-level consent banner or Google Consent Mode call was found before Analytics initialization. The Analytics dashboard’s retention, regional, Signals, or advertising-link settings are not stored in this repository and therefore are not described as configured.

Pages also request Google Fonts, Tailwind’s CDN script, Font Awesome through cdnjs, and jQuery from code.jquery.com. Those hosts receive ordinary browser request information such as IP address, user agent, and referrer data. The shared template uses a strict-origin-when-cross-origin referrer policy.

Advertising status: no Google AdSense, Ezoic, or other advertising-network script was found in the current application layout. This notice therefore does not describe advertising as an active application feature.

09 / SERVICES_AND_PROVIDERS

Named services in the current data path

Cloudflare

Processes site traffic and provides visitor IP/location headers; Cloudflare Radar receives a target IP for explicit IP research. Review the Cloudflare Privacy Policy.

SpiderFoot.tools scan endpoints

workers.spiderfoot.tools receives username/site-definition batches; apiv2.spiderfoot.tools receives username or email category requests. Their deployed runtime and log settings are outside this repository and need separate operational confirmation.

Serper

Receives the raw research query and country code when indexed public-web results run. Review the Serper Privacy Policy.

Google

Provides Analytics, optional OAuth profile data, result favicons, Fonts, and the implemented Gemini AI path. The data involved depends on the feature described above. Review the Google Privacy Policy and Analytics data-collection documentation.

Public sites and browser resource hosts

Username results can involve the public sites represented by the selected definitions. Opening a result sends a new request directly from your browser to that site. The home page also loads badge images from Findly.tools, LovableApp.org, SaaSTool.site, Fazier, and DANG. Tailwind, cdnjs, and code.jquery.com provide shared browser resources.

Third parties apply their own policies and retention practices. This repository does not establish where every provider processes data or how long it keeps a request.

10 / BLOG_ACTIONS

Aggregate blog interaction counters

Opening a blog detail page posts the numeric article ID to increment its aggregate view count. Pressing the like button posts the same kind of ID to increment the aggregate like count. The blog API does not create a per-reader like/view table, although normal cookies and request metadata still accompany the HTTP request and infrastructure logs may exist.

11 / RETENTION

Confirmed expiries and unresolved retention

Confirmed in code

  • Target-IP cache: 15 minutes
  • Public-web result cache: 7 days
  • Successful AI caches: 7 days
  • OAuth return cookie: 1 hour
  • Browser identifier cookie: 30 days, refreshed on requests
  • Pending-query session storage: consumed on restore or limited to the session

Not confirmed in this repository

  • Automatic expiry of user and history database rows
  • Production web, CDN, and application-log rotation
  • Backup deletion schedules
  • Google Analytics property retention settings
  • Scan-service and external-provider retention
  • Mailbox retention for support email

An application cache entry becoming unavailable after its TTL does not, by itself, prove when an expired cache file, provider copy, or backup is physically erased.

12 / CONTROLS_AND_REQUESTS

What you can control now

  • Use the history pages to review username and email tasks linked to the current browser identifier.
  • Clear spiderfoot_pending_search by closing the browser session or clearing this site’s session storage. Browser controls can also clear the site identifier and Analytics cookies.
  • Logout deletes the current site identifier cookie. A later page request creates a new identifier; logout does not delete the existing database user or history rows.
  • Google-account authorization can be managed through Google, but revoking it does not automatically erase the profile fields or history already stored by SpiderFoot.tools.

No working self-service history-delete, account-delete, or Google-unbind backend was found in the current application. For an access, correction, or deletion request, email [email protected]. We may need enough information to locate the relevant record and verify that the requester is entitled to control it; do not send a password or authentication secret.

Depending on the law that applies to your circumstances, you may have rights concerning personal data, such as requesting access, correction, deletion, restriction, or objection. Those rights and any exceptions depend on the applicable law; this notice does not state that every right applies identically to every visitor.

13 / LOGS_AND_SECURITY

What the source review can—and cannot—show

The application validates email and IP formats, checks task ownership against the current browser/account identifier, and uses timestamped signed-request validation on public-web and AI endpoints. These are implementation controls, not a guarantee that the service or any provider is risk-free.

File-based application logging is enabled. The reviewed code explicitly logs provider errors, Google OAuth binding errors, Radar errors, and AI provider responses when the AI worker runs. Web servers, reverse proxies, and Cloudflare may separately process or record visitor IP, timestamps, paths, referrers, user agent, routing data, and other request metadata for delivery, debugging, reliability, or abuse response.

No production log-retention schedule, backup policy, encryption-at-rest configuration, security-audit schedule, or complete access-control policy was found in the repository. This policy therefore does not claim any of those practices.

14 / PUBLIC_SOURCES_EXTERNAL_SITES_CONTACT

Public availability does not remove privacy impact

SpiderFoot.tools primarily processes user-submitted identifiers and signals from public or provider-accessible sources. Public availability does not mean that information is accurate, harmless to combine, or free of legal and ethical limits. Use the smallest amount of data needed, respect source rules, and verify important findings against original evidence. The Methodology explains result limits.

When you open a returned profile, research link, social link, or other external website, that site receives the new browser request and may process your IP, cookies, referrer, browser information, and activity under its own privacy practices. This Privacy Policy does not control those sites.

Questions, data requests, or corrections can be sent to [email protected]. The Contact page provides reporting guidance.

This notice may be updated when product data flows, storage, infrastructure, or providers change. The date at the top records the latest published revision.