Home

The Blind Spots of AI

Treasure troves of information your agent can't reach

Noir illustration of an investigator with a magnifying glass on a dark wet street under a starry sky

Why AI misses public information

A web-enabled agent does not search the whole internet. It sends a few queries to a search index, opens a small selection of the highest-ranked results, and builds its answer from what fits into its context.

Much of the useful web does sit outside that path. Company registers, regulatory filings, public records, and specialist archives often live inside searchable databases rather than ordinary pages. They may require a form submission, a date range, or a CAPTCHA before showing a result. A person can search them, but a crawler may never see what is inside.

Localization creates another blind spot. Pages can change with the visitor's location or language, while local search engines, local scripts, and regional terminology lead to sources that an English query never reaches. Even information that has been indexed can disappear below newer, more popular, or better-optimized pages.

Mainstream sources may be easier to trust, but breakthroughs often come from the most unexpected corners.

But we can do something about its blind spots. We can step outside the default search and look in the places it routinely passes by: from ordinary marketing tools, through geography and local knowledge, to the more technical traces businesses leave behind.

This is only a panorama of open-source intelligence, usually shortened to OSINT: collecting and connecting information that is lawfully available to the public. Later articles will unpack individual categories in more detail. Follow me on LinkedIn if you do not want to miss them.

Business contacts: email addresses by role

A company website often gives you info@company.com, which is the email equivalent of putting a message in a bottle. If you are offering a service, reporting a problem, or proposing a partnership, the useful recipient is usually the CEO, procurement lead, head of marketing, or whoever owns the problem.

Hunter's Domain Search gathers professional addresses associated with a domain and shows roles, source pages, likely address patterns, and confidence scores. If you already know a person's name, its Email Finder combines that name with the company domain. This lets you address the relevant person instead of hoping a general inbox forwards your message before the next ice age.

However, an inferred address is still an inference. A deliverable mailbox does not imply consent to receive bulk marketing, and a high confidence score is not a relationship. Check the source, verify the role, write a message that is actually relevant, and follow the direct-marketing rules that apply where the recipient lives.

Company records: filings and registries

Marketing pages tell you what a company hopes to become. Regulatory filings tell you what it must disclose while lawyers are watching.

For US public companies and other SEC filers, the SEC's EDGAR system provides millions of filings for free. A 10-K annual report describes the business, markets, competition, risks, subsidiaries, and financial results. A 10-Q updates the picture quarterly. An 8-K reports significant events such as acquisitions, leadership changes, new debt, or major agreements. The full-text search reaches filings submitted since 2001.

From a sales perspective, this is a list of expensive problems written by the prospect. Search a target's filings for phrases such as "major customer", "supply agreement", cybersecurity, expansion, or the name of your service category. Compare how several companies describe the same risk. Their wording reveals what the market worries about before it becomes a slogan on LinkedIn.

EDGAR is powerful, but it does not cover every American business. Private companies are generally registered at state level. Other regions have their own doors:

A registry record confirms that an entity exists and what it filed. It does not confirm that the company is solvent, honest, or good at its job. Still, comparing the legal name, registration number, address, officers, and filing dates with a proposal or invoice can prevent expensive misunderstandings.

Documents and spreadsheets: advanced Google search

Before reaching for an exotic tool, learn to make a normal search engine uncomfortable. Google-fu is the informal name for using operators and careful phrasing to narrow a search instead of adding more hopeful words.

Quotation marks demand an exact phrase. site: limits results to a domain. filetype: asks for a format such as PDF or XLSX. A minus sign removes noisy terms. Google's search controls document these operators, which also work with indexed spreadsheets, presentations, and office documents.

site:example.com "customer case study"
site:example.com filetype:pdf "supplier code"
site:gov.uk filetype:xlsx "contract award" "Company Name"
"Company Name" (distributor OR reseller) "Poland"
"Company Name" -jobs -careers

The useful material is often boring enough to escape ordinary rankings: procurement spreadsheets, grant recipients, price lists, product catalogues, conference presentations, and technical manuals. One public Excel file can answer a market question that ten glossy reports dance around.

However, accidental exposure is not permission. If a search finds passwords, personal records, or obviously confidential files, do not collect or exploit them. Tell the owner when doing so is safe and useful.

Regional information: alternative search engines

Search engines do not keep identical copies of the web. They run different crawlers, discover different links, refresh at different speeds, and rank with different incentives. Consequently, a page buried on Google may be ordinary on another engine.

Bing has its own large index and useful operators including site:, filetype:, and contains:. Brave Search and Mojeek maintain independent indexes, so they can return genuinely different pages.

Then there are engines that live closer to the information:

  • Baidu is essential for the mainland-Chinese web. Search using a company's Chinese legal name or Unified Social Credit Code, not an English brand approximation.
  • Naver reaches South Korean news, blogs, cafés, maps, shopping, and Knowledge iN material that global engines often underrepresent. Korean search terms transform the results.
  • Yandex Search is valuable for sources from Russia, Belarus, Kazakhstan, Uzbekistan, and other Russian-speaking markets, especially when you try Cyrillic names and local place spellings.
  • Seznam runs a Czech-focused index and can surface local companies, mentions, and map results differently from Google.

The deeper lesson: localized information is often hidden by our query, not by secrecy. Search for the legal entity instead of the brand, the local script instead of English, and the regulator's vocabulary instead of your industry's marketing term.

Images: reverse search and provenance

An image can reveal a product, a place, a copied campaign, or the first appearance of a claim. No visual search engine covers everything, so use a sequence rather than declaring one winner.

Start with the cleanest, highest-quality version you can obtain. Search the whole image, then crop around distinctive details such as a logo, package, façade, sign, artwork, or unusual piece of furniture. A useful route is:

  • For a product, object, plant, or place, try Google Lens first, then Bing Visual Search, then Yandex Images. They have different indexes and different ideas about what “similar” means.
  • For provenance, copyright questions, or an earlier upload, use TinEye. Sort by biggest to seek a better copy and by oldest to find the earliest version TinEye crawled. “Oldest” does not prove who created it, though.
  • For illustration, anime, manga, or game art, use the specialist indexes SauceNAO and IQDB before returning to a general engine.

A crop should isolate a distinctive part, but do not crop blindly. SauceNAO performs better with the complete artwork when possible, while Google Lens often benefits from a tight selection.

Places: maps and historical imagery

The obvious choice is Google Maps. Street View can answer mundane but valuable questions: Does the supplier appear to occupy the address? Is there loading access? Has the storefront changed names? Where is the entrance?

In locations Google has photographed more than once, enter Street View and choose See more dates. Google's historical Street View instructions show how to move through the dated thumbnails. The archive is not continuous or available everywhere, but it can establish how a street looked at several moments instead of only today.

Google Earth pulls the camera back. Historical satellite and aerial imagery lets you compare development, land use, roofs, yards, roads, and construction over time. The timeline can reach back decades where imagery exists, but there is no universal 1984 starting line and some select archives are older. Earth also adds 3D terrain, elevation profiles in Google Earth Pro, distance and polygon measurements, KML/KMZ overlays (map files containing locations, routes, shapes, and image layers), and collaborative geographic projects.

Yandex Maps adds another perspective across Russia, Belarus, Kazakhstan, Turkey, and Central Asia. Its Panoramas include street, interior, and some aerial imagery. More distinctively, Yandex Mirrors arranges user-contributed car and pedestrian snapshots into dated, playable routes. This can fill a road-level gap or show a journey rather than a single photosphere, although coverage and quality are uneven. Its detailed map data can also be particularly useful in the regions where Yandex is the local default.

Mapillary goes further off the official road. Anyone can upload sequences from a phone, dashcam, bike, or action camera, creating coverage on trails, pedestrian streets, rural roads, and construction routes missed by commercial mapping cars.

The point is not just to see a place, but to see it from many angles. An industrial complex might look impressive from the polished road captured by Google Street View, while a dirt road behind the site could draw a very different picture.

Websites: subdomains and hidden properties

A company's main domain is its reception desk. Subdomains are the doors down the side: shop.example.com, partners.example.com, de.example.com, status.example.com, or a campaign everyone forgot after the launch party.

C99's Subdomain Finder and Pentest-Tools' Subdomain Finder combine sources such as DNS data, certificate-transparency logs, search engines, and cached observations. The names alone can reveal regional operations, acquired brands, customer portals, documentation, hiring systems, and products not linked from the corporate homepage.

Servers: Shodan and Censys

Google indexes pages. Shodan and Censys index servers: observed ports, software banners, TLS certificates, hostnames, network owners, and other technical metadata.

This sounds like something hackers would use, and yes, they do. So do defenders, insurers, researchers, vendors, and companies trying to understand their own external footprint. You can look for public services associated with your domain, see whether an old product name remains on a certificate, estimate where a technology is used, or spot a supplier's forgotten remote-access portal before trusting a security claim.

Shodan's query language uses a readable name:value form:

hostname:example.com
hostname:example.com port:443
org:"Example Corporation"
hostname:example.com http.component:wordpress

Censys Query Language searches structured fields and relationships:

host.dns.names: "example.com"
host.services: (protocol="HTTP" and port=443) and host.dns.names: "example.com"
web.software.product: "WordPress"

These queries search the providers' existing indexes. Results can be old, misattributed, or shared between unrelated customers on the same infrastructure. They are leads, not proof, and certainly not authorization to connect, test a password, or exploit a weakness.

More OSINT sources

Sources you can check:

If you find something genuinely helpful that belongs in this article, please send it to me. I would be glad to keep the list useful rather than merely long.

I've intentionally left out tools for researching the private lives of people, and social media in that context. Stalk businesses and nation states, but never individuals.