Who lets AI read them?
Finance publishers are quietly deciding whether AI assistants may read their work, and almost nobody announces it. The decision is public anyway: it sits in each site’s robots.txt. This table reads that file for each site and records what it asks for, with the date. Every input is a file you can fetch yourself — including ours, which is the row most worth checking.
| Site | GPTBot | OAI-SearchBot | ClaudeBot | Claude-SearchBot | PerplexityBot | Google-Extended | Applebot-Extended | CCBot | meta-externalagent | Bytespider | Open | Named |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| StockToolsus | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | 9/10 | 10 |
| Moby | ✕ | ✓ | ✕ | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | 3/10 | 7 |
| StockAnalysis | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 10/10 | none |
| Investing.com | · | · | · | · | · | · | · | · | · | · | unread | — |
| Finviz | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 10/10 | none |
| MarketBeat | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 10/10 | none |
| TipRanks | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 10/10 | none |
| Seeking Alpha | ✓ | ✓ | ✕ | ✓ | ✕ | ✓ | ✕ | ✕ | ✕ | ✕ | 4/10 | 6 |
| Motley Fool | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | 9/10 | 1 |
| Macrotrends | ✕ | ✓ | ✕ | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | ✕ | 2/10 | 8 |
| CompaniesMarketCap | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 10/10 | none |
| SEC EDGAR (the source) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 10/10 | none |
StockTools: search=yes, ai-input=yes, ai-train=yes
Moby: search=yes,ai-train=no,use=reference
MarketBeat: search=yes, ai-input=yes, ai-train=yes
Macrotrends: search=yes,ai-train=no,use=reference
How to read this, honestly
- robots.txt is a request, not a fence. It records what a publisher askscrawlers to do. We are not observing crawler behaviour and do not claim to.
- A file we could not read is “unread”, never “allowed”. Silence is not permission, and recording it as permission would flatter exactly the sites that block hardest.
- Blocking is not wrongdoing. A publisher who spent years building a corpus has a real reason to keep it out of a training set. This table records a choice; it does not grade it.
- An open row is not always a decision. A tick means the file does not ask that crawler to stay out. It does not mean the publisher considered it. The Named column separates the two: a site that lists these crawlers by name has taken a position, while
nonemeans the file has never mentioned them and access follows from the absence of a rule. Both are honest outcomes; they are not the same outcome, and reading silence as endorsement would flatter every site that has simply not looked. - Parsed per the standard. The most specific
User-agentgroup wins over*, then the longest matching path, with Allow beating Disallow on a tie — so a site that blocks one bad actor is not reported as blocking everyone.
Our own row is in the table, not in a footnote. We allow the named AI crawlers and refuse one, and the reasoning is on the methodology page. If that ever stops being true, this table will say so before we do.