用 Python 写一个 GEO 可见性检查脚本:你的网站现在能被 AI 引用吗
一、GEO 的第一步,先确认 AI 能不能爬到你
GEO(生成式引擎优化)的目标,是让 ChatGPT、Perplexity、Google AI Overviews 这类 AI 引擎在回答用户问题时引用你的内容。要做这件事,第一道门槛很朴素:这些 AI 的爬虫能不能访问你的网站。
很多网站出于隐私或版权的考虑,会在 robots.txt 里屏蔽 AI 爬虫。屏蔽本身没有对错,但它会直接决定你的内容能不能进入 AI 的语料和引用范围。这篇文章写一个脚本,帮你把这件事查清楚,顺带检查页面本身的"可引用性"。
二、要检查的三件事
- AI 爬虫访问权限:robots.txt 里有没有屏蔽主流 AI 爬虫。
- 页面正文:去掉脚本和导航后,还有没有干净、够长的正文。
- 可引用元信息:标题、作者、发布时间、结构化数据这些 AI 引用时需要的字段齐不齐。
三、先认识几个 AI 爬虫
下面这几个 User-Agent 是主流 AI 引擎的爬虫,判断规则来自各家的官方文档:
- GPTBot:OpenAI 的 ChatGPT,说明见 OpenAI 爬虫文档。
- OAI-SearchBot:OpenAI 搜索产品。
- ClaudeBot:Anthropic 的 Claude。
- PerplexityBot:Perplexity。
- Google-Extended:Google 用于训练 Gemini,允许网站选择退出,说明见 Google 爬虫文档。
robots.txt 的语法标准可以参考 robotstxt.org。
四、代码实现
先写 robots.txt 检查,把主流 AI 爬虫的访问规则列出来:
python
# -*- coding: utf-8 -*-
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
HEADERS = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/120.0 Safari/537.36"
}
AI_BOTS = {
"GPTBot": "OpenAI ChatGPT",
"OAI-SearchBot": "OpenAI 搜索",
"ClaudeBot": "Anthropic Claude",
"PerplexityBot": "Perplexity",
"Google-Extended": "Google Gemini 训练",
}
def fetch(url, timeout=10):
resp = requests.get(url, headers=HEADERS, timeout=timeout)
resp.raise_for_status()
return resp.text
def check_robots(site_url):
try:
txt = fetch(urljoin(site_url, "/robots.txt"))
except Exception as e:
print("robots.txt 获取失败:%s" % e)
return
blocks = {}
current = None
for raw in txt.splitlines():
line = raw.split("#")[0].strip()
if not line:
continue
if line.lower().startswith("user-agent:"):
current = line.split(":", 1)[1].strip().lower()
blocks.setdefault(current, [])
elif current and line.lower().startswith("disallow:"):
val = line.split(":", 1)[1].strip()
blocks[current].append(val)
print("主流 AI 爬虫访问情况:")
for bot, desc in AI_BOTS.items():
key = bot.lower()
if key in blocks and "/" in blocks[key]:
status = "被屏蔽"
else:
status = "允许(或未提及)"
print(" - %s(%s):%s" % (bot, desc, status))
再写页面可引用性检查,看正文和元信息是否齐全:
python
def check_page(url):
html = fetch(url)
soup = BeautifulSoup(html, "html.parser")
# 去掉脚本、样式、导航,估算干净正文
for tag in soup(["script", "style", "nav", "header", "footer", "noscript"]):
tag.decompose()
text = soup.get_text(separator=" ", strip=True)
title = soup.title.get_text(strip=True) if soup.title else ""
author = soup.find("meta", attrs={"name": "author"})
published = soup.find("meta", attrs={"property": "article:published_time"})
has_jsonld = bool(soup.find("script", type="application/ld+json"))
print("正文纯文本长度:%d 字" % len(text))
print("结构化数据(JSON-LD):%s" % ("有" if has_jsonld else "无"))
print("标题:%s" % (title or "缺失"))
print("作者:%s" % (author.get("content") if author else "缺失"))
print("发布时间:%s" % (published.get("content") if published else "缺失"))
if __name__ == "__main__":
site = input("请输入网址: ").strip()
check_robots(site)
check_page(site)
五、运行结果怎么看
跑一遍,你会得到两类信息:
- AI 爬虫访问情况。如果某个爬虫显示"被屏蔽",说明这个 AI 引擎暂时抓不到你的内容。
- 页面可引用性。正文太短、没有标题、没有作者、没有发布时间、没有结构化数据,这些都会降低内容被 AI 引用的概率。
六、怎么改
- 想让某个 AI 引用你,就在 robots.txt 里放行对应爬虫。判断标准很简单:内容被引用带来的价值大,就放行;涉及敏感或版权内容,就屏蔽。
- 正文别全被脚本和图片占掉,保证去掉噪音后还有足够的纯文本。
- 把作者、发布时间这类元信息补上,结构化数据也可以加上,规范参考 Schema.org 和 Google 结构化数据文档。
七、写在最后
GEO 里"能不能被 AI 看到"是地基,后面还有内容质量、品牌实体、全网信息一致性这些更重的活。这个脚本的价值在于,先帮你排除掉最基础、也最容易忽略的那一层问题。