用 Python 写一个 GEO 可见性检查脚本:你的网站现在能被 AI 引用吗

用 Python 写一个 GEO 可见性检查脚本:你的网站现在能被 AI 引用吗

一、GEO 的第一步,先确认 AI 能不能爬到你

GEO(生成式引擎优化)的目标,是让 ChatGPT、Perplexity、Google AI Overviews 这类 AI 引擎在回答用户问题时引用你的内容。要做这件事,第一道门槛很朴素:这些 AI 的爬虫能不能访问你的网站。

很多网站出于隐私或版权的考虑,会在 robots.txt 里屏蔽 AI 爬虫。屏蔽本身没有对错,但它会直接决定你的内容能不能进入 AI 的语料和引用范围。这篇文章写一个脚本,帮你把这件事查清楚,顺带检查页面本身的"可引用性"。

二、要检查的三件事

  1. AI 爬虫访问权限:robots.txt 里有没有屏蔽主流 AI 爬虫。
  2. 页面正文:去掉脚本和导航后,还有没有干净、够长的正文。
  3. 可引用元信息:标题、作者、发布时间、结构化数据这些 AI 引用时需要的字段齐不齐。

三、先认识几个 AI 爬虫

下面这几个 User-Agent 是主流 AI 引擎的爬虫,判断规则来自各家的官方文档:

  • GPTBot:OpenAI 的 ChatGPT,说明见 OpenAI 爬虫文档。
  • OAI-SearchBot:OpenAI 搜索产品。
  • ClaudeBot:Anthropic 的 Claude。
  • PerplexityBot:Perplexity。
  • Google-Extended:Google 用于训练 Gemini,允许网站选择退出,说明见 Google 爬虫文档。

robots.txt 的语法标准可以参考 robotstxt.org。

四、代码实现

先写 robots.txt 检查,把主流 AI 爬虫的访问规则列出来:

python 复制代码
# -*- coding: utf-8 -*-
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                  "(KHTML, like Gecko) Chrome/120.0 Safari/537.36"
}

AI_BOTS = {
    "GPTBot": "OpenAI ChatGPT",
    "OAI-SearchBot": "OpenAI 搜索",
    "ClaudeBot": "Anthropic Claude",
    "PerplexityBot": "Perplexity",
    "Google-Extended": "Google Gemini 训练",
}


def fetch(url, timeout=10):
    resp = requests.get(url, headers=HEADERS, timeout=timeout)
    resp.raise_for_status()
    return resp.text


def check_robots(site_url):
    try:
        txt = fetch(urljoin(site_url, "/robots.txt"))
    except Exception as e:
        print("robots.txt 获取失败:%s" % e)
        return
    blocks = {}
    current = None
    for raw in txt.splitlines():
        line = raw.split("#")[0].strip()
        if not line:
            continue
        if line.lower().startswith("user-agent:"):
            current = line.split(":", 1)[1].strip().lower()
            blocks.setdefault(current, [])
        elif current and line.lower().startswith("disallow:"):
            val = line.split(":", 1)[1].strip()
            blocks[current].append(val)
    print("主流 AI 爬虫访问情况:")
    for bot, desc in AI_BOTS.items():
        key = bot.lower()
        if key in blocks and "/" in blocks[key]:
            status = "被屏蔽"
        else:
            status = "允许(或未提及)"
        print("  - %s(%s):%s" % (bot, desc, status))

再写页面可引用性检查,看正文和元信息是否齐全:

python 复制代码
def check_page(url):
    html = fetch(url)
    soup = BeautifulSoup(html, "html.parser")

    # 去掉脚本、样式、导航,估算干净正文
    for tag in soup(["script", "style", "nav", "header", "footer", "noscript"]):
        tag.decompose()
    text = soup.get_text(separator=" ", strip=True)

    title = soup.title.get_text(strip=True) if soup.title else ""
    author = soup.find("meta", attrs={"name": "author"})
    published = soup.find("meta", attrs={"property": "article:published_time"})
    has_jsonld = bool(soup.find("script", type="application/ld+json"))

    print("正文纯文本长度:%d 字" % len(text))
    print("结构化数据(JSON-LD):%s" % ("有" if has_jsonld else "无"))
    print("标题:%s" % (title or "缺失"))
    print("作者:%s" % (author.get("content") if author else "缺失"))
    print("发布时间:%s" % (published.get("content") if published else "缺失"))


if __name__ == "__main__":
    site = input("请输入网址: ").strip()
    check_robots(site)
    check_page(site)

五、运行结果怎么看

跑一遍,你会得到两类信息:

  1. AI 爬虫访问情况。如果某个爬虫显示"被屏蔽",说明这个 AI 引擎暂时抓不到你的内容。
  2. 页面可引用性。正文太短、没有标题、没有作者、没有发布时间、没有结构化数据,这些都会降低内容被 AI 引用的概率。

六、怎么改

  1. 想让某个 AI 引用你,就在 robots.txt 里放行对应爬虫。判断标准很简单:内容被引用带来的价值大,就放行;涉及敏感或版权内容,就屏蔽。
  2. 正文别全被脚本和图片占掉,保证去掉噪音后还有足够的纯文本。
  3. 把作者、发布时间这类元信息补上,结构化数据也可以加上,规范参考 Schema.org 和 Google 结构化数据文档。

七、写在最后

GEO 里"能不能被 AI 看到"是地基,后面还有内容质量、品牌实体、全网信息一致性这些更重的活。这个脚本的价值在于,先帮你排除掉最基础、也最容易忽略的那一层问题。

相关推荐
小易老师AI实战3 小时前
RLHF深度详解(超通俗+原理+工程+对比):大模型对齐的核心基石
人工智能·大模型·sft·rlhf·ppo·人类反馈强化学习·llm 对齐
资深电气设计3 小时前
高压直流母线系统测试是什么?宜迈思液冷直流负载方案技术说明
人工智能
数智工坊3 小时前
视觉SLAM第12讲|地图构建:单目稠密重建、RGB-D点云与八叉树地图全解析
人工智能·深度学习·矩阵·机器人
西柚小萌新3 小时前
【LLM&&AI应用开发 八股文】--4.3.Agent智能体(下)
java·开发语言·数据库
北极有牛4 小时前
cpp学习笔记--常量指针
java·开发语言·算法
回眸&啤酒鸭4 小时前
【回眸】OpenSwarm 多智能体协作系统实战指南
大数据·前端·人工智能
博图光电4 小时前
Libra 27105相关技术参数
人工智能·数码相机
半杯咖啡半行码4 小时前
C++ 编程与 STL 模板:从泛型编程到内存安全详解
开发语言·c++
caoerzhong4 小时前
JeeWMS 开源 WMS 部署避坑指南:Java 仓库管理系统的环境基线、四类根因与可复现交付
java·开发语言·开源