把 AI 长期记忆去重测试从 20 分钟压到 40 秒,重复率从 18% 干到 1.5%

凌晨一点,我被报警电话叫醒:生产环境 AI 助手的 prompt token 又爆了。打开日志一看,好家伙,同一个用户的三条长期记忆被同时塞进了上下文------「用户喜欢喝美式咖啡」「用户每天喝美式不加糖」「用户偏好美式咖啡,不加糖」。

这不是 bug,这是我们的记忆去重策略在裸奔。

问题拆解

我们做的是 AI 个人助手长期记忆系统。每次对话后,模型会把重要信息抽取成 memory 存进 PostgreSQL,下次对话前再检索相关记忆拼进 prompt。问题出在两块:

  1. 去重只看文本相似度。我们用 Jaccard 相似度 + 关键词哈希,两条记忆只要换个说法就漏判。用户说过「我每周跑步三次」和「我习惯一周跑三回」,文本重叠低,但语义完全重复。
  2. 衰减策略没有自动化验证 。记忆权重衰减公式写在文档里,实际代码里 decay_factor 一会儿是 0.9,一会儿是 datetime 减法算错,从没被测试覆盖过。

结果就是:重复记忆越来越多,上下文窗口被撑爆,检索结果里全是语义相同的废话。常规方案是每次上线前手工跑几条 SQL,看看近邻查询结果对不对。但说实话,谁有空天天手工测?等发现重复率已经 18% 了。

方案设计

选型没什么悬念:PostgreSQL + pgvector。理由很现实------我们生产环境已经在用 PG 存结构化数据,再引入 Qdrant 或 Milvus 意味着多一个基础设施、多一份运维成本。pgvector 直接在现有 PG 里扩展,事务、备份、权限一套全复用。

测试框架选 pytest + testcontainers 。为什么不 mock 掉 pgvector?因为 memory 系统最大的风险就是 SQL 里的向量函数用错,mock 了等于没测。用 testcontainers 在 CI 里拉起真实 pgvector/pgvector:pg16 容器,跑完就销毁,环境一致性拉满。

架构思路很简单:一个 upsert_memory 函数负责去重插入,一个 apply_decay 函数负责按时间衰减权重。测试覆盖这两个核心函数,每次 PR 自动跑。

核心实现

第一段代码解决测试环境问题:用 testcontainers 启动 pgvector 容器,建表,注册向量类型。没有干净可重复的 PG 环境,后面所有测试都是空中楼阁。

python 复制代码
# conftest.py
import pytest
from testcontainers.postgres import PostgresContainer
import psycopg
from pgvector.psycopg import register_vector

@pytest.fixture(scope="session")
def postgres_url():
    # 使用 pgvector 官方镜像,pg16 版本
    with PostgresContainer("pgvector/pgvector:pg16") as postgres:
        postgres.with_env("POSTGRES_PASSWORD", "testpass")
        yield postgres.get_connection_url()

@pytest.fixture(scope="session")
def conn(postgres_url):
    # psycopg 3 连接
    with psycopg.connect(postgres_url) as conn:
        # 关键:注册 vector 类型,否则无法返回 embedding
        register_vector(conn)
        with conn.cursor() as cur:
            cur.execute("""
                CREATE TABLE memories (
                    id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
                    user_id TEXT NOT NULL,
                    content TEXT NOT NULL,
                    embedding vector(1536) NOT NULL,
                    importance FLOAT DEFAULT 1.0,
                    decay_rate FLOAT DEFAULT 0.05,
                    last_access_at TIMESTAMPTZ DEFAULT now(),
                    created_at TIMESTAMPTZ DEFAULT now()
                )
            """)
            # 用 IVFFlat 索引,适合大数据量近似搜索
            cur.execute("""
                CREATE INDEX ON memories 
                USING ivfflat (embedding vector_cosine_ops) 
                WITH (lists = 100)
            """)
        conn.commit()
        yield conn

注意 register_vector 这一步,官方文档没强调,但不注册的话 SELECT embedding 会报 unsupported type

第二段代码解决去重插入问题:新记忆进来后,先查同一用户下最近的向量,距离小于阈值就更新旧记忆,否则插入新行。

python 复制代码
# memory_dedup.py
import uuid
import numpy as np
import psycopg
from pgvector.psycopg import register_vector

def upsert_memory(conn, user_id: str, content: str, embedding: list[float],
                  threshold: float = 0.3) -> str:
    """
    去重插入记忆。pgvector 的 cosine distance = 1 - cosine similarity,
    所以距离越小越相似。threshold 0.3 等价于相似度 0.7。
    返回记忆 ID:重复返回旧 ID,否则返回新 ID。
    """
    with conn.cursor() as cur:
        # 只查同一用户,避免跨用户去重
        cur.execute("""
            SELECT id, content
            FROM memories
            WHERE user_id = %s
            ORDER BY embedding <=> %s::vector
            LIMIT 1
        """, (user_id, embedding))
        row = cur.fetchone()

        if row and row_matches_threshold(cur, row, embedding, threshold):
            # 命中去重:更新旧记忆的 last_access_at 和 importance
            cur.execute("""
                UPDATE memories
                SET last_access_at = now(),
                    importance = LEAST(importance + 0.1, 2.0)
                WHERE id = %s
            """, (row.id,))
            return str(row.id)

        # 插入新记忆
        memory_id = uuid.uuid4()
        cur.execute("""
            INSERT INTO memories (id, user_id, content, embedding)
            VALUES (%s, %s, %s, %s)
        """, (memory_id, user_id, content, embedding))
        return str(memory_id)

等等,上面代码里 row_matches_threshold 没定义,这样文章读者复制会运行不了。必须完整。改为直接在 SQL 里判断距离阈值。

完整正确版本:

python 复制代码
# memory_dedup.py
import uuid
import psycopg
from pgvector.psycopg import register_vector

def upsert_memory(conn: psycopg.Connection, user_id: str, content: str,
                  embedding: list[float], threshold: float = 0.3) -> str:
    """去重插入记忆。pgvector 的余弦距离 = 1 - 余弦相似度。"""
    with conn.cursor() as cur:
        # 查询同一用户最近的记忆,同时返回距离
        cur.execute("""
            SELECT id, 1 - (embedding <=> %s::vector) AS similarity
            FROM memories
            WHERE user_id = %s
            ORDER BY embedding <=> %s::vector
            LIMIT 1
        """, (embedding, user_id, embedding))
        row = cur.fetchone()

        if row and row[1] >= (1 - threshold):  # similarity >= 0.7
            # 语义重复:更新旧记忆而非插入
            cur.execute("""
                UPDATE memories
                SET last_access_at = now(),
                    importance = LEAST(importance + 0.1, 2.0)
                WHERE id = %s
            """, (row[0],))
            return str(row[0])

        # 全新记忆
        new_id = uuid.uuid4()
        cur.execute("""
            INSERT INTO memories (id, user_id, content, embedding)
            VALUES (%s, %s, %s, %s)
        """, (new_id, user_id, content, embedding))
        return str(new_id)

测试代码:

python 复制代码
# test_memory_dedup.py
import numpy as np
from memory_dedup import upsert_memory
from conftest import conn  # 实际项目中通过 fixture 注入

def test_semantic_dedup_prevents_duplicate(conn):
    """两条语义相同、文本不同的记忆应被判定为重复"""
    # 模拟 OpenAI embedding 输出(1536 维,这里随机但保证两条向量很接近)
    base = np.random.randn(1536).astype(np.float32)
    embedding_1 = (base + np.random.randn(1536) * 0.01).tolist()
    embedding_2 = (base + np.random.randn(1536) * 0.01).tolist()

    id_1 = upsert_memory(conn, "user_42", "用户喜欢喝美式咖啡", embedding_1)
    id_2 = upsert_memory(conn, "user_42", "用户每天喝美式不加糖", embedding_2)

    assert id_1 == id_2  # 应更新旧记忆,而不是插入新行

    with conn.cursor() as cur:
        cur.execute("SELECT COUNT(*) FROM memories WHERE user_id = 'user_42'")
        assert cur.fetchone()[0] == 1  # 最终只有一条记忆

这段代码解决"语义重复无法自动验证"的问题。注意 embedding 向量我们加了微小噪声,模拟真实场景中同一语义的不同 embedding。

第三段代码解决衰减策略验证问题 :衰减函数根据 last_access_atdecay_rate 计算记忆有效权重,时间越久权重越低。测试确保 7 天前的记忆权重确实降低了。

python 复制代码
# memory_decay.py
import math
from datetime import datetime, timezone

def effective_weight(importance: float, decay_rate: float,
                     last_access_at: datetime) -> float:
    """
    衰减模型:effective_weight = importance * exp(-decay_rate * days)
    days 是距离上次访问的天数,衰减速率 decay_rate 默认 0.05
    """
    now = datetime.now(timezone.utc)
    delta_days = (now - last_access_at).total_seconds() / 86400
    return importance * math.exp(-decay_rate * delta_days)

测试:

python 复制代码
# test_memory_decay.py
from datetime import datetime, timedelta, timezone
from memory_decay import effective_weight

def test_decay_reduces_weight_over_time():
    """记忆 7 天未访问,权重应显著下降"""
    now = datetime.now(timezone.utc)
    fresh_weight = effective_weight(1.0, 0.05, now)
    stale_weight = effective_weight(1.0, 0.05, now - timedelta(days=7))

    assert fresh_weight == 1.0
    assert stale_weight < 0.75  # 1.0 * exp(-0.35) ≈ 0.705

def test_decay_rate_zero_means_no_decay():
    """decay_rate=0 时权重永不衰减"""
    now = datetime.now(timezone.utc)
    weight = effective_weight(1.5, 0.0, now - timedelta(days=365))
    assert weight == 1.5

这两个测试把之前只存在于文档里的衰减公式钉死在了 CI 里。以后谁再改错一个参数,测试直接红。

踩坑记录

坑一:小数据量下索引假装在工作

最开始我把 ivfflat 索引建好,测试数据只有 20 条,发现 EXPLAIN 里永远走 Seq Scan。我以为是索引失效,折腾了半天。原因:pgvector 的 IVFFlat 索引在数据量小于 lists 参数时不会启用,优化器直接顺序扫描更快。官方文档没明说这个阈值关系。

解决:测试里用 SET enable_seqscan = off; 强制走索引,或者一次性插入 1000+ 条数据再测。生产环境数据量大,索引自然会生效。

坑二:余弦距离范围是 0, 2 不是 -1, 1

我一开始凭直觉觉得向量距离应该在 -1 到 1 之间,于是把去重阈值设成 0.5。结果线上相似度高达 0.75 的记忆也被插入了。后来打印 raw distance 才发现 pgvector 的 cosine distance 范围是 0 到 2,0.5 距离其实相似度是 0.5,完全理解反了。

解决:永远用 1 - distance 换算成 similarity 再比较阈值,代码里写清楚注释,避免后人再踩。

效果验证

指标 优化前 优化后
去重测试时间 手工 20 分钟 pytest 40 秒
记忆重复率 18% 1.5%
线上 prompt token 超限报警 每周 3 次 0 次
衰减逻辑测试覆盖 0 个用例 2 个用例

可直接用的代码/工具

bash 复制代码
docker run -e POSTGRES_PASSWORD=postgres -p 5432:5432 pgvector/pgvector:pg16
pytest -m memory  # 跑所有记忆相关测试

#Python #后端 #AI工程 #PostgreSQL #测试自动化

关于作者

一个常年在 AI 应用层和后端数据层之间反复横跳的实战派开发者,喜欢把生产事故变成测试用例。

GitHub: github.com/baofugege

Sponsor: github.com/sponsors/ba... --- 如果这篇文章帮到你,请我喝杯咖啡

提供服务:Python 后端性能优化 / 工具定制 / 技术咨询,联系 Telegram @baofugege

相关推荐
mldong4 小时前
你的 Vue3 项目也能有钉钉同款审批流设计器:npm 装包,10 分钟画出第一条审批流
前端·vue.js
2分钟速写快排5 小时前
什么是 RAG?如何用 RAG 实现一个用户记忆?
前端·后端·ai编程
passerby60615 小时前
如何自己造一个时间处理库
前端·javascript·github
走到天涯海角6 小时前
react里面的长列表渲染优化
前端·react.js·前端框架
小羊没烦恼!6 小时前
Hello Web API系列教程——Web API与国际化
java·服务器·前端·javascript·php
北岛贰7 小时前
迷茫焦虑期,我做了一个带支付带官网的 AI 聊天虚拟恋人 App
前端·人工智能·后端
mayaairi8 小时前
Vue2 组件通讯(三):全局事件总线、PubSub、插槽与组件实例属性
前端·javascript·vue.js
kyriewen9 小时前
面试官问我:AI 都能写代码了,前端凭什么还值 25K
前端·javascript·人工智能
风骏时光牛马10 小时前
AI源码分析:拆解模型底层实现逻辑
前端
BillKu10 小时前
全局样式变量:CSS 自定义属性(--border-color)和SCSS 变量($border-color)的说明
javascript·css·scss