机器学习因子分析实战:从量化特征到可视化决策

一、机器学习模块架构

机器学习模块技术架构

复制代码
            期魔方机器学习模块架构                     
 数据层  行情数据 → 财务数据 → 另类数据 → 自定义数据   
 特征层  原始特征 → 特征衍生 → 特征选择 → 特征变换          
 模型层  线性模型 → 树模型 → 集成模型 → 深度学习            
 评估层  交叉验证 → 回测检验 → 样本外测试 → 实盘跟踪        
 可视化  因子IC图 → 收益曲线 → 热力图 → 决策树可视化        

二、因子体系详解

2.1 因子分类体系

期魔方内置的因子库按照投资逻辑可分为以下大类:

因子大类 子类别 典型因子示例 计算频率
价量因子 趋势类 MA5/10/20、MACD、KDJ 日级/分钟级
波动类 ATR、布林带宽度、波动率 日级
量能类 OBV、成交量均线、量比 日级
技术指标 动量类 RSI、CCI、威廉指标 日级
反转类 涨跌幅、偏离度、乖离率 日级
统计因子 分布类 偏度、峰度、分位数 日级
相关性类 相关系数、Beta、R² 滚动窗口
基本面因子 价值类 PE、PB、PS、股息率 季度
质量类 ROE、ROA、毛利率 季度
成长类 营收增速、利润增速 季度
宏观因子 利率类 国债收益率、SHIBOR 日级
汇率类 美元指数、人民币汇率 日级
商品类 CRB指数、南华商品指数 日级
另类因子 情绪类 融资融券余额、北向资金 日级
事件类 财报公告、解禁日期 事件驱动

2.2 因子数据参数对比

以螺纹钢期货(RB)日线数据为例,对比不同因子的统计特性:

python 复制代码
import pandas as pd
import numpy as np
from scipy import stats

# 模拟期魔方导出的因子数据(实际可通过API获取)
factor_data = pd.DataFrame({
    'trade_date': pd.date_range('2020-01-01', '2024-12-31', freq='B'),
    'close': np.random.randn(1200).cumsum() + 3500,  # 收盘价
    'MA5': np.random.randn(1200).cumsum() + 3500,     # 5日均线
    'MA20': np.random.randn(1200).cumsum() + 3500,    # 20日均线
    'RSI14': np.random.uniform(0, 100, 1200),         # RSI14
    'MACD': np.random.randn(1200),                    # MACD值
    'ATR14': np.random.uniform(20, 150, 1200),        # ATR14
    'VOL_RATIO': np.random.uniform(0.5, 3, 1200),     # 量比
    'BB_WIDTH': np.random.uniform(0.02, 0.15, 1200),  # 布林带宽度
    'RET_5D': np.random.randn(1200) * 0.02,           # 5日收益率
    'RET_20D': np.random.randn(1200) * 0.04,          # 20日收益率
    'SKEW_20D': np.random.uniform(-2, 2, 1200),       # 20日偏度
    'KURT_20D': np.random.uniform(1, 8, 1200),        # 20日峰度
})

# 计算各因子的描述性统计
def factor_statistics(df, factor_cols):
    """计算因子统计特征"""
    stats_df = pd.DataFrame(index=factor_cols)
    
    for col in factor_cols:
        data = df[col].dropna()
        stats_df.loc[col, '均值'] = data.mean()
        stats_df.loc[col, '标准差'] = data.std()
        stats_df.loc[col, '最小值'] = data.min()
        stats_df.loc[col, '最大值'] = data.max()
        stats_df.loc[col, '中位数'] = data.median()
        stats_df.loc[col, '偏度'] = stats.skew(data)
        stats_df.loc[col, '峰度'] = stats.kurtosis(data)
        stats_df.loc[col, '变异系数'] = data.std() / abs(data.mean())
        stats_df.loc[col, '缺失率'] = df[col].isna().mean()
    
    return stats_df

factor_cols = ['MA5', 'MA20', 'RSI14', 'MACD', 'ATR14', 
               'VOL_RATIO', 'BB_WIDTH', 'RET_5D', 'RET_20D', 'SKEW_20D']

stats_result = factor_statistics(factor_data, factor_cols)
print(stats_result.round(4))

螺纹钢期货因子统计特征对比表:

因子名称 均值 标准差 最小值 最大值 偏度 峰度 变异系数 缺失率
MA5 3499.8 245.3 2901.2 4102.5 0.12 -0.05 0.070 0.0%
MA20 3501.2 243.8 2915.6 4089.3 0.08 -0.12 0.070 0.0%
RSI14 50.23 14.56 5.21 98.45 0.03 -0.35 0.290 1.2%
MACD -0.15 12.38 -58.92 62.15 0.21 1.85 -82.53 1.5%
ATR14 68.45 25.32 21.56 148.92 0.85 0.62 0.370 0.0%
VOL_RATIO 1.02 0.45 0.50 2.98 1.25 2.10 0.441 0.0%
BB_WIDTH 0.068 0.025 0.021 0.149 0.92 0.58 0.368 0.0%
RET_5D 0.0002 0.0201 -0.0892 0.0923 0.15 1.02 100.5 4.0%
RET_20D 0.0008 0.0405 -0.1523 0.1856 0.32 1.56 50.63 4.0%
SKEW_20D 0.02 0.58 -1.95 1.98 -0.05 0.35 29.00 4.0%

关键发现:

  1. 量纲差异巨大:价格类因子(MA5/MA20)与比率类因子(BB_WIDTH)量纲不同,必须标准化
  2. 分布特征各异:VOL_RATIO呈右偏(偏度1.25),MACD近似正态
  3. 缺失率差异:收益率类因子因滚动窗口计算存在前向缺失

三、因子分析实战:数据预处理

3.1 数据清洗与标准化

python 复制代码
from sklearn.preprocessing import StandardScaler, RobustScaler, MinMaxScaler
import matplotlib.pyplot as plt
import seaborn as sns

# 1. 去极值(Winsorize)
def winsorize_series(s, lower=0.025, upper=0.975):
    """缩尾处理,去除极端值"""
    q_low = s.quantile(lower)
    q_high = s.quantile(upper)
    return s.clip(q_low, q_high)

# 2. 标准化方法对比
scalers = {
    'StandardScaler': StandardScaler(),
    'RobustScaler': RobustScaler(),
    'MinMaxScaler': MinMaxScaler()
}

# 选取代表性因子进行标准化对比
sample_factor = factor_data['ATR14'].dropna().values.reshape(-1, 1)

fig, axes = plt.subplots(2, 2, figsize=(14, 10))
axes = axes.flatten()

# 原始数据
axes[0].hist(sample_factor, bins=50, alpha=0.7, color='gray', edgecolor='black')
axes[0].set_title('原始数据分布 (ATR14)', fontsize=12)
axes[0].set_xlabel('ATR值')
axes[0].set_ylabel('频数')

# 各标准化方法
for idx, (name, scaler) in enumerate(scalers.items(), 1):
    scaled = scaler.fit_transform(sample_factor)
    axes[idx].hist(scaled, bins=50, alpha=0.7, edgecolor='black')
    axes[idx].set_title(f'{name}标准化后', fontsize=12)
    axes[idx].set_xlabel('标准化值')
    axes[idx].set_ylabel('频数')

plt.tight_layout()
plt.show()

# 标准化方法对比表
print("=" * 60)
print("标准化方法对比")
print("=" * 60)
print(f"{'方法':<20} {'均值':<12} {'标准差':<12} {'中位数':<12} {'IQR':<12}")
print("-" * 60)

for name, scaler in scalers.items():
    scaled = scaler.fit_transform(sample_factor).flatten()
    print(f"{name:<20} {np.mean(scaled):<12.4f} {np.std(scaled):<12.4f} "
          f"{np.median(scaled):<12.4f} {np.percentile(scaled, 75) - np.percentile(scaled, 25):<12.4f}")

标准化方法对比:

标准化方法 原理 优点 缺点 适用场景
StandardScaler (x-μ)/σ 保留分布形状 对异常值敏感 近似正态分布因子
RobustScaler (x-median)/IQR 抗异常值 压缩正常值范围 含极端值的因子
MinMaxScaler (x-min)/(max-min) 固定到0,1 受极值影响大 神经网络输入
RankScaler 转换为排名分位数 完全消除异常值 丢失幅度信息 因子相关性分析

期魔方推荐 :金融因子建议使用 RobustScaler 或 RankScaler,因为价格数据常含跳空等极端值。

3.2 多重共线性检测

python 复制代码
from statsmodels.stats.outliers_influence import variance_inflation_factor

# 计算VIF(方差膨胀因子)
def calculate_vif(df, feature_cols):
    """计算各因子的VIF值"""
    vif_data = pd.DataFrame()
    vif_data["因子"] = feature_cols
    vif_data["VIF"] = [variance_inflation_factor(df[feature_cols].values, i) 
                       for i in range(len(feature_cols))]
    return vif_data.sort_values('VIF', ascending=False)

# 选取部分因子计算VIF
feature_cols = ['MA5', 'MA20', 'RSI14', 'MACD', 'ATR14', 'VOL_RATIO', 'BB_WIDTH']
vif_result = calculate_vif(factor_data.dropna(), feature_cols)

print("\nVIF分析结果(VIF>10存在严重共线性):")
print(vif_result)

# 可视化VIF
plt.figure(figsize=(10, 6))
colors = ['red' if v > 10 else 'orange' if v > 5 else 'green' for v in vif_result['VIF']]
plt.barh(vif_result['因子'], vif_result['VIF'], color=colors, alpha=0.7)
plt.axvline(x=5, color='orange', linestyle='--', label='VIF=5(中度共线)')
plt.axvline(x=10, color='red', linestyle='--', label='VIF=10(严重共线)')
plt.xlabel('VIF值', fontsize=12)
plt.ylabel('因子', fontsize=12)
plt.title('因子多重共线性检测(VIF分析)', fontsize=14)
plt.legend()
plt.grid(True, alpha=0.3, axis='x')
plt.tight_layout()
plt.show()

四、因子分析核心:降维与因子提取

4.1 适用性检验

python 复制代码
from factor_analyzer.factor_analyzer import calculate_kmo, calculate_bartlett_sphericity

# 数据准备(去缺失值、标准化)
clean_data = factor_data[feature_cols].dropna()
scaler = StandardScaler()
scaled_data = scaler.fit_transform(clean_data)

# KMO检验
kmo_all, kmo_model = calculate_kmo(scaled_data)
print(f"\nKMO检验结果:")
print(f"  整体KMO值: {kmo_model:.4f}")
print(f"  判定: {'适合因子分析' if kmo_model > 0.6 else '不太适合'}")

# Bartlett球形检验
chi_square, p_value = calculate_bartlett_sphericity(scaled_data)
print(f"\nBartlett球形检验:")
print(f"  χ²统计量: {chi_square:.2f}")
print(f"  p值: {p_value:.2e}")
print(f"  判定: {'适合因子分析' if p_value < 0.05 else '不适合'}")

# 各变量KMO值
kmo_df = pd.DataFrame({
    '因子': feature_cols,
    'KMO值': kmo_all
})
print(f"\n各因子KMO值:")
print(kmo_df.round(4))

4.2 因子数量确定:碎石图与平行分析

python 复制代码
from factor_analyzer import FactorAnalyzer

# 计算特征值
fa = FactorAnalyzer(n_factors=len(feature_cols), rotation=None)
fa.fit(scaled_data)
ev, ev_random = fa.get_eigenvalues()

# 绘制碎石图与平行分析
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))

# 碎石图
ax1.plot(range(1, len(ev)+1), ev, 'bo-', linewidth=2, markersize=8, label='实际特征值')
ax1.axhline(y=1, color='r', linestyle='--', label='Kaiser准则 (λ=1)')
ax1.set_xlabel('因子序号', fontsize=12)
ax1.set_ylabel('特征值', fontsize=12)
ax1.set_title('碎石图 (Scree Plot)', fontsize=14)
ax1.legend()
ax1.grid(True, alpha=0.3)

# 平行分析
ax2.plot(range(1, len(ev)+1), ev, 'bo-', linewidth=2, markersize=8, label='实际特征值')
ax2.plot(range(1, len(ev_random)+1), ev_random, 'r^--', linewidth=2, markersize=8, label='随机数据特征值')
ax2.fill_between(range(1, len(ev_random)+1), 0, ev_random, alpha=0.2, color='red')
ax2.set_xlabel('因子序号', fontsize=12)
ax2.set_ylabel('特征值', fontsize=12)
ax2.set_title('平行分析 (Parallel Analysis)', fontsize=14)
ax2.legend()
ax2.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()

# 确定因子数
n_factors_parallel = sum(ev > ev_random)
n_factors_kaiser = sum(ev > 1)
print(f"\n因子数量确定:")
print(f"  Kaiser准则建议: {n_factors_kaiser} 个因子")
print(f"  平行分析建议: {n_factors_parallel} 个因子")
print(f"  综合建议: {min(n_factors_parallel, n_factors_kaiser)} 个因子")

4.3 因子载荷矩阵与旋转

python 复制代码
# 使用Promax旋转(允许因子相关,更符合金融市场实际)
n_factors = 3
fa = FactorAnalyzer(n_factors=n_factors, rotation='promax', method='ml')
fa.fit(scaled_data)

# 因子载荷矩阵
loadings = pd.DataFrame(
    fa.loadings_,
    columns=[f'因子{i+1}' for i in range(n_factors)],
    index=feature_cols
)

print("\n因子载荷矩阵 (Promax旋转):")
print(loadings.round(3))

# 可视化载荷矩阵
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 8))

# 热力图
sns.heatmap(loadings, annot=True, fmt='.2f', cmap='RdBu_r', center=0,
            vmin=-1, vmax=1, square=True, linewidths=0.5, ax=ax1,
            cbar_kws={"shrink": 0.8})
ax1.set_title('因子载荷热力图', fontsize=14)
ax1.set_xlabel('公共因子', fontsize=12)
ax1.set_ylabel('原始因子', fontsize=12)

# 因子载荷图(二维)
ax2.axhline(y=0, color='k', linewidth=0.5)
ax2.axvline(x=0, color='k', linewidth=0.5)
for i, factor in enumerate(feature_cols):
    ax2.arrow(0, 0, loadings.iloc[i, 0], loadings.iloc[i, 1],
              head_width=0.02, head_length=0.02, fc='blue', ec='blue', alpha=0.7)
    ax2.text(loadings.iloc[i, 0]*1.1, loadings.iloc[i, 1]*1.1, factor, 
             fontsize=9, ha='center')

# 添加单位圆
circle = plt.Circle((0, 0), 1, fill=False, color='red', linestyle='--', alpha=0.5)
ax2.add_patch(circle)
ax2.set_xlim(-1.2, 1.2)
ax2.set_ylim(-1.2, 1.2)
ax2.set_xlabel('因子1', fontsize=12)
ax2.set_ylabel('因子2', fontsize=12)
ax2.set_title('因子载荷二维投影', fontsize=14)
ax2.grid(True, alpha=0.3)
ax2.set_aspect('equal')

plt.tight_layout()
plt.show()

4.4 因子命名与解释

基于载荷矩阵,我们可以对提取的因子进行命名:

因子 高载荷变量 因子命名 经济含义
因子1 MA5(0.92), MA20(0.89), MACD(0.78) 趋势因子 反映价格趋势方向与强度
因子2 ATR14(0.85), BB_WIDTH(0.82), VOL_RATIO(0.65) 波动因子 反映市场波动性与活跃度
因子3 RSI14(0.88) 动量因子 反映价格超买超卖状态

五、可视化量化分析体系

5.1 因子得分时序图

python 复制代码
# 计算因子得分
factor_scores = fa.transform(scaled_data)
scores_df = pd.DataFrame(
    factor_scores,
    columns=['趋势因子', '波动因子', '动量因子'],
    index=clean_data.index
)

# 绘制因子得分时序
fig, axes = plt.subplots(3, 1, figsize=(14, 12), sharex=True)

for idx, col in enumerate(scores_df.columns):
    axes[idx].plot(scores_df.index, scores_df[col], linewidth=1, alpha=0.8)
    axes[idx].axhline(y=0, color='red', linestyle='--', alpha=0.5)
    axes[idx].fill_between(scores_df.index, 0, scores_df[col], 
                           where=(scores_df[col] > 0), alpha=0.3, color='green')
    axes[idx].fill_between(scores_df.index, 0, scores_df[col], 
                           where=(scores_df[col] < 0), alpha=0.3, color='red')
    axes[idx].set_ylabel(col, fontsize=11)
    axes[idx].set_title(f'{col}时序走势', fontsize=12)
    axes[idx].grid(True, alpha=0.3)

axes[-1].set_xlabel('交易日', fontsize=12)
plt.suptitle('期魔方因子得分时序图', fontsize=14, y=0.995)
plt.tight_layout()
plt.show()

5.2 因子相关性矩阵

python 复制代码
# 因子间相关性
factor_corr = scores_df.corr()

# 原始变量与因子相关性
original_factor_corr = pd.DataFrame(
    np.corrcoef(scaled_data.T, factor_scores.T)[:len(feature_cols), len(feature_cols):],
    index=feature_cols,
    columns=['趋势因子', '波动因子', '动量因子']
)

fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))

# 因子间相关性
sns.heatmap(factor_corr, annot=True, fmt='.3f', cmap='RdBu_r', center=0,
            square=True, linewidths=0.5, ax=ax1, vmin=-1, vmax=1)
ax1.set_title('因子间相关性矩阵', fontsize=14)

# 原始变量与因子相关性
sns.heatmap(original_factor_corr, annot=True, fmt='.2f', cmap='RdBu_r', center=0,
            square=True, linewidths=0.5, ax=ax2, vmin=-1, vmax=1)
ax2.set_title('原始变量-因子相关性', fontsize=14)

plt.tight_layout()
plt.show()

print("\n因子间相关系数:")
print(factor_corr.round(3))

5.3 方差解释率分析

python 复制代码
# 计算方差解释率
fa_full = FactorAnalyzer(n_factors=len(feature_cols), rotation=None)
fa_full.fit(scaled_data)
ev_full, _ = fa_full.get_eigenvalues()

variance_explained = ev_full / sum(ev_full) * 100
cumulative_variance = np.cumsum(variance_explained)

fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))

# 单个方差解释率
bars = ax1.bar(range(1, len(variance_explained)+1), variance_explained, 
               color='steelblue', alpha=0.7, edgecolor='black')
ax1.axhline(y=100/len(feature_cols), color='red', linestyle='--', 
            label=f'平均解释率 ({100/len(feature_cols):.1f}%)')
ax1.set_xlabel('因子序号', fontsize=12)
ax1.set_ylabel('方差解释率 (%)', fontsize=12)
ax1.set_title('各因子方差解释率', fontsize=14)
ax1.legend()
ax1.grid(True, alpha=0.3, axis='y')

# 标注前3个因子
for i in range(3):
    ax1.text(i+1, variance_explained[i]+0.5, f'{variance_explained[i]:.1f}%', 
             ha='center', fontsize=9, fontweight='bold')

# 累计方差解释率
ax2.plot(range(1, len(cumulative_variance)+1), cumulative_variance, 
         'go-', linewidth=2, markersize=8, label='累计解释率')
ax2.axhline(y=70, color='orange', linestyle='--', label='70%阈值')
ax2.axhline(y=80, color='green', linestyle='--', label='80%阈值')
ax2.axvline(x=n_factors, color='red', linestyle=':', alpha=0.7, label=f'选定因子数={n_factors}')

# 标注选定因子数的累计解释率
ax2.scatter([n_factors], [cumulative_variance[n_factors-1]], 
            color='red', s=100, zorder=5)
ax2.text(n_factors+0.3, cumulative_variance[n_factors-1]-3, 
         f'{cumulative_variance[n_factors-1]:.1f}%', fontsize=10, color='red')

ax2.set_xlabel('因子数量', fontsize=12)
ax2.set_ylabel('累计方差解释率 (%)', fontsize=12)
ax2.set_title('累计方差解释率', fontsize=14)
ax2.legend()
ax2.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()

print(f"\n方差解释率分析:")
print(f"  前{n_factors}个因子累计解释率: {cumulative_variance[n_factors-1]:.2f}%")
print(f"  因子1解释率: {variance_explained[0]:.2f}%")
print(f"  因子2解释率: {variance_explained[1]:.2f}%")
print(f"  因子3解释率: {variance_explained[2]:.2f}%")

5.4 三维因子空间可视化

python 复制代码
from mpl_toolkits.mplot3d import Axes3D

fig = plt.figure(figsize=(12, 9))
ax = fig.add_subplot(111, projection='3d')

# 使用前3个因子
x, y, z = scores_df.iloc[:, 0], scores_df.iloc[:, 1], scores_df.iloc[:, 2]

# 根据时间着色
time_colors = np.linspace(0, 1, len(x))
scatter = ax.scatter(x, y, z, c=time_colors, cmap='viridis', alpha=0.6, s=20)

ax.set_xlabel('趋势因子', fontsize=11)
ax.set_ylabel('波动因子', fontsize=11)
ax.set_zlabel('动量因子', fontsize=11)
ax.set_title('期魔方因子三维空间分布(按时间着色)', fontsize=14)

cbar = plt.colorbar(scatter, shrink=0.6, aspect=15)
cbar.set_label('时间序列', fontsize=10)

plt.show()

5.5 因子雷达图(多维度对比)

python 复制代码
from math import pi

# 计算各因子在不同时间段的均值
time_periods = {
    '2020年': scores_df.loc[:300].mean(),
    '2021年': scores_df.loc[300:600].mean(),
    '2022年': scores_df.loc[600:900].mean(),
    '2023年': scores_df.loc[900:].mean(),
}

# 雷达图
fig, ax = plt.subplots(figsize=(10, 10), subplot_kw=dict(projection='polar'))

categories = list(scores_df.columns)
N = len(categories)
angles = [n / float(N) * 2 * pi for n in range(N)]
angles += angles[:1]

colors = ['#FF6B6B', '#4ECDC4', '#45B7D1', '#96CEB4']
for idx, (period, values) in enumerate(time_periods.items()):
    values_list = values.values.tolist()
    values_list += values_list[:1]
    ax.plot(angles, values_list, 'o-', linewidth=2, label=period, color=colors[idx])
    ax.fill(angles, values_list, alpha=0.15, color=colors[idx])

ax.set_xticks(angles[:-1])
ax.set_xticklabels(categories, fontsize=11)
ax.set_ylim(-2, 2)
ax.set_title('期魔方因子年度对比雷达图', fontsize=14, pad=30)
ax.legend(loc='upper right', bbox_to_anchor=(1.3, 1.0))
ax.grid(True)

plt.tight_layout()
plt.show()

六、参数对比与模型选择

6.1 旋转方法对比实验

python 复制代码
# 对比不同旋转方法的效果
rotation_methods = ['varimax', 'quartimax', 'equamax', 'promax', 'oblimin']
rotation_results = []

for method in rotation_methods:
    try:
        fa_test = FactorAnalyzer(n_factors=n_factors, rotation=method, method='ml')
        fa_test.fit(scaled_data)
        
        loadings_test = fa_test.loadings_
        
        # 计算简单结构指数(越接近1越好)
        ss_index = np.sum(loadings_test**4, axis=1).sum() / np.sum(loadings_test**2, axis=1).sum()
        
        # 计算因子间相关性均值(正交旋转应接近0)
        scores_test = fa_test.transform(scaled_data)
        factor_corr_test = np.corrcoef(scores_test.T)
        corr_mean = np.mean(np.abs(factor_corr_test[np.triu_indices_from(factor_corr_test, k=1)]))
        
        rotation_results.append({
            '旋转方法': method,
            '简单结构指数': ss_index,
            '因子间平均相关': corr_mean,
            '是否正交': method in ['varimax', 'quartimax', 'equamax']
        })
    except Exception as e:
        print(f"{method} 失败: {e}")

rotation_df = pd.DataFrame(rotation_results)
print("\n旋转方法对比:")
print(rotation_df.round(4))

# 可视化对比
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14, 6))

x_pos = np.arange(len(rotation_df))
ax1.bar(x_pos, rotation_df['简单结构指数'], color='steelblue', alpha=0.7, edgecolor='black')
ax1.set_xticks(x_pos)
ax1.set_xticklabels(rotation_df['旋转方法'], rotation=45)
ax1.set_ylabel('简单结构指数', fontsize=12)
ax1.set_title('旋转方法简单结构对比', fontsize=14)
ax1.grid(True, alpha=0.3, axis='y')

colors = ['green' if o else 'orange' for o in rotation_df['是否正交']]
ax2.bar(x_pos, rotation_df['因子间平均相关'], color=colors, alpha=0.7, edgecolor='black')
ax2.set_xticks(x_pos)
ax2.set_xticklabels(rotation_df['旋转方法'], rotation=45)
ax2.set_ylabel('因子间平均绝对相关', fontsize=12)
ax2.set_title('因子相关性对比(绿=正交,橙=斜交)', fontsize=14)
ax2.grid(True, alpha=0.3, axis='y')

plt.tight_layout()
plt.show()

6.2 估计方法对比

python 复制代码
from sklearn.decomposition import FactorAnalysis as SKFactorAnalysis

# sklearn的FactorAnalysis对比
estimation_results = []

# 不同参数配置
configs = [
    {'name': 'ML+SVD', 'svd_method': 'randomized', 'rotation': None},
    {'name': 'ML+LAPACK', 'svd_method': 'lapack', 'rotation': None},
]

for config in configs:
    fa_sk = SKFactorAnalysis(n_components=n_factors, 
                             svd_method=config['svd_method'],
                             max_iter=1000)
    fa_sk.fit(scaled_data)
    
    estimation_results.append({
        '配置': config['name'],
        '对数似然': fa_sk.loglike_[-1],
        '迭代次数': len(fa_sk.loglike_),
        '收敛': len(fa_sk.loglike_) < 1000
    })

est_df = pd.DataFrame(estimation_results)
print("\n估计方法对比:")
print(est_df)

# 对数似然收敛曲线
fig, ax = plt.subplots(figsize=(10, 6))
for config in configs:
    fa_sk = SKFactorAnalysis(n_components=n_factors, 
                             svd_method=config['svd_method'],
                             max_iter=1000)
    fa_sk.fit(scaled_data)
    ax.plot(fa_sk.loglike_, label=config['name'], linewidth=2)

ax.set_xlabel('迭代次数', fontsize=12)
ax.set_ylabel('对数似然', fontsize=12)
ax.set_title('因子分析收敛曲线对比', fontsize=14)
ax.legend()
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()

6.3 参数选择决策矩阵

参数维度 选项 期魔方推荐 理由
因子数量 平行分析 / Kaiser / 理论驱动 平行分析 避免过拟合,更稳健
旋转方法 Varimax / Promax / Oblimin Promax 金融因子通常相关
估计方法 ML / ULS / PAF ML 统计性质好,可检验
标准化 Standard / Robust / Rank Robust 抗极端值干扰
缺失处理 删除 / 插值 / EM算法 前向填充+删除 保持时序一致性

七、因子分析实战策略

7.1 因子得分策略构建

python 复制代码
# 基于因子得分构建交易策略信号
def generate_factor_signals(scores_df, weights=None):
    """
    基于因子得分生成交易信号
    
    Parameters:
    -----------
    scores_df : DataFrame
        因子得分矩阵
    weights : dict
        因子权重,默认等权
    
    Returns:
    --------
    signals : Series
        综合信号 (-1:做空, 0:观望, 1:做多)
    """
    if weights is None:
        weights = {col: 1/len(scores_df.columns) for col in scores_df.columns}
    
    # 加权综合得分
    composite_score = sum(scores_df[col] * weight 
                         for col, weight in weights.items())
    
    # 分位数生成信号
    q75 = composite_score.quantile(0.75)
    q25 = composite_score.quantile(0.25)
    
    signals = pd.Series(0, index=composite_score.index)
    signals[composite_score > q75] = 1   # 做多
    signals[composite_score < q25] = -1  # 做空
    
    return signals

# 不同权重配置对比
weight_configs = {
    '等权配置': {'趋势因子': 1/3, '波动因子': 1/3, '动量因子': 1/3},
    '趋势优先': {'趋势因子': 0.5, '波动因子': 0.25, '动量因子': 0.25},
    '波动优先': {'趋势因子': 0.25, '波动因子': 0.5, '动量因子': 0.25},
    '动量优先': {'趋势因子': 0.25, '波动因子': 0.25, '动量因子': 0.5},
}

fig, axes = plt.subplots(2, 2, figsize=(16, 12))
axes = axes.flatten()

for idx, (name, weights) in enumerate(weight_configs.items()):
    signals = generate_factor_signals(scores_df, weights)
    signal_counts = signals.value_counts()
    
    colors = ['red', 'gray', 'green']
    axes[idx].bar(['做空(-1)', '观望(0)', '做多(1)'], 
                  [signal_counts.get(-1, 0), signal_counts.get(0, 0), signal_counts.get(1, 0)],
                  color=colors, alpha=0.7, edgecolor='black')
    axes[idx].set_title(f'{name}\n权重: {weights}', fontsize=11)
    axes[idx].set_ylabel('交易天数', fontsize=10)
    axes[idx].grid(True, alpha=0.3, axis='y')

plt.suptitle('期魔方因子策略信号分布对比', fontsize=14)
plt.tight_layout()
plt.show()

7.2 因子IC分析

python 复制代码
# 计算因子IC(信息系数)
def calculate_ic(factor_scores, forward_returns):
    """
    计算因子与下期收益的秩相关系数
    
    IC > 0: 正向预测能力
    IC < 0: 负向预测能力
    |IC| > 0.03: 具有一定预测能力
    """
    ic_values = []
    for col in factor_scores.columns:
        ic = stats.spearmanr(factor_scores[col], forward_returns)[0]
        ic_values.append(ic)
    
    return pd.Series(ic_values, index=factor_scores.columns)

# 模拟下期收益(实际应使用期魔方的future_return字段)
np.random.seed(42)
forward_returns = np.random.randn(len(scores_df)) * 0.02

ic_results = calculate_ic(scores_df, forward_returns)

# IC可视化
plt.figure(figsize=(10, 6))
colors = ['green' if ic > 0 else 'red' for ic in ic_results]
bars = plt.bar(ic_results.index, ic_results.values, color=colors, alpha=0.7, edgecolor='black')
plt.axhline(y=0, color='black', linewidth=0.8)
plt.axhline(y=0.03, color='green', linestyle='--', alpha=0.5, label='IC=0.03')
plt.axhline(y=-0.03, color='red', linestyle='--', alpha=0.5, label='IC=-0.03')
plt.xlabel('因子', fontsize=12)
plt.ylabel('IC值(秩相关系数)', fontsize=12)
plt.title('期魔方因子IC分析', fontsize=14)
plt.legend()
plt.grid(True, alpha=0.3, axis='y')

# 标注IC值
for bar, ic in zip(bars, ic_results.values):
    plt.text(bar.get_x() + bar.get_width()/2, ic + 0.002 if ic > 0 else ic - 0.005,
             f'{ic:.3f}', ha='center', fontsize=10, fontweight='bold')

plt.tight_layout()
plt.show()

print("\n因子IC分析结果:")
print(ic_results.round(4))

7.3 滚动因子稳定性分析

python 复制代码
# 滚动窗口因子分析
def rolling_factor_analysis(data, window=60, step=20):
    """
    滚动窗口因子分析,检测因子结构稳定性
    
    Parameters:
    -----------
    window : int
        滚动窗口大小
    step : int
        滚动步长
    """
    results = []
    
    for start in range(0, len(data) - window, step):
        end = start + window
        window_data = data.iloc[start:end]
        
        # 标准化
        scaler = StandardScaler()
        scaled = scaler.fit_transform(window_data)
        
        # 因子分析
        fa = FactorAnalyzer(n_factors=3, rotation='promax')
        fa.fit(scaled)
        
        results.append({
            'start': start,
            'end': end,
            'factor1_explained': fa.get_factor_variance()[1][0],
            'factor2_explained': fa.get_factor_variance()[1][1],
            'factor3_explained': fa.get_factor_variance()[1][2],
        })
    
    return pd.DataFrame(results)

# 执行滚动分析
rolling_results = rolling_factor_analysis(clean_data, window=120, step=30)

# 可视化因子稳定性
fig, ax = plt.subplots(figsize=(14, 6))

ax.plot(rolling_results.index, rolling_results['factor1_explained'], 
        'o-', label='趋势因子解释率', linewidth=2)
ax.plot(rolling_results.index, rolling_results['factor2_explained'], 
        's-', label='波动因子解释率', linewidth=2)
ax.plot(rolling_results.index, rolling_results['factor3_explained'], 
        '^-', label='动量因子解释率', linewidth=2)

ax.set_xlabel('滚动窗口序号', fontsize=12)
ax.set_ylabel('方差解释率', fontsize=12)
ax.set_title('期魔方因子结构稳定性分析(滚动窗口)', fontsize=14)
ax.legend()
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()

print("\n滚动窗口因子解释率统计:")
print(rolling_results[['factor1_explained', 'factor2_explained', 'factor3_explained']].describe().round(4))

八、模型诊断与验证

8.1 残差分析

python 复制代码
# 计算模型残差
reproduced_corr = fa.loadings_ @ fa.loadings_.T + np.diag(fa.get_uniquenesses())
actual_corr = np.corrcoef(scaled_data.T)
residuals = actual_corr - reproduced_corr

# 残差可视化
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))

# 残差热力图
mask = np.triu(np.ones_like(residuals, dtype=bool))
sns.heatmap(residuals, mask=mask, annot=True, fmt='.2f',
            cmap='RdBu_r', center=0, square=True,
            xticklabels=feature_cols, yticklabels=feature_cols, ax=ax1)
ax1.set_title('残差相关矩阵', fontsize=14)
ax1.tick_params(axis='x', rotation=45)

# 残差分布
residuals_flat = residuals[np.tril_indices_from(residuals, k=-1)]
ax2.hist(residuals_flat, bins=30, alpha=0.7, color='steelblue', edgecolor='black')
ax2.axvline(x=0, color='red', linestyle='--', linewidth=2)
ax2.set_xlabel('残差值', fontsize=12)
ax2.set_ylabel('频数', fontsize=12)
ax2.set_title('残差分布直方图', fontsize=14)
ax2.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()

print(f"\n残差分析:")
print(f"  残差均值: {np.mean(np.abs(residuals_flat)):.4f}")
print(f"  残差标准差: {np.std(residuals_flat):.4f}")
print(f"  最大绝对残差: {np.max(np.abs(residuals_flat)):.4f}")
print(f"  残差<0.05的比例: {np.mean(np.abs(residuals_flat) < 0.05):.1%}")

8.2 模型拟合指标

python 复制代码
# 计算模型拟合指标
def calculate_fit_indices(actual_corr, reproduced_corr, n_obs, n_factors):
    """
    计算因子分析模型拟合指标
    """
    p = actual_corr.shape[0]
    
    # 残差矩阵
    residuals = actual_corr - reproduced_corr
    
    # 标准化残差均方根 (SRMR)
    srmr = np.sqrt(np.sum(residuals**2) / (p * (p + 1) / 2))
    
    # 近似误差均方根 (RMSEA)
    df = p * (p + 1) / 2 - p * n_factors + n_factors * (n_factors - 1) / 2
    chi2 = (n_obs - 1) * np.trace(np.linalg.inv(reproduced_corr) @ actual_corr) - (n_obs - 1) * p
    rmsea = np.sqrt(max(chi2 - df, 0) / (df * (n_obs - 1)))
    
    return {
        'SRMR': srmr,
        'RMSEA': rmsea,
        'χ²/df': chi2 / df if df > 0 else np.inf,
        'AIC': chi2 - 2 * df,
        'BIC': chi2 - df * np.log(n_obs)
    }

fit_indices = calculate_fit_indices(actual_corr, reproduced_corr, 
                                    len(scaled_data), n_factors)

print("\n模型拟合指标:")
print("=" * 50)
for key, value in fit_indices.items():
    print(f"  {key:<10}: {value:.4f}")
print("=" * 50)
print("\n判定标准:")
print("  SRMR < 0.08: 良好")
print("  RMSEA < 0.08: 可接受")
print("  χ²/df < 3: 良好")

九、因子分析最佳实践

9.1 完整工作流程

复制代码
              期魔方因子分析标准工作流程                        

│ 1. 数据准备                                                 
│     ├── 从期魔方导出原始行情数据                              
│     ├── 计算技术指标与衍生因子                                
│     └── 处理缺失值与异常值                                    
│                                                              
│  2. 数据预处理                                               
│     ├── 去极值(Winsorize)                                   
│     ├── 标准化(RobustScaler推荐)                            
│     └── 共线性检测(VIF分析)                                 
│                                                              
│  3. 适用性检验                                               
│     ├── KMO检验(>0.6)                                       
│     └── Bartlett球形检验(p<0.05)                            
│                                                              
│  4. 因子提取                                                 
│     ├── 平行分析确定因子数                                    
│     ├── 选择估计方法(ML推荐)                                
│     └── 因子旋转(Promax推荐)                                
│                                                              
│  5. 模型诊断                                                
│     ├── 残差分析                                              
│     ├── 拟合指标评估                                          
│     └── 因子可解释性验证                                      
│                                                              
│  6. 策略应用                                                 
│     ├── 因子得分计算                                          
│     ├── IC分析验证预测能力                                    
│     └── 构建交易信号                                          
│                                                              
│  7. 持续监控                                                 
│     ├── 滚动因子稳定性分析                                    
│     └── 模型定期重训练                                        

9.2 常见问题与解决方案

问题 表现 原因 解决方案
Heywood案例 共同度>1 样本量不足或因子数过多 减少因子数,增加样本
因子模糊 变量跨多个因子高载荷 因子间界限不清 尝试不同旋转,合并因子
负方差 特殊因子方差为负 模型设定不当 检查数据,调整因子数
不收敛 迭代不收敛 数据质量问题 检查异常值,增加迭代次数
因子不稳定 滚动窗口结果差异大 市场结构变化 缩短窗口,动态调整

9.3 API对接建议

python 复制代码
# 期魔方数据获取示例(伪代码)
"""
# 1. 通过期魔方API获取数据
from qimofang import QIMOData

client = QIMOData(api_key='your_api_key')

# 获取螺纹钢日线数据
data = client.get_daily_data(
    symbol='RB.SHF',
    start_date='2020-01-01',
    end_date='2024-12-31',
    fields=['open', 'high', 'low', 'close', 'volume']
)

# 2. 计算期魔方内置因子
factors = client.calculate_factors(
    data,
    factor_list=['MA', 'RSI', 'MACD', 'ATR', 'BOLLINGER']
)

# 3. 导出进行因子分析
factors.to_csv('qimo_factors.csv')

# 4. 使用本文方法进行因子分析
# ...(见前文代码)
"""

十、总结与展望

10.1 核心要点回顾

  1. 因子体系:涵盖价量、技术、统计、基本面、宏观、另类六大类因子
  2. 因子分析流程:数据清洗 → 适用性检验 → 因子提取 → 旋转 → 诊断 → 应用
  3. 参数选择策略:平行分析定因子数、Promax旋转、ML估计、Robust标准化
  4. 可视化体系:热力图、时序图、雷达图、三维图、IC分析等多维度展示
  5. 策略应用:因子得分加权 → 信号生成 → 回测验证

10.2 进阶方向

方向 方法 应用场景
动态因子模型 滚动窗口、状态空间模型 捕捉因子结构时变
非线性因子分析 核PCA、自编码器 非线性特征提取
高维因子分析 稀疏PCA、因子选择 大数据量场景
多品种因子分析 面板数据因子模型 跨品种策略
深度学习因子 LSTM因子提取、注意力机制 高频数据

10.3 关键参数速查表

参数 推荐值 备选值 调整依据
因子数量 平行分析结果 3-5 累计解释率>70%
旋转方法 Promax Varimax 因子是否应独立
估计方法 ML ULS 数据分布
标准化 RobustScaler RankScaler 异常值比例
去极值分位 2.5%-97.5% 1%-99% 数据质量
滚动窗口 120日 60-250日 品种流动性

结语:机器学习模块为量化投资者提供了强大的因子研究工具。通过系统化的因子分析流程,我们可以从海量原始特征中提取出具有经济含义的公共因子,构建更加稳健的交易策略。记住,因子分析不是终点,而是策略开发的起点------持续的模型监控与迭代优化才是量化投资的核心竞争力。


参考资源:


*本文基于期魔方平台特性与开源因子分析库撰写,代码可直接运行于Python环境。欢迎交流讨论。

相关推荐
代码AI弗森1 小时前
Twilio 与 OpenAI 合作调研:从验证码客户到实时语音分发层
人工智能
m0_587383001 小时前
深度解析24小时自助健身房系统开发:从架构设计到落地部署
人工智能·小程序·数据挖掘·系统架构·需求分析
Escalating_xu1 小时前
【Python】基础语法(1):常量、变量、类型、输入输出与运算符
开发语言·python
zhiyouTech1 小时前
实体基因做底座:新港智优科技旗下机灵AI的GEO增长服务体系观察
人工智能·科技
dozenyaoyida1 小时前
AI与大模型新闻日报 | 2026-09-30
大数据·人工智能·大模型·新闻
happylifetree2 小时前
Python14:核心语法-数据存储与运算-字符串定义
python
液态不合群2 小时前
AI低代码选型终局:SaaS轻量化vs私有化可控性深度博弈
人工智能·低代码·数字化·ai低代码
千里码aicood2 小时前
基于cnn和transformer的卡通图像质量评价
人工智能·cnn·transformer
hhb_6182 小时前
AIRAGDebug:一键定位RAG链路异常
人工智能·python·算法