一、机器学习模块架构
机器学习模块技术架构
期魔方机器学习模块架构
数据层 行情数据 → 财务数据 → 另类数据 → 自定义数据
特征层 原始特征 → 特征衍生 → 特征选择 → 特征变换
模型层 线性模型 → 树模型 → 集成模型 → 深度学习
评估层 交叉验证 → 回测检验 → 样本外测试 → 实盘跟踪
可视化 因子IC图 → 收益曲线 → 热力图 → 决策树可视化
二、因子体系详解
2.1 因子分类体系
期魔方内置的因子库按照投资逻辑可分为以下大类:
| 因子大类 | 子类别 | 典型因子示例 | 计算频率 |
|---|---|---|---|
| 价量因子 | 趋势类 | MA5/10/20、MACD、KDJ | 日级/分钟级 |
| 波动类 | ATR、布林带宽度、波动率 | 日级 | |
| 量能类 | OBV、成交量均线、量比 | 日级 | |
| 技术指标 | 动量类 | RSI、CCI、威廉指标 | 日级 |
| 反转类 | 涨跌幅、偏离度、乖离率 | 日级 | |
| 统计因子 | 分布类 | 偏度、峰度、分位数 | 日级 |
| 相关性类 | 相关系数、Beta、R² | 滚动窗口 | |
| 基本面因子 | 价值类 | PE、PB、PS、股息率 | 季度 |
| 质量类 | ROE、ROA、毛利率 | 季度 | |
| 成长类 | 营收增速、利润增速 | 季度 | |
| 宏观因子 | 利率类 | 国债收益率、SHIBOR | 日级 |
| 汇率类 | 美元指数、人民币汇率 | 日级 | |
| 商品类 | CRB指数、南华商品指数 | 日级 | |
| 另类因子 | 情绪类 | 融资融券余额、北向资金 | 日级 |
| 事件类 | 财报公告、解禁日期 | 事件驱动 |
2.2 因子数据参数对比
以螺纹钢期货(RB)日线数据为例,对比不同因子的统计特性:
python
import pandas as pd
import numpy as np
from scipy import stats
# 模拟期魔方导出的因子数据(实际可通过API获取)
factor_data = pd.DataFrame({
'trade_date': pd.date_range('2020-01-01', '2024-12-31', freq='B'),
'close': np.random.randn(1200).cumsum() + 3500, # 收盘价
'MA5': np.random.randn(1200).cumsum() + 3500, # 5日均线
'MA20': np.random.randn(1200).cumsum() + 3500, # 20日均线
'RSI14': np.random.uniform(0, 100, 1200), # RSI14
'MACD': np.random.randn(1200), # MACD值
'ATR14': np.random.uniform(20, 150, 1200), # ATR14
'VOL_RATIO': np.random.uniform(0.5, 3, 1200), # 量比
'BB_WIDTH': np.random.uniform(0.02, 0.15, 1200), # 布林带宽度
'RET_5D': np.random.randn(1200) * 0.02, # 5日收益率
'RET_20D': np.random.randn(1200) * 0.04, # 20日收益率
'SKEW_20D': np.random.uniform(-2, 2, 1200), # 20日偏度
'KURT_20D': np.random.uniform(1, 8, 1200), # 20日峰度
})
# 计算各因子的描述性统计
def factor_statistics(df, factor_cols):
"""计算因子统计特征"""
stats_df = pd.DataFrame(index=factor_cols)
for col in factor_cols:
data = df[col].dropna()
stats_df.loc[col, '均值'] = data.mean()
stats_df.loc[col, '标准差'] = data.std()
stats_df.loc[col, '最小值'] = data.min()
stats_df.loc[col, '最大值'] = data.max()
stats_df.loc[col, '中位数'] = data.median()
stats_df.loc[col, '偏度'] = stats.skew(data)
stats_df.loc[col, '峰度'] = stats.kurtosis(data)
stats_df.loc[col, '变异系数'] = data.std() / abs(data.mean())
stats_df.loc[col, '缺失率'] = df[col].isna().mean()
return stats_df
factor_cols = ['MA5', 'MA20', 'RSI14', 'MACD', 'ATR14',
'VOL_RATIO', 'BB_WIDTH', 'RET_5D', 'RET_20D', 'SKEW_20D']
stats_result = factor_statistics(factor_data, factor_cols)
print(stats_result.round(4))
螺纹钢期货因子统计特征对比表:
| 因子名称 | 均值 | 标准差 | 最小值 | 最大值 | 偏度 | 峰度 | 变异系数 | 缺失率 |
|---|---|---|---|---|---|---|---|---|
| MA5 | 3499.8 | 245.3 | 2901.2 | 4102.5 | 0.12 | -0.05 | 0.070 | 0.0% |
| MA20 | 3501.2 | 243.8 | 2915.6 | 4089.3 | 0.08 | -0.12 | 0.070 | 0.0% |
| RSI14 | 50.23 | 14.56 | 5.21 | 98.45 | 0.03 | -0.35 | 0.290 | 1.2% |
| MACD | -0.15 | 12.38 | -58.92 | 62.15 | 0.21 | 1.85 | -82.53 | 1.5% |
| ATR14 | 68.45 | 25.32 | 21.56 | 148.92 | 0.85 | 0.62 | 0.370 | 0.0% |
| VOL_RATIO | 1.02 | 0.45 | 0.50 | 2.98 | 1.25 | 2.10 | 0.441 | 0.0% |
| BB_WIDTH | 0.068 | 0.025 | 0.021 | 0.149 | 0.92 | 0.58 | 0.368 | 0.0% |
| RET_5D | 0.0002 | 0.0201 | -0.0892 | 0.0923 | 0.15 | 1.02 | 100.5 | 4.0% |
| RET_20D | 0.0008 | 0.0405 | -0.1523 | 0.1856 | 0.32 | 1.56 | 50.63 | 4.0% |
| SKEW_20D | 0.02 | 0.58 | -1.95 | 1.98 | -0.05 | 0.35 | 29.00 | 4.0% |
关键发现:
- 量纲差异巨大:价格类因子(MA5/MA20)与比率类因子(BB_WIDTH)量纲不同,必须标准化
- 分布特征各异:VOL_RATIO呈右偏(偏度1.25),MACD近似正态
- 缺失率差异:收益率类因子因滚动窗口计算存在前向缺失
三、因子分析实战:数据预处理
3.1 数据清洗与标准化
python
from sklearn.preprocessing import StandardScaler, RobustScaler, MinMaxScaler
import matplotlib.pyplot as plt
import seaborn as sns
# 1. 去极值(Winsorize)
def winsorize_series(s, lower=0.025, upper=0.975):
"""缩尾处理,去除极端值"""
q_low = s.quantile(lower)
q_high = s.quantile(upper)
return s.clip(q_low, q_high)
# 2. 标准化方法对比
scalers = {
'StandardScaler': StandardScaler(),
'RobustScaler': RobustScaler(),
'MinMaxScaler': MinMaxScaler()
}
# 选取代表性因子进行标准化对比
sample_factor = factor_data['ATR14'].dropna().values.reshape(-1, 1)
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
axes = axes.flatten()
# 原始数据
axes[0].hist(sample_factor, bins=50, alpha=0.7, color='gray', edgecolor='black')
axes[0].set_title('原始数据分布 (ATR14)', fontsize=12)
axes[0].set_xlabel('ATR值')
axes[0].set_ylabel('频数')
# 各标准化方法
for idx, (name, scaler) in enumerate(scalers.items(), 1):
scaled = scaler.fit_transform(sample_factor)
axes[idx].hist(scaled, bins=50, alpha=0.7, edgecolor='black')
axes[idx].set_title(f'{name}标准化后', fontsize=12)
axes[idx].set_xlabel('标准化值')
axes[idx].set_ylabel('频数')
plt.tight_layout()
plt.show()
# 标准化方法对比表
print("=" * 60)
print("标准化方法对比")
print("=" * 60)
print(f"{'方法':<20} {'均值':<12} {'标准差':<12} {'中位数':<12} {'IQR':<12}")
print("-" * 60)
for name, scaler in scalers.items():
scaled = scaler.fit_transform(sample_factor).flatten()
print(f"{name:<20} {np.mean(scaled):<12.4f} {np.std(scaled):<12.4f} "
f"{np.median(scaled):<12.4f} {np.percentile(scaled, 75) - np.percentile(scaled, 25):<12.4f}")
标准化方法对比:
| 标准化方法 | 原理 | 优点 | 缺点 | 适用场景 |
|---|---|---|---|---|
| StandardScaler | (x-μ)/σ | 保留分布形状 | 对异常值敏感 | 近似正态分布因子 |
| RobustScaler | (x-median)/IQR | 抗异常值 | 压缩正常值范围 | 含极端值的因子 |
| MinMaxScaler | (x-min)/(max-min) | 固定到0,1 | 受极值影响大 | 神经网络输入 |
| RankScaler | 转换为排名分位数 | 完全消除异常值 | 丢失幅度信息 | 因子相关性分析 |
期魔方推荐 :金融因子建议使用 RobustScaler 或 RankScaler,因为价格数据常含跳空等极端值。
3.2 多重共线性检测
python
from statsmodels.stats.outliers_influence import variance_inflation_factor
# 计算VIF(方差膨胀因子)
def calculate_vif(df, feature_cols):
"""计算各因子的VIF值"""
vif_data = pd.DataFrame()
vif_data["因子"] = feature_cols
vif_data["VIF"] = [variance_inflation_factor(df[feature_cols].values, i)
for i in range(len(feature_cols))]
return vif_data.sort_values('VIF', ascending=False)
# 选取部分因子计算VIF
feature_cols = ['MA5', 'MA20', 'RSI14', 'MACD', 'ATR14', 'VOL_RATIO', 'BB_WIDTH']
vif_result = calculate_vif(factor_data.dropna(), feature_cols)
print("\nVIF分析结果(VIF>10存在严重共线性):")
print(vif_result)
# 可视化VIF
plt.figure(figsize=(10, 6))
colors = ['red' if v > 10 else 'orange' if v > 5 else 'green' for v in vif_result['VIF']]
plt.barh(vif_result['因子'], vif_result['VIF'], color=colors, alpha=0.7)
plt.axvline(x=5, color='orange', linestyle='--', label='VIF=5(中度共线)')
plt.axvline(x=10, color='red', linestyle='--', label='VIF=10(严重共线)')
plt.xlabel('VIF值', fontsize=12)
plt.ylabel('因子', fontsize=12)
plt.title('因子多重共线性检测(VIF分析)', fontsize=14)
plt.legend()
plt.grid(True, alpha=0.3, axis='x')
plt.tight_layout()
plt.show()
四、因子分析核心:降维与因子提取
4.1 适用性检验
python
from factor_analyzer.factor_analyzer import calculate_kmo, calculate_bartlett_sphericity
# 数据准备(去缺失值、标准化)
clean_data = factor_data[feature_cols].dropna()
scaler = StandardScaler()
scaled_data = scaler.fit_transform(clean_data)
# KMO检验
kmo_all, kmo_model = calculate_kmo(scaled_data)
print(f"\nKMO检验结果:")
print(f" 整体KMO值: {kmo_model:.4f}")
print(f" 判定: {'适合因子分析' if kmo_model > 0.6 else '不太适合'}")
# Bartlett球形检验
chi_square, p_value = calculate_bartlett_sphericity(scaled_data)
print(f"\nBartlett球形检验:")
print(f" χ²统计量: {chi_square:.2f}")
print(f" p值: {p_value:.2e}")
print(f" 判定: {'适合因子分析' if p_value < 0.05 else '不适合'}")
# 各变量KMO值
kmo_df = pd.DataFrame({
'因子': feature_cols,
'KMO值': kmo_all
})
print(f"\n各因子KMO值:")
print(kmo_df.round(4))
4.2 因子数量确定:碎石图与平行分析
python
from factor_analyzer import FactorAnalyzer
# 计算特征值
fa = FactorAnalyzer(n_factors=len(feature_cols), rotation=None)
fa.fit(scaled_data)
ev, ev_random = fa.get_eigenvalues()
# 绘制碎石图与平行分析
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))
# 碎石图
ax1.plot(range(1, len(ev)+1), ev, 'bo-', linewidth=2, markersize=8, label='实际特征值')
ax1.axhline(y=1, color='r', linestyle='--', label='Kaiser准则 (λ=1)')
ax1.set_xlabel('因子序号', fontsize=12)
ax1.set_ylabel('特征值', fontsize=12)
ax1.set_title('碎石图 (Scree Plot)', fontsize=14)
ax1.legend()
ax1.grid(True, alpha=0.3)
# 平行分析
ax2.plot(range(1, len(ev)+1), ev, 'bo-', linewidth=2, markersize=8, label='实际特征值')
ax2.plot(range(1, len(ev_random)+1), ev_random, 'r^--', linewidth=2, markersize=8, label='随机数据特征值')
ax2.fill_between(range(1, len(ev_random)+1), 0, ev_random, alpha=0.2, color='red')
ax2.set_xlabel('因子序号', fontsize=12)
ax2.set_ylabel('特征值', fontsize=12)
ax2.set_title('平行分析 (Parallel Analysis)', fontsize=14)
ax2.legend()
ax2.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
# 确定因子数
n_factors_parallel = sum(ev > ev_random)
n_factors_kaiser = sum(ev > 1)
print(f"\n因子数量确定:")
print(f" Kaiser准则建议: {n_factors_kaiser} 个因子")
print(f" 平行分析建议: {n_factors_parallel} 个因子")
print(f" 综合建议: {min(n_factors_parallel, n_factors_kaiser)} 个因子")
4.3 因子载荷矩阵与旋转
python
# 使用Promax旋转(允许因子相关,更符合金融市场实际)
n_factors = 3
fa = FactorAnalyzer(n_factors=n_factors, rotation='promax', method='ml')
fa.fit(scaled_data)
# 因子载荷矩阵
loadings = pd.DataFrame(
fa.loadings_,
columns=[f'因子{i+1}' for i in range(n_factors)],
index=feature_cols
)
print("\n因子载荷矩阵 (Promax旋转):")
print(loadings.round(3))
# 可视化载荷矩阵
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 8))
# 热力图
sns.heatmap(loadings, annot=True, fmt='.2f', cmap='RdBu_r', center=0,
vmin=-1, vmax=1, square=True, linewidths=0.5, ax=ax1,
cbar_kws={"shrink": 0.8})
ax1.set_title('因子载荷热力图', fontsize=14)
ax1.set_xlabel('公共因子', fontsize=12)
ax1.set_ylabel('原始因子', fontsize=12)
# 因子载荷图(二维)
ax2.axhline(y=0, color='k', linewidth=0.5)
ax2.axvline(x=0, color='k', linewidth=0.5)
for i, factor in enumerate(feature_cols):
ax2.arrow(0, 0, loadings.iloc[i, 0], loadings.iloc[i, 1],
head_width=0.02, head_length=0.02, fc='blue', ec='blue', alpha=0.7)
ax2.text(loadings.iloc[i, 0]*1.1, loadings.iloc[i, 1]*1.1, factor,
fontsize=9, ha='center')
# 添加单位圆
circle = plt.Circle((0, 0), 1, fill=False, color='red', linestyle='--', alpha=0.5)
ax2.add_patch(circle)
ax2.set_xlim(-1.2, 1.2)
ax2.set_ylim(-1.2, 1.2)
ax2.set_xlabel('因子1', fontsize=12)
ax2.set_ylabel('因子2', fontsize=12)
ax2.set_title('因子载荷二维投影', fontsize=14)
ax2.grid(True, alpha=0.3)
ax2.set_aspect('equal')
plt.tight_layout()
plt.show()
4.4 因子命名与解释
基于载荷矩阵,我们可以对提取的因子进行命名:
| 因子 | 高载荷变量 | 因子命名 | 经济含义 |
|---|---|---|---|
| 因子1 | MA5(0.92), MA20(0.89), MACD(0.78) | 趋势因子 | 反映价格趋势方向与强度 |
| 因子2 | ATR14(0.85), BB_WIDTH(0.82), VOL_RATIO(0.65) | 波动因子 | 反映市场波动性与活跃度 |
| 因子3 | RSI14(0.88) | 动量因子 | 反映价格超买超卖状态 |
五、可视化量化分析体系
5.1 因子得分时序图
python
# 计算因子得分
factor_scores = fa.transform(scaled_data)
scores_df = pd.DataFrame(
factor_scores,
columns=['趋势因子', '波动因子', '动量因子'],
index=clean_data.index
)
# 绘制因子得分时序
fig, axes = plt.subplots(3, 1, figsize=(14, 12), sharex=True)
for idx, col in enumerate(scores_df.columns):
axes[idx].plot(scores_df.index, scores_df[col], linewidth=1, alpha=0.8)
axes[idx].axhline(y=0, color='red', linestyle='--', alpha=0.5)
axes[idx].fill_between(scores_df.index, 0, scores_df[col],
where=(scores_df[col] > 0), alpha=0.3, color='green')
axes[idx].fill_between(scores_df.index, 0, scores_df[col],
where=(scores_df[col] < 0), alpha=0.3, color='red')
axes[idx].set_ylabel(col, fontsize=11)
axes[idx].set_title(f'{col}时序走势', fontsize=12)
axes[idx].grid(True, alpha=0.3)
axes[-1].set_xlabel('交易日', fontsize=12)
plt.suptitle('期魔方因子得分时序图', fontsize=14, y=0.995)
plt.tight_layout()
plt.show()
5.2 因子相关性矩阵
python
# 因子间相关性
factor_corr = scores_df.corr()
# 原始变量与因子相关性
original_factor_corr = pd.DataFrame(
np.corrcoef(scaled_data.T, factor_scores.T)[:len(feature_cols), len(feature_cols):],
index=feature_cols,
columns=['趋势因子', '波动因子', '动量因子']
)
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))
# 因子间相关性
sns.heatmap(factor_corr, annot=True, fmt='.3f', cmap='RdBu_r', center=0,
square=True, linewidths=0.5, ax=ax1, vmin=-1, vmax=1)
ax1.set_title('因子间相关性矩阵', fontsize=14)
# 原始变量与因子相关性
sns.heatmap(original_factor_corr, annot=True, fmt='.2f', cmap='RdBu_r', center=0,
square=True, linewidths=0.5, ax=ax2, vmin=-1, vmax=1)
ax2.set_title('原始变量-因子相关性', fontsize=14)
plt.tight_layout()
plt.show()
print("\n因子间相关系数:")
print(factor_corr.round(3))
5.3 方差解释率分析
python
# 计算方差解释率
fa_full = FactorAnalyzer(n_factors=len(feature_cols), rotation=None)
fa_full.fit(scaled_data)
ev_full, _ = fa_full.get_eigenvalues()
variance_explained = ev_full / sum(ev_full) * 100
cumulative_variance = np.cumsum(variance_explained)
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))
# 单个方差解释率
bars = ax1.bar(range(1, len(variance_explained)+1), variance_explained,
color='steelblue', alpha=0.7, edgecolor='black')
ax1.axhline(y=100/len(feature_cols), color='red', linestyle='--',
label=f'平均解释率 ({100/len(feature_cols):.1f}%)')
ax1.set_xlabel('因子序号', fontsize=12)
ax1.set_ylabel('方差解释率 (%)', fontsize=12)
ax1.set_title('各因子方差解释率', fontsize=14)
ax1.legend()
ax1.grid(True, alpha=0.3, axis='y')
# 标注前3个因子
for i in range(3):
ax1.text(i+1, variance_explained[i]+0.5, f'{variance_explained[i]:.1f}%',
ha='center', fontsize=9, fontweight='bold')
# 累计方差解释率
ax2.plot(range(1, len(cumulative_variance)+1), cumulative_variance,
'go-', linewidth=2, markersize=8, label='累计解释率')
ax2.axhline(y=70, color='orange', linestyle='--', label='70%阈值')
ax2.axhline(y=80, color='green', linestyle='--', label='80%阈值')
ax2.axvline(x=n_factors, color='red', linestyle=':', alpha=0.7, label=f'选定因子数={n_factors}')
# 标注选定因子数的累计解释率
ax2.scatter([n_factors], [cumulative_variance[n_factors-1]],
color='red', s=100, zorder=5)
ax2.text(n_factors+0.3, cumulative_variance[n_factors-1]-3,
f'{cumulative_variance[n_factors-1]:.1f}%', fontsize=10, color='red')
ax2.set_xlabel('因子数量', fontsize=12)
ax2.set_ylabel('累计方差解释率 (%)', fontsize=12)
ax2.set_title('累计方差解释率', fontsize=14)
ax2.legend()
ax2.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
print(f"\n方差解释率分析:")
print(f" 前{n_factors}个因子累计解释率: {cumulative_variance[n_factors-1]:.2f}%")
print(f" 因子1解释率: {variance_explained[0]:.2f}%")
print(f" 因子2解释率: {variance_explained[1]:.2f}%")
print(f" 因子3解释率: {variance_explained[2]:.2f}%")
5.4 三维因子空间可视化
python
from mpl_toolkits.mplot3d import Axes3D
fig = plt.figure(figsize=(12, 9))
ax = fig.add_subplot(111, projection='3d')
# 使用前3个因子
x, y, z = scores_df.iloc[:, 0], scores_df.iloc[:, 1], scores_df.iloc[:, 2]
# 根据时间着色
time_colors = np.linspace(0, 1, len(x))
scatter = ax.scatter(x, y, z, c=time_colors, cmap='viridis', alpha=0.6, s=20)
ax.set_xlabel('趋势因子', fontsize=11)
ax.set_ylabel('波动因子', fontsize=11)
ax.set_zlabel('动量因子', fontsize=11)
ax.set_title('期魔方因子三维空间分布(按时间着色)', fontsize=14)
cbar = plt.colorbar(scatter, shrink=0.6, aspect=15)
cbar.set_label('时间序列', fontsize=10)
plt.show()
5.5 因子雷达图(多维度对比)
python
from math import pi
# 计算各因子在不同时间段的均值
time_periods = {
'2020年': scores_df.loc[:300].mean(),
'2021年': scores_df.loc[300:600].mean(),
'2022年': scores_df.loc[600:900].mean(),
'2023年': scores_df.loc[900:].mean(),
}
# 雷达图
fig, ax = plt.subplots(figsize=(10, 10), subplot_kw=dict(projection='polar'))
categories = list(scores_df.columns)
N = len(categories)
angles = [n / float(N) * 2 * pi for n in range(N)]
angles += angles[:1]
colors = ['#FF6B6B', '#4ECDC4', '#45B7D1', '#96CEB4']
for idx, (period, values) in enumerate(time_periods.items()):
values_list = values.values.tolist()
values_list += values_list[:1]
ax.plot(angles, values_list, 'o-', linewidth=2, label=period, color=colors[idx])
ax.fill(angles, values_list, alpha=0.15, color=colors[idx])
ax.set_xticks(angles[:-1])
ax.set_xticklabels(categories, fontsize=11)
ax.set_ylim(-2, 2)
ax.set_title('期魔方因子年度对比雷达图', fontsize=14, pad=30)
ax.legend(loc='upper right', bbox_to_anchor=(1.3, 1.0))
ax.grid(True)
plt.tight_layout()
plt.show()
六、参数对比与模型选择
6.1 旋转方法对比实验
python
# 对比不同旋转方法的效果
rotation_methods = ['varimax', 'quartimax', 'equamax', 'promax', 'oblimin']
rotation_results = []
for method in rotation_methods:
try:
fa_test = FactorAnalyzer(n_factors=n_factors, rotation=method, method='ml')
fa_test.fit(scaled_data)
loadings_test = fa_test.loadings_
# 计算简单结构指数(越接近1越好)
ss_index = np.sum(loadings_test**4, axis=1).sum() / np.sum(loadings_test**2, axis=1).sum()
# 计算因子间相关性均值(正交旋转应接近0)
scores_test = fa_test.transform(scaled_data)
factor_corr_test = np.corrcoef(scores_test.T)
corr_mean = np.mean(np.abs(factor_corr_test[np.triu_indices_from(factor_corr_test, k=1)]))
rotation_results.append({
'旋转方法': method,
'简单结构指数': ss_index,
'因子间平均相关': corr_mean,
'是否正交': method in ['varimax', 'quartimax', 'equamax']
})
except Exception as e:
print(f"{method} 失败: {e}")
rotation_df = pd.DataFrame(rotation_results)
print("\n旋转方法对比:")
print(rotation_df.round(4))
# 可视化对比
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14, 6))
x_pos = np.arange(len(rotation_df))
ax1.bar(x_pos, rotation_df['简单结构指数'], color='steelblue', alpha=0.7, edgecolor='black')
ax1.set_xticks(x_pos)
ax1.set_xticklabels(rotation_df['旋转方法'], rotation=45)
ax1.set_ylabel('简单结构指数', fontsize=12)
ax1.set_title('旋转方法简单结构对比', fontsize=14)
ax1.grid(True, alpha=0.3, axis='y')
colors = ['green' if o else 'orange' for o in rotation_df['是否正交']]
ax2.bar(x_pos, rotation_df['因子间平均相关'], color=colors, alpha=0.7, edgecolor='black')
ax2.set_xticks(x_pos)
ax2.set_xticklabels(rotation_df['旋转方法'], rotation=45)
ax2.set_ylabel('因子间平均绝对相关', fontsize=12)
ax2.set_title('因子相关性对比(绿=正交,橙=斜交)', fontsize=14)
ax2.grid(True, alpha=0.3, axis='y')
plt.tight_layout()
plt.show()
6.2 估计方法对比
python
from sklearn.decomposition import FactorAnalysis as SKFactorAnalysis
# sklearn的FactorAnalysis对比
estimation_results = []
# 不同参数配置
configs = [
{'name': 'ML+SVD', 'svd_method': 'randomized', 'rotation': None},
{'name': 'ML+LAPACK', 'svd_method': 'lapack', 'rotation': None},
]
for config in configs:
fa_sk = SKFactorAnalysis(n_components=n_factors,
svd_method=config['svd_method'],
max_iter=1000)
fa_sk.fit(scaled_data)
estimation_results.append({
'配置': config['name'],
'对数似然': fa_sk.loglike_[-1],
'迭代次数': len(fa_sk.loglike_),
'收敛': len(fa_sk.loglike_) < 1000
})
est_df = pd.DataFrame(estimation_results)
print("\n估计方法对比:")
print(est_df)
# 对数似然收敛曲线
fig, ax = plt.subplots(figsize=(10, 6))
for config in configs:
fa_sk = SKFactorAnalysis(n_components=n_factors,
svd_method=config['svd_method'],
max_iter=1000)
fa_sk.fit(scaled_data)
ax.plot(fa_sk.loglike_, label=config['name'], linewidth=2)
ax.set_xlabel('迭代次数', fontsize=12)
ax.set_ylabel('对数似然', fontsize=12)
ax.set_title('因子分析收敛曲线对比', fontsize=14)
ax.legend()
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
6.3 参数选择决策矩阵
| 参数维度 | 选项 | 期魔方推荐 | 理由 |
|---|---|---|---|
| 因子数量 | 平行分析 / Kaiser / 理论驱动 | 平行分析 | 避免过拟合,更稳健 |
| 旋转方法 | Varimax / Promax / Oblimin | Promax | 金融因子通常相关 |
| 估计方法 | ML / ULS / PAF | ML | 统计性质好,可检验 |
| 标准化 | Standard / Robust / Rank | Robust | 抗极端值干扰 |
| 缺失处理 | 删除 / 插值 / EM算法 | 前向填充+删除 | 保持时序一致性 |
七、因子分析实战策略
7.1 因子得分策略构建
python
# 基于因子得分构建交易策略信号
def generate_factor_signals(scores_df, weights=None):
"""
基于因子得分生成交易信号
Parameters:
-----------
scores_df : DataFrame
因子得分矩阵
weights : dict
因子权重,默认等权
Returns:
--------
signals : Series
综合信号 (-1:做空, 0:观望, 1:做多)
"""
if weights is None:
weights = {col: 1/len(scores_df.columns) for col in scores_df.columns}
# 加权综合得分
composite_score = sum(scores_df[col] * weight
for col, weight in weights.items())
# 分位数生成信号
q75 = composite_score.quantile(0.75)
q25 = composite_score.quantile(0.25)
signals = pd.Series(0, index=composite_score.index)
signals[composite_score > q75] = 1 # 做多
signals[composite_score < q25] = -1 # 做空
return signals
# 不同权重配置对比
weight_configs = {
'等权配置': {'趋势因子': 1/3, '波动因子': 1/3, '动量因子': 1/3},
'趋势优先': {'趋势因子': 0.5, '波动因子': 0.25, '动量因子': 0.25},
'波动优先': {'趋势因子': 0.25, '波动因子': 0.5, '动量因子': 0.25},
'动量优先': {'趋势因子': 0.25, '波动因子': 0.25, '动量因子': 0.5},
}
fig, axes = plt.subplots(2, 2, figsize=(16, 12))
axes = axes.flatten()
for idx, (name, weights) in enumerate(weight_configs.items()):
signals = generate_factor_signals(scores_df, weights)
signal_counts = signals.value_counts()
colors = ['red', 'gray', 'green']
axes[idx].bar(['做空(-1)', '观望(0)', '做多(1)'],
[signal_counts.get(-1, 0), signal_counts.get(0, 0), signal_counts.get(1, 0)],
color=colors, alpha=0.7, edgecolor='black')
axes[idx].set_title(f'{name}\n权重: {weights}', fontsize=11)
axes[idx].set_ylabel('交易天数', fontsize=10)
axes[idx].grid(True, alpha=0.3, axis='y')
plt.suptitle('期魔方因子策略信号分布对比', fontsize=14)
plt.tight_layout()
plt.show()
7.2 因子IC分析
python
# 计算因子IC(信息系数)
def calculate_ic(factor_scores, forward_returns):
"""
计算因子与下期收益的秩相关系数
IC > 0: 正向预测能力
IC < 0: 负向预测能力
|IC| > 0.03: 具有一定预测能力
"""
ic_values = []
for col in factor_scores.columns:
ic = stats.spearmanr(factor_scores[col], forward_returns)[0]
ic_values.append(ic)
return pd.Series(ic_values, index=factor_scores.columns)
# 模拟下期收益(实际应使用期魔方的future_return字段)
np.random.seed(42)
forward_returns = np.random.randn(len(scores_df)) * 0.02
ic_results = calculate_ic(scores_df, forward_returns)
# IC可视化
plt.figure(figsize=(10, 6))
colors = ['green' if ic > 0 else 'red' for ic in ic_results]
bars = plt.bar(ic_results.index, ic_results.values, color=colors, alpha=0.7, edgecolor='black')
plt.axhline(y=0, color='black', linewidth=0.8)
plt.axhline(y=0.03, color='green', linestyle='--', alpha=0.5, label='IC=0.03')
plt.axhline(y=-0.03, color='red', linestyle='--', alpha=0.5, label='IC=-0.03')
plt.xlabel('因子', fontsize=12)
plt.ylabel('IC值(秩相关系数)', fontsize=12)
plt.title('期魔方因子IC分析', fontsize=14)
plt.legend()
plt.grid(True, alpha=0.3, axis='y')
# 标注IC值
for bar, ic in zip(bars, ic_results.values):
plt.text(bar.get_x() + bar.get_width()/2, ic + 0.002 if ic > 0 else ic - 0.005,
f'{ic:.3f}', ha='center', fontsize=10, fontweight='bold')
plt.tight_layout()
plt.show()
print("\n因子IC分析结果:")
print(ic_results.round(4))
7.3 滚动因子稳定性分析
python
# 滚动窗口因子分析
def rolling_factor_analysis(data, window=60, step=20):
"""
滚动窗口因子分析,检测因子结构稳定性
Parameters:
-----------
window : int
滚动窗口大小
step : int
滚动步长
"""
results = []
for start in range(0, len(data) - window, step):
end = start + window
window_data = data.iloc[start:end]
# 标准化
scaler = StandardScaler()
scaled = scaler.fit_transform(window_data)
# 因子分析
fa = FactorAnalyzer(n_factors=3, rotation='promax')
fa.fit(scaled)
results.append({
'start': start,
'end': end,
'factor1_explained': fa.get_factor_variance()[1][0],
'factor2_explained': fa.get_factor_variance()[1][1],
'factor3_explained': fa.get_factor_variance()[1][2],
})
return pd.DataFrame(results)
# 执行滚动分析
rolling_results = rolling_factor_analysis(clean_data, window=120, step=30)
# 可视化因子稳定性
fig, ax = plt.subplots(figsize=(14, 6))
ax.plot(rolling_results.index, rolling_results['factor1_explained'],
'o-', label='趋势因子解释率', linewidth=2)
ax.plot(rolling_results.index, rolling_results['factor2_explained'],
's-', label='波动因子解释率', linewidth=2)
ax.plot(rolling_results.index, rolling_results['factor3_explained'],
'^-', label='动量因子解释率', linewidth=2)
ax.set_xlabel('滚动窗口序号', fontsize=12)
ax.set_ylabel('方差解释率', fontsize=12)
ax.set_title('期魔方因子结构稳定性分析(滚动窗口)', fontsize=14)
ax.legend()
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
print("\n滚动窗口因子解释率统计:")
print(rolling_results[['factor1_explained', 'factor2_explained', 'factor3_explained']].describe().round(4))
八、模型诊断与验证
8.1 残差分析
python
# 计算模型残差
reproduced_corr = fa.loadings_ @ fa.loadings_.T + np.diag(fa.get_uniquenesses())
actual_corr = np.corrcoef(scaled_data.T)
residuals = actual_corr - reproduced_corr
# 残差可视化
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6))
# 残差热力图
mask = np.triu(np.ones_like(residuals, dtype=bool))
sns.heatmap(residuals, mask=mask, annot=True, fmt='.2f',
cmap='RdBu_r', center=0, square=True,
xticklabels=feature_cols, yticklabels=feature_cols, ax=ax1)
ax1.set_title('残差相关矩阵', fontsize=14)
ax1.tick_params(axis='x', rotation=45)
# 残差分布
residuals_flat = residuals[np.tril_indices_from(residuals, k=-1)]
ax2.hist(residuals_flat, bins=30, alpha=0.7, color='steelblue', edgecolor='black')
ax2.axvline(x=0, color='red', linestyle='--', linewidth=2)
ax2.set_xlabel('残差值', fontsize=12)
ax2.set_ylabel('频数', fontsize=12)
ax2.set_title('残差分布直方图', fontsize=14)
ax2.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
print(f"\n残差分析:")
print(f" 残差均值: {np.mean(np.abs(residuals_flat)):.4f}")
print(f" 残差标准差: {np.std(residuals_flat):.4f}")
print(f" 最大绝对残差: {np.max(np.abs(residuals_flat)):.4f}")
print(f" 残差<0.05的比例: {np.mean(np.abs(residuals_flat) < 0.05):.1%}")
8.2 模型拟合指标
python
# 计算模型拟合指标
def calculate_fit_indices(actual_corr, reproduced_corr, n_obs, n_factors):
"""
计算因子分析模型拟合指标
"""
p = actual_corr.shape[0]
# 残差矩阵
residuals = actual_corr - reproduced_corr
# 标准化残差均方根 (SRMR)
srmr = np.sqrt(np.sum(residuals**2) / (p * (p + 1) / 2))
# 近似误差均方根 (RMSEA)
df = p * (p + 1) / 2 - p * n_factors + n_factors * (n_factors - 1) / 2
chi2 = (n_obs - 1) * np.trace(np.linalg.inv(reproduced_corr) @ actual_corr) - (n_obs - 1) * p
rmsea = np.sqrt(max(chi2 - df, 0) / (df * (n_obs - 1)))
return {
'SRMR': srmr,
'RMSEA': rmsea,
'χ²/df': chi2 / df if df > 0 else np.inf,
'AIC': chi2 - 2 * df,
'BIC': chi2 - df * np.log(n_obs)
}
fit_indices = calculate_fit_indices(actual_corr, reproduced_corr,
len(scaled_data), n_factors)
print("\n模型拟合指标:")
print("=" * 50)
for key, value in fit_indices.items():
print(f" {key:<10}: {value:.4f}")
print("=" * 50)
print("\n判定标准:")
print(" SRMR < 0.08: 良好")
print(" RMSEA < 0.08: 可接受")
print(" χ²/df < 3: 良好")
九、因子分析最佳实践
9.1 完整工作流程
期魔方因子分析标准工作流程
│ 1. 数据准备
│ ├── 从期魔方导出原始行情数据
│ ├── 计算技术指标与衍生因子
│ └── 处理缺失值与异常值
│
│ 2. 数据预处理
│ ├── 去极值(Winsorize)
│ ├── 标准化(RobustScaler推荐)
│ └── 共线性检测(VIF分析)
│
│ 3. 适用性检验
│ ├── KMO检验(>0.6)
│ └── Bartlett球形检验(p<0.05)
│
│ 4. 因子提取
│ ├── 平行分析确定因子数
│ ├── 选择估计方法(ML推荐)
│ └── 因子旋转(Promax推荐)
│
│ 5. 模型诊断
│ ├── 残差分析
│ ├── 拟合指标评估
│ └── 因子可解释性验证
│
│ 6. 策略应用
│ ├── 因子得分计算
│ ├── IC分析验证预测能力
│ └── 构建交易信号
│
│ 7. 持续监控
│ ├── 滚动因子稳定性分析
│ └── 模型定期重训练
9.2 常见问题与解决方案
| 问题 | 表现 | 原因 | 解决方案 |
|---|---|---|---|
| Heywood案例 | 共同度>1 | 样本量不足或因子数过多 | 减少因子数,增加样本 |
| 因子模糊 | 变量跨多个因子高载荷 | 因子间界限不清 | 尝试不同旋转,合并因子 |
| 负方差 | 特殊因子方差为负 | 模型设定不当 | 检查数据,调整因子数 |
| 不收敛 | 迭代不收敛 | 数据质量问题 | 检查异常值,增加迭代次数 |
| 因子不稳定 | 滚动窗口结果差异大 | 市场结构变化 | 缩短窗口,动态调整 |
9.3 API对接建议
python
# 期魔方数据获取示例(伪代码)
"""
# 1. 通过期魔方API获取数据
from qimofang import QIMOData
client = QIMOData(api_key='your_api_key')
# 获取螺纹钢日线数据
data = client.get_daily_data(
symbol='RB.SHF',
start_date='2020-01-01',
end_date='2024-12-31',
fields=['open', 'high', 'low', 'close', 'volume']
)
# 2. 计算期魔方内置因子
factors = client.calculate_factors(
data,
factor_list=['MA', 'RSI', 'MACD', 'ATR', 'BOLLINGER']
)
# 3. 导出进行因子分析
factors.to_csv('qimo_factors.csv')
# 4. 使用本文方法进行因子分析
# ...(见前文代码)
"""
十、总结与展望
10.1 核心要点回顾
- 因子体系:涵盖价量、技术、统计、基本面、宏观、另类六大类因子
- 因子分析流程:数据清洗 → 适用性检验 → 因子提取 → 旋转 → 诊断 → 应用
- 参数选择策略:平行分析定因子数、Promax旋转、ML估计、Robust标准化
- 可视化体系:热力图、时序图、雷达图、三维图、IC分析等多维度展示
- 策略应用:因子得分加权 → 信号生成 → 回测验证
10.2 进阶方向
| 方向 | 方法 | 应用场景 |
|---|---|---|
| 动态因子模型 | 滚动窗口、状态空间模型 | 捕捉因子结构时变 |
| 非线性因子分析 | 核PCA、自编码器 | 非线性特征提取 |
| 高维因子分析 | 稀疏PCA、因子选择 | 大数据量场景 |
| 多品种因子分析 | 面板数据因子模型 | 跨品种策略 |
| 深度学习因子 | LSTM因子提取、注意力机制 | 高频数据 |
10.3 关键参数速查表
| 参数 | 推荐值 | 备选值 | 调整依据 |
|---|---|---|---|
| 因子数量 | 平行分析结果 | 3-5 | 累计解释率>70% |
| 旋转方法 | Promax | Varimax | 因子是否应独立 |
| 估计方法 | ML | ULS | 数据分布 |
| 标准化 | RobustScaler | RankScaler | 异常值比例 |
| 去极值分位 | 2.5%-97.5% | 1%-99% | 数据质量 |
| 滚动窗口 | 120日 | 60-250日 | 品种流动性 |
结语:机器学习模块为量化投资者提供了强大的因子研究工具。通过系统化的因子分析流程,我们可以从海量原始特征中提取出具有经济含义的公共因子,构建更加稳健的交易策略。记住,因子分析不是终点,而是策略开发的起点------持续的模型监控与迭代优化才是量化投资的核心竞争力。
参考资源:
- 期魔方官方文档: https://www.qimofang.com/
- FactorAnalyzer文档: https://factor-analyzer.readthedocs.io/
- sklearn FactorAnalysis: https://scikit-learn.org/stable/modules/decomposition.html
- 《量化投资:以Python为工具》--- 蔡立耑
*本文基于期魔方平台特性与开源因子分析库撰写,代码可直接运行于Python环境。欢迎交流讨论。