外观
列联表分析
考察两个分类变量是否有关联,例如不同教育水平的使用意愿构成。
计算口径
每名对象只能贡献一条独立记录。仅使用两个变量均非缺失的记录。Fisher精确检验限2×2表。
分别指定行变量与列变量,选择行、列、总百分比或只显示频数。默认使用无连续性校正的Pearson卡方;2×2稀疏表可主动选择Fisher。
行/列百分比回答不同问题,报告时注明分母。卡方检验判断整体关联,Cramér V描述关联强度;原始输出保留期望频数。
同源实现
以下片段来自 core/statistical/descriptive.py 的 crosstab,由镜像脚本按语法树提取。共享辅助函数和分发逻辑包含在完整下载包中。
py
def crosstab(data, options):
selected = columns(data, options.get('variables'), 2, 2)
sample = enough(data[selected].dropna(), 2)
table = pd.crosstab(sample[selected[0]], sample[selected[1]], dropna=True)
if min(table.shape) < 2 or max(table.shape) > 50:
raise ValueError(_('列联表的行、列变量都需要2至50个有效类别'))
test = choice(options, 'test', 'chi2', ('chi2', 'fisher'))
percent = choice(options, 'percent', 'row', ('row', 'column', 'total', 'none'))
chi2, p_value, df, expected = stats.chi2_contingency(table, correction=False)
association = float(np.sqrt(chi2/(table.values.sum()*min(table.shape[0]-1, table.shape[1]-1))))
raw = {'n': len(sample), 'observed': table, 'expected': expected, 'chi2': chi2, 'df': df,
'chi2_p': p_value, 'cramers_v': association, 'expected_below_five': int((expected < 5).sum())}
if test == 'fisher':
if table.shape != (2, 2):
raise ValueError(_('Fisher精确检验首版适用于2×2列联表'))
odds, p_value = stats.fisher_exact(table.values, alternative='two-sided')
raw.update(odds_ratio=odds, fisher_p=p_value)
rows, headers = [], [Column('category', selected[0], 'text')]
for i, label in enumerate(table.columns):
headers.append(Column('c'+str(i), str(label), 'text'))
headers.append(Column('total', _('合计'), 'integer'))
for label, cells in table.iterrows():
row = {'category': str(label), 'total': int(cells.sum())}
for i, count in enumerate(cells):
denominator = cells.sum() if percent == 'row' else table.iloc[:, i].sum() if percent == 'column' else len(sample)
row['c'+str(i)] = str(count) if percent == 'none' else '%d (%.1f%%)' % (count, 100*count/denominator)
rows.append(row)
p_text = '<0.001' if p_value < .001 else '=%.3f' % p_value
test_text = 'Fisher' if test == 'fisher' else 'χ²(%d)=%.3f' % (df, chi2)
note = _('N=%(n)s;%(test)s,p%(p)s;Cramér’s V=%(v)s。', n=len(sample), test=test_text, p=p_text, v='%.3f' % association)
return StatisticalResult([PaperTable(_('列联表分析'), headers, rows, note)], raw)复现本方法
在解压目录安装 requirements.txt 后执行。样例为固定种子的模拟数据,仅供验证;输出不得冒充真实研究结果。
python
import examples._bootstrap
from examples.statistics_cases import build_data, method_cases
from core.statistical.runner import analyze
method = "stat_crosstab"
options = dict(method_cases())[method]
result = analyze(build_data(), method, options)
print(result.tables[0].html())