外观
多重响应分析
整理问卷类别构成,回答各选项有多少人选择。
计算口径
一行代表一名受访者。缺失回答与“未选中”含义不同,编码前先区分。
二分编码模式:每个选项一列,指定选中与未选中编码。位置编码模式:每个选择位置一列,同一人重复选择同一选项只计一次。有效个案的定义随编码模式记录在原始输出中。
先核对有效人数与分母,再报告频数和百分比;多选题的个案百分比之和可能超过100%。
同源实现
以下片段来自 core/statistical/descriptive.py 的 multi_response,由镜像脚本按语法树提取。共享辅助函数和分发逻辑包含在完整下载包中。
py
def multi_response(data, options):
selected = columns(data, options.get('variables'))
mode = choice(options, 'response_mode', 'binary', ('binary', 'slots'))
frame = data[selected]
# 整行缺失不进入分母;填写了全部0的个案仍属于有效个案。
frame = enough(frame.loc[frame.notna().any(axis=1)], 1)
counts = {}
if mode == 'binary':
yes, no = options.get('selected_code', 1), options.get('unselected_code', 0)
if _same_code(yes, no):
raise ValueError(_('选中与未选中的编码必须不同'))
for name in selected:
values = frame[name].dropna()
if any(not _same_code(v, yes) and not _same_code(v, no) for v in values):
raise ValueError(_('多选题变量“%(name)s”包含未定义的编码', name=name))
counts[name] = sum(_same_code(v, yes) for v in values)
else:
for values in frame.itertuples(index=False, name=None):
# 同一个受访者重复填写同一选项只计一次。
chosen = {str(v) for v in values if not pd.isna(v) and str(v).strip()}
for category in sorted(chosen):
counts[category] = counts.get(category, 0) + 1
responses = sum(counts.values())
if not responses:
raise ValueError(_('有效个案没有选择任何选项,无法计算响应百分比'))
rows = [{'category': key, 'n': count, 'response_percent': 100*count/responses,
'case_percent': 100*count/len(frame)} for key, count in counts.items()]
table = PaperTable(_('多重响应分析'), [Column('category', _('选项'), 'text'), Column('n', _('响应数'), 'integer'),
Column('response_percent', _('响应百分比'), 'percent'),
Column('case_percent', _('个案百分比'), 'percent')], rows,
_('有效个案N=%(n)s;响应总数=%(responses)s。', n=len(frame), responses=responses))
return StatisticalResult([table], {'valid_cases': len(frame), 'total_responses': responses,
'excluded_cases': len(data)-len(frame), 'coding': mode})复现本方法
在解压目录安装 requirements.txt 后执行。样例为固定种子的模拟数据,仅供验证;输出不得冒充真实研究结果。
python
import examples._bootstrap
from examples.statistics_cases import build_data, method_cases
from core.statistical.runner import analyze
method = "stat_multi_response"
options = dict(method_cases())[method]
result = analyze(build_data(), method, options)
print(result.tables[0].html())