Code Generation On Res Q
评估指标
pass@1
评测结果
各个模型在此基准测试上的表现结果
比较表格
模型名称 | pass@1 |
---|---|
res-q-evaluating-code-editing-large-language | 30.0 |
res-q-evaluating-code-editing-large-language | 58.0 |
res-q-evaluating-code-editing-large-language | 20.0 |
res-q-evaluating-code-editing-large-language | 18.0 |
res-q-evaluating-code-editing-large-language | 30.0 |
res-q-evaluating-code-editing-large-language | 36.0 |
res-q-evaluating-code-editing-large-language | 46.0 |
res-q-evaluating-code-editing-large-language | 29.0 |
res-q-evaluating-code-editing-large-language | 37.0 |