Skip to content

Add DeBERTa-ConPara-v2.17: 1M leakage-free stratified training set - #181

Open
MohamedMady19 wants to merge 1 commit into
liamdugan:mainfrom
MohamedMady19:submission/conpara-v217
Open

Add DeBERTa-ConPara-v2.17: 1M leakage-free stratified training set#181
MohamedMady19 wants to merge 1 commit into
liamdugan:mainfrom
MohamedMady19:submission/conpara-v217

Conversation

@MohamedMady19

Copy link
Copy Markdown

DeBERTa-ConPara v2.17 - retrained on a 1M-sample leakage-free stratified dataset.

Changes vs v2.16:

  • Training data rebuilt: RAID splits grouped by adv_source_id to prevent attack-variant leakage between train and val
  • Source-stratified composition (no source above 45% of either class)
  • Attack-augmented human training (~11 RAID attack variants per human article)
  • Corrected Unicode homoglyph normalisation (U+0440 was mapping to 'r' instead of 'p', corrupting rather than restoring attacked text)

Same architecture as v2.16 (DeBERTa-v3-large + 62 linguistic features);
Only the training data and preprocessing changed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant