CodeBERT for Flaky Test Categorisation (FlakeBench)

Classifies a Java/Kotlin test method into one of six categories: five kinds of flaky test plus non-flaky.

What this is

A fine-tune of microsoft/codebert-base on the FlakeBench dataset from Understanding and Improving Flaky Test Classification (OOPSLA 2025), trained as a reproduction exercise on a single 8 GB consumer GPU. The uploaded weights are the "Balanced" configuration below.

Training configurations

Parameter Baseline lr 2e-5 Balanced Augmented Paper
Encoder codebert-base codebert-base codebert-base codebert-base codebert-base
Learning rate 1e-5 2e-5 1e-5 1e-5 1e-5
Batch size 8 8 8 8 8
Max length 512 512 512 512 512
Loss focal γ=2.0 focal γ=2.0 focal γ=2.0 focal γ=2.0 focal γ=2.0
Class weights balanced balanced balanced balanced balanced
Optimizer AdamW wd 0.01 AdamW wd 0.01 AdamW wd 0.01 AdamW wd 0.01 AdamW wd 0.01
Precision fp16 fp16 fp16 fp16 fp32
Non-flaky rows 4,972 4,972 800 800 full
Minority handling none none ×160 copies ×200 variants none
Train rows 5,114 5,114 1,600 1,800 5,114
Epochs run 8 8 18 13 40
Dynamic padding no no no no no

Hardware: 1× RTX 4060 Laptop (8 GB). Class rebalancing is the one deviation from the paper's method, which trains on the raw distribution (97% non-flaky).

Results (per-category F1)

Category Baseline lr 2e-5 Balanced Augmented Paper
Async Wait 76.92% 78.26% 74.07% 64.52% 58.37%
Concurrency 0.00% 0.00% 0.00% 0.00% 35.92%
Time 57.14% 66.67% 66.67% 40.00% 72.73%
Unordered Coll. 75.00% 83.33% 83.33% 72.73% 73.63%
Order Dep. 82.35% 86.96% 95.24% 73.68% 64.35%
Non-flaky 100.00% 99.92% 99.51% 100.00% 100.00%
Macro F1 65.24% 69.19% 69.89% 58.49% 65.79%

The Balanced configuration (uploaded weights) achieves the best macro-F1 of 69.89% .

Usage

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

name = "Ariful1904129/codebert-flakytest-fold2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name, trust_remote_code=True).eval()

code = """@Test
public void testConnect() throws Exception {
    Thread.sleep(1000);
    assertTrue(client.isConnected());
}"""

x = tok(code, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    pred = model(**x).logits.argmax(-1).item()
print(model.config.id2label[pred])

Scope: Java/Kotlin test methods; inputs longer than 512 tokens are truncated.

Citation

Please cite the original paper. This model is a third-party reproduction and is not endorsed by its authors.

@inproceedings{flakylens2025,
  title     = {Understanding and Improving Flaky Test Classification},
  booktitle = {OOPSLA},
  year      = {2025}
}

Dataset and method: UT-SE-Research/FlakyLens.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ariful1904129/codebert-flakytest-fold2

Finetuned
(147)
this model