I built MOLT, a Windows-first, thermally aware QLoRA fine-tuning runtime for consumer NVIDIA GPUs.
Its newest qualified benchmark used one million Qwen 1.5B training targets on an RTX 4060 Laptop GPU, with both MOLT and Unsloth using the same initial adapter, prepared token order, AC power, and a 140W GPU power limit.
Results:
| Metric | MOLT | Tested Unsloth |
|—|—:|—
| End-to-end session | 1,103 s | 1,433 s |
| Training throughput | 1,063 targets/s | 768 targets/s |
| Peak allocated VRAM | 1.63 GB | 1.85 GB |
| Measured board energy | 63.1 kJ | 63.8 kJ |
That was 23% lower total training time, 38% higher throughput, and 12% lower allocated VRAM in this specific test.
MOLT also includes GPU telemetry, thermal pacing, resumable checkpoints, safe recovery, and adapter export.
This is single-machine, single-seed development evidence—not a claim that MOLT universally beats Unsloth. I’m looking for Windows NVIDIA GPU users to reproduce the benchmark and test their own workloads.