YOLO-KAN: What the Ablation Experiments Taught Me
A concise research note on introducing Kolmogorov-Arnold Network modules into YOLO11n, testing flatten strategies, and reading the resulting trade-offs.

The goal of this research was deliberately narrow: improve YOLO11n accuracy while reducing network depth. Kolmogorov-Arnold Networks (KANs) were a promising candidate because they replace fixed node activations with learnable functions on edges, potentially extracting richer features without simply stacking more layers.
From an MLP block to a KAN block
I introduced KAN modules into the YOLO backbone, then compared several ways of serializing feature maps before they entered the KAN layer:
- depthwise flattening;
- max-pool flattening; and
- convolutional flattening with different kernel sizes.
This detail mattered more than expected. The flatten structure materially changed both the accuracy curve and the spatial features emphasized by the detector.
What the experiments showed
The baseline YOLO11n reached 63.99% precision. The highest-precision configuration, KAN-2-5, reached 65.83%. In the full ablation table, KAN-1-7 produced the highest mAP@50 at 54.51%; KAN-2-7 reached 54.48%. Across the experiments, the largest precision gain was 1.84 percentage points.
The simplified architecture also reduced the layer count from 319 to 299. That result supported the original hypothesis: the KAN structure could improve feature extraction without depending on a deeper network.
Heatmaps revealed the trade-off
The heatmaps made the model behavior easier to interpret. YOLO11n focused mainly on the train's most obvious features. KAN-2-5 attended to both the train and useful background context. With the larger KAN-2-7 receptive field, attention spread further into non-primary regions.
That made KAN-2-5 the best overall balance in this study—not because it won every metric, but because it improved accuracy without losing focus.
The practical lesson
Adding a new module is only the beginning of an architecture experiment. How data is reshaped before the module can determine whether the module works at all. In this project, flatten-layer design was as important as the KAN block itself.
Reading and reproducing the results
These figures are the observations reported in the original research poster for Microsoft COCO and YOLO11n. KAN-2-5 uses two KAN modules with a 5 × 5 convolutional flatten layer. Its precision gain came with 49.24% recall, compared with 49.49% for the baseline. The poster also reports increased computational cost for the KAN variants despite their reduced layer count. Accuracy, recall, and computation therefore need to be assessed together.
The public YOLO-KAN repository contains model checkpoints and an entry point for evaluation. In the linked revision, val.py uses a 640-pixel image size, batch size 64, and JSON prediction export. Before rerunning, update the local dataset and checkpoint paths, verify the intended evaluation split, and review the resume=True setting in train.py. The checked-in COCO configuration points both train and val to train2017.txt. This repository snapshot does not establish which split was used for the original poster. Reproducing that table requires confirming the original run configuration and matching checkpoints.
Explore the project page and complete research poster, or open the original PDF.