Preface
In embedded development, screen refresh rate is the “first lifeline” of user experience. I have always been curious, in order to pursue high frame rates, should we “optimize the 60MHz SPI to the extreme”, or should we directly “upgrade the dimension” and use an 8-bit parallel port? In order to clarify this problem, I decided to purchase hardware at my own expense to conduct an “Original Showdown”
test items
I deliberately chose two tests with “different dimensions”:
- TFT_eSPI’s graphicstest is a ‘raw I/O’ test, which only measures the ‘bus’ throughput speed.
- lv_demo_music is a ‘real world’ test which measures the ‘overall performance’ of ‘CPU rendering + bus I/O’
Test results
Run the lv_demo_music that comes with LVGL
| Control group number | Screen parameters | ESP32 operating frequency | LVGL buffer parameters | SPI rate | TFT_eSPI version | LVGL version | Frame rate |
|---|---|---|---|---|---|---|---|
| 1 | 240x240 1.54-inch SPI LCD | 240MHz | 240*120 DMA enabled, double buffering not used | 60MHz | 2.3.70 | 8.1.1-dev | 26 frames |
| 2 | 240x240 1.33-inch 8-bit parallel LCD | 240MHz | 240*120 Not using double buffering | \ | 2.3.70 | 8.1.1-dev | 28 frames |
It feels like the frame rate difference isn’t that big.
Run the Viewport_graphicstest that comes with TFT_eSPI
| Control group number | Screen parameters | ESP32 operating frequency | SPI rate | TFT_eSPI version |
|---|---|---|---|---|
| 1 | 240x240 1.54-inch SPI LCD | 240MHz | 60MHz | 2.3.70 |
| 2 | 240x240 1.33-inch 8-bit parallel LCD | 240MHz | \ | 2.3.70 |
When running, it is obvious that the drawing speed of the parallel port is faster, but the data still has to speak, so there is the following Log
8-bit parallel port results
TFT_eSPI library test!
Benchmark Time (microseconds)
Screen fill 53073
Text 21126
Lines 36424
Horiz/Vert Lines 10469
Rectangles (outline) 5918
Rectangles (filled) 122742
Circles (filled) 28151
Circles (outline) 22372
Triangles (outline) 14111
Triangles (filled) 46193
Rounded rects (outline) 16475
Rounded rects (filled) 127122
Done!
serial port results
TFT_eSPI library test!
Benchmark Time (microseconds)
Screen fill 102725
Text 12506
Lines 38321
Horiz/Vert Lines 9293
Rectangles (outline) 6594
Rectangles (filled) 234378
Circles (filled) 30350
Circles (outline) 16884
Triangles (outline) 12893
Triangles (filled) 79688
Rounded rects (outline) 11730
Rounded rects (filled) 240258
Done!
The drawing time under the parallel port driver is almost half of the drawing time under the serial port driver.
Sure enough, there is still improvement.
As for why the gap cannot be widened in LVGL, after being asked by group friend Principle, I realized that there is a bottleneck in CPU performance when running LVGL. My ESP32 is the first-generation version. If I switch to the latest ESP32 S3, the gap will widen.
Conclusion analysis: CPU bottleneck vs I/O bottleneck
Looking at the logs of TFT_eSPI (raw I/O), the results are clear at a glance: the parallel port is almost twice as fast (e.g. Screen fill 53,073μs vs 102,725μs). This proves that the ‘parallel port bus’ has obvious physical advantages. So why is the difference so small (26 vs 28 frames) in the ‘real world’ (LVGL)? This reveals the real bottleneck of the ESP32 (original generation): CPU rendering speed. LVGL’s ‘lv_demo_music’ effect requires a lot of CPU calculations, causing the CPU to be 100% full, while the screen bus (whether SPI or parallel port) is actually ‘idle waiting’.
Summary: So if your project is ‘heavy rendering’, blindly upgrading the screen interface is useless, Then you should first upgrade the MCU (such as ESP32-S3).
Extended reading
Later, I read an article by Mr. P. He compared the lvgl effect from the chip dimension. He conducted an in-depth analysis of some CPU and Flash-related content and benefited a lot. I recommend that all readers also read it ((Note: The picture bed of Mr. P’s blog seems to be down, but it does not affect the gold content of the text)) AT32F403A demo-transplanted LVGL and compared with ESP32-S3