To compare Embedded Proto to alternatives such as NanoPB and ArduinoJson in terms of performance. Performance in this case is specified as flash, RAM, CPU, and buffer usage. ArduinoJson is not a protobuf library but is considered a likely alternative to protobuf, so it is used as a comparison.
Two sets live in Scenarios/<Library>/. The published set is what the website shows, six
messages shaped like real device traffic with the same values in every library
(Scenarios/scenario_values.h):
| Scenario | Message | Shows |
|---|---|---|
telemetry_serialize, telemetry_serialize_deserialize |
Telemetry, 7 mixed fields |
the small message every device sends |
telemetry_batch_serialize_deserialize |
TelemetryBatch, 20 nested samples |
the buffered upload |
configuration_serialize_deserialize |
Configuration, oneofs, string, optional |
the complex message |
command_ack_serialize_deserialize |
DeviceCommand and DeviceAck |
fixed cost per call |
firmware_chunk_serialize_deserialize |
FirmwareChunk, 256 resident bytes |
the payload |
firmware_stream_serialize_deserialize |
FirmwareStream, streamed in 32 byte windows |
the RAM story, protobuf libraries only |
schema_growth_01, _05, _20 |
1, 5 or 20 Growth_NN types round tripped |
flash per message type in the schema |
The other scenarios (simple_types_*, string_*, oneof_*, optional_*, repeated_*,
complete_demo_*, *_callback_*) are the diagnostic set for the Embedded Proto developers.
They probe single features and encoder edge cases and are not meant for the website.
All C++ is compiled with -Os -ffunction-sections -fdata-sections -fno-exceptions -fno-rtti -fno-threadsafe-statics -fno-use-cxa-atexit and linked with --gc-sections, the flags an
embedded IDE sets for a C++ project. Without -fno-exceptions -fno-rtti GCC links the exception
unwinder and RTTI into every C++ build, about 4.5 kB that nanopb (plain C) never pays. See
CMakeLists.txt, the flags are stored with every result in the database.
text,dataandbssfromarm-none-eabi-size, compare against theBaseline/blinkyscenario to get the cost of a library.cpu_cyclesfrom the DWT cycle counter aroundrun_scenario(), read over SWD.n_bytes_in_buffer, the number of bytes serialized. Libraries of the same format must agree on it, Embedded Proto and nanopb do byte for byte; the other formats produce their own sizes.message_size_bytes, the RAM the scenario's message object takes. A scenario reports each message it creates withreport_message_size()and the largest is kept. For Embedded Proto, nanopb and zcbor that issizeof()of the message. ArduinoJson keeps its data on the heap, so its documents use a counting allocator and report the peak heap use plussizeof(). FlatBuffers has no message object, a buffer is read in place, so it reports nothing here.peak_ram_bytes, the RAM the scenario takes at its worst moment: the stack it uses belowmain()plus the heap it makes newlib grow. Before the scenario runs the firmware paints the free RAM with a pattern, afterwards it finds the deepest word the pattern is gone from. This includes the message objects, the buffers, the flatcc builder and the library's own stack use. SysTick is suspended while the scenario runs so its interrupt frames stay out of it.
- Embedded Proto (
External/EmbeddedProto), header only C++. - nanopb (
External/Nanopb), C, built withPB_NO_ERRMSG. - FlatBuffers through flatcc (
External/flatcc), the C implementation. Its compiler is built for the host once intobuild/flatcc-host, its runtime (builder, emitter, verifier, no JSON) is compiled for the target. The builder allocates its stacks and pages from the heap, which shows up inpeak_ram_bytes. A FlatBuffer is read in place and has no message object, so FlatBuffers reports nomessage_size_bytes. - zcbor (
External/zcbor), Nordic's CBOR library, C code generated from CDDL schemas. The generator is installed from the submodule intoExternal/zcbor/venvon first use. The schemas keep protobuf's numbered, optional fields as map entries keyed by field number. - ArduinoJson (
External/ArduinoJson), JSON, not protobuf, the alternative people consider.
The .proto files are shared by all libraries. Options which are specific to one library live
next to them in Protofiles/:
<name>.optionsfor nanopb (max_count,max_size,FT_CALLBACK); nanopb is built withPB_NO_ERRMSG, the one size option its documentation recommends for a small target<name>.options.jsonfor Embedded Proto (maxLength,callbackStorage), passed to the generator as--eams_opt=options_file=...<name>.fbsfor FlatBuffers, a hand written schema equivalent to the.proto. The flatcc builder emits into pages on the heap; each scenario setsFLATCC_EMITTER_PAGE_SIZEto its buffer size, rounded up to a multiple of 128, and the CMake builds the runtime with it. That is the counterpart of the other libraries sizing their buffers per scenario, and what flatcc's documentation advises for constrained devices.- ArduinoJson has no schema. Each scenario sets
ARDUINOJSON_POOL_CAPACITYto the number of slots its largest document needs, rounded up to eight, so a document takes one pool sized to it instead of the default pool of 128 slots. Its documentation describes this setting for small targets. The library is header only, so the define in the scenario file is enough. <name>.cddlfor zcbor, the same, all generated together intoGenerated/cbor_messages_*
- Clean previous build results:
rm -rf build/ - Create build scripts:
cmake -B build -DLIBRARY=TARGET_LIBRARY -DSCENARIO=TARGET_SCENARIO - Build the code:
cmake --build build
Use benchmark.py for build + measurement + database storage.
Example (single scenario):
python3 benchmark.py --library EmbeddedProto --scenario simple_types_serialize_maxBenchmark mode selection:
python3 benchmark.py --library EmbeddedProto --scenario simple_types_serialize_max --mode normal
python3 benchmark.py --library EmbeddedProto --scenario simple_types_serialize_max --mode partial--mode normaluses regular serialization.--mode partialenables partial serialization for EmbeddedProto scenarios.- The selected mode is stored in
benchmark_results.dbin columnbenchmark_mode.
Show latest EmbeddedProto results with mode:
sqlite3 benchmark_results.db "SELECT timestamp, library, scenario, benchmark_mode, text_bytes, data_bytes, bss_bytes, cpu_cycles FROM benchmark_results WHERE library='EmbeddedProto' ORDER BY timestamp DESC LIMIT 10;"Filter a specific scenario by mode:
sqlite3 benchmark_results.db "SELECT timestamp, scenario, benchmark_mode, text_bytes, data_bytes, bss_bytes FROM benchmark_results WHERE library='EmbeddedProto' AND scenario='simple_types_serialize_max' AND benchmark_mode='partial' ORDER BY timestamp DESC LIMIT 5;"Run ./flash.sh
LIBRARYselects which test set is built (see./Externaland./Scenarios).SCENARIOselects which file is built from./Scenarios/<LIBRARY>/<SCENARIO>.cpp.- Arm GNU EABI tools are used to extract flash and static RAM usage.
- OpenOCD is used to flash a Nucleo-F446RE and read dynamic metrics (CPU cycles and buffer usage).
- The green LED LD2 blinks (100 ms every 2 s, driven by TIM2 in hardware so the cycle count is unaffected) while the board runs a scenario or holds results not yet read. After reading the results the tool resets the peripherals and halts the core, so a dark LED means the tool is done with the board.
- All results are stored in a SQLite database.