CPU benchmarks for comparing m68k compilers, built from code that Amiga programs actually run.
Everything works on memory buffers, so the timed region contains only compiler-generated code: no disk, display or OS calls.1
Numbers for the release binaries on a reference target are in RESULTS.md.
| benchmark | source | what it measures |
|---|---|---|
| dhrystone | Dhrystone 2.1, as used by xSysInfo | the classic integer mix |
| backdrop | p96cts's dithered landscape scene | per-pixel integer arithmetic, divides, byte stores |
| lha-pack | LHa for UNIX 1.14i, -lh5- | hash-chain match search, Huffman coding, K&R-era C |
| lha-unpack | LHa for UNIX 1.14i, -lh5- | Huffman decoding, sliding-window copies |
| zlib-deflate | zlib 1.3.2 | deflate's longest_match() |
| zlib-inflate | zlib 1.3.2 | inflate's bit-level state machine |
| png-encode | libpng 1.6.58 + zlib | filter selection byte loops |
| png-decode | libpng 1.6.58 + zlib | filter reconstruction byte loops |
| ftgrays | FreeType 2.12.1 smooth rasterizer | fixed-point curve subdivision, cell sweep, switch-heavy code |
| memcpy-small | constant-size copies up to 128 bytes | by-pieces expansion2 |
| memcpy-large | constant-size copies from 256 to 4096 bytes | an expander taking over from the memcpy call |
| memcpy-var-small | small copies sized at run time | the memcpy call1 |
| memcpy-var-large | large copies sized at run time | memcpy's copy loop |
| memmove-small | overlapping moves up to 128 bytes, both ways | the backwards path a block-move expander gets wrong |
| memmove-large | overlapping moves of 2048 and 4096 bytes, both ways | the same past the by-pieces limit |
| wipeout-tris | render_push_tris() from arczi84's Wipeout port | a game's per-triangle path: float UV scaling, byte clamps, 96-byte struct copies |
The compiler defaults to cc; override CC to compare toolchains,
and BUILD to keep their objects apart:
make CC=gcc-6.5/bin/m68k-amigaos-gcc BUILD=build-gcc6 TARGET=benchwork-gcc6
CPUFLAGS (default -m68020-60) and OPT (default -O2 -fomit-frame-pointer)
are the knobs a comparison usually turns.
vbcc from the same toolchain builds benchwork-vbcc with make vbcc, and
benchwork-vbcc-000/020/040 with make vbcc-release; the flags are
spelled its way:
make vbcc CPUFLAGS=-cpu=68020 OPT=-O2
NOTE: vbcc 0.9i miscompiles two of the benchmarks, so until the compiler
is fixed ftgrays reports run 1 failed with all CPUs and memcpy-small
gives a wrong checksum and scribbles over memory with -cpu=68040.
We emailed both bug reports to the vbcc maintainer.
SAS/C 6.58 builds benchwork-sasc with make sasc, and the release
binaries with make sasc-release; SC=sc-volamos names a wrapper that
runs sc under volamos, with slink reached the same way.
benchwork [-n iterations] [--fastmem] [-l] [name ...]
The FreeType rasterizer keeps its cell pool on the stack, so the program
refuses to start on the shell's default 4 KiB stack: stack 65536 first.
Each benchmark reports its fastest and mean iteration in milliseconds, and a checksum of its output that must be identical across compilers.
The harness is 0BSD (LICENSE); the vendored code keeps its own licenses, listed in third_party/LICENSE.md, and the LHa terms mean binaries must be distributed together with this source.
Footnotes
-
Out-of-line
memcpy,memsetandmemmoveare each compiler's own C library's, so these rows compare the compiler together with its runtime. gcc links the toolchain's default newlib, whose memcpy is assembler; libnix's hands the work to execCopyMem, which an emulator may run for free, so the gcc build does not use-noixemul. ↩ ↩2 -
128 bytes is the size up to which the AmigaOS by-pieces hook expands a copy inline. ↩