Conversation
Collaborator
Member
Author
|
@davidrohr any comments on this? |
davidrohr
reviewed
Sep 30, 2026
Member
Author
|
Ok, updated and cleaned up. notice the value_t cast is needed because the ternary will not work with two different types. |
Collaborator
Member
Author
|
@davidrohr any further objections? @vkucera I see a bunch of spurious: can you turn off the check, please? We routinely use VLAs for performance reason. |
davidrohr
requested changes
Oct 1, 2026
Metal Shading Language has no double, and the tracking needs one: the track parametrisation, the propagator and the material budget all do double arithmetic. The Metal entry point in dev already includes GPUCommonDoubleBinary64.h; this adds it. The class implements binary64 over a 64-bit integer. Round to nearest even, with subnormals, infinities and NaNs, and NaN propagation in the ARM64 order, so an Apple host is a bit-exact reference down to the payload. Addition, subtraction, multiplication, division and the conversions to and from float and the 32-bit integers are exact; sin and cos are fdlibm's and land within 2 ulp of libm. There is no fused multiply-add and no square root. The entry point aliases the double keyword to the class, so the shared headers go on saying double, and the header refuses to build where a real double exists. Abs<double> needs a Metal specialisation of its own, since fabs would otherwise resolve to metal::fabs(float) through the implicit conversion and round the mantissa away at 17 call sites, with no error and no warning, among them the SMatrixGPU pivot selection. A static_assert keeps it that way, since metal::fabs is not constant-evaluable. SinCosd goes the other way: nothing on the device path needs it, so math_utils::sincosd is excluded for Metal as it already is for OpenCL, and the stale GPUCommonDouble.h include goes with it. The one remaining line is a ternary that needs an explicit value_t to stay unambiguous once value_t is a class. It costs of the order of a hundred times plain float on an M-series GPU, which the tracking can afford because double is a small fraction of its floating point work.
davidrohr
approved these changes
Oct 1, 2026
Collaborator
It's already off. https://github.com/alisw/alidist/blob/dea51bc9ac4b9322d5377d31df8837964610aa69/o2checkcode.sh#L67 Isn't this the misbehaviour investigated by @sawenzel ? |
Collaborator
Member
Author
|
@davidrohr ok, updated. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Metal Shading Language has no double, and the tracking needs one: the track
parametrisation, the propagator and the material budget all do double
arithmetic. The Metal entry point in dev already includes
GPUCommonDoubleBinary64.h; this adds it.
The class implements binary64 over a 64-bit integer. Round to nearest even,
with subnormals, infinities and NaNs, and NaN propagation in the ARM64 order,
so an Apple host is a bit-exact reference down to the payload. Addition,
subtraction, multiplication, division and the conversions to and from float
and the 32-bit integers are exact; sin and cos are fdlibm's and land within
2 ulp of libm. There is no fused multiply-add and no square root. The entry
point aliases the double keyword to the class, so the shared headers go on
saying double, and the header refuses to build where a real double exists.
Abs needs a Metal specialisation of its own, since fabs would
otherwise resolve to metal::fabs(float) through the implicit conversion and
round the mantissa away at 17 call sites, with no error and no warning, among
them the SMatrixGPU pivot selection. A static_assert keeps it that way, since
metal::fabs is not constant-evaluable. SinCosd goes the other way: nothing on
the device path needs it, so math_utils::sincosd is excluded for Metal as it
already is for OpenCL, and the stale GPUCommonDouble.h include goes with it.
The one remaining line is a ternary that needs an explicit value_t to stay
unambiguous once value_t is a class.
It costs of the order of a hundred times plain float on an M-series GPU, which
the tracking can afford because double is a small fraction of its floating
point work.