Skip to content

Relax the Geant4 field epsilons outside the muon spectrometer and support local field parameters - #68

Closed
sawenzel wants to merge 28 commits into
g4-field-params-basefrom
g4-field-params
Closed

sawenzel wants to merge 28 commits into
g4-field-params-basefrom
g4-field-params

Conversation

@sawenzel

@sawenzel sawenzel commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

This PR relaxes the minimum epsilon of the Geant4 field integration, keeps the tight value in the muon spectrometer through local fields, and adds the O2 support that local field parameters need. It reduces the transport CPU time in pp by 10%, while the tracking accuracy stays unchanged.

The tight settings in g4config.in go back to the ALIROOT-7121 Geant4 validation of 2018. There, deltaIntersection and both epsilons were tightened together to remove a 1/pT bias of the TrackRef refit. With Geant4 11.2 the bias is still present at the default deltaIntersection, but only deltaIntersection controls it. The epsilons do not affect it. Geant4 uses for each step ε = deltaOneStep / L, clamped to [ε_min, ε_max] (G4PropagatorInField::ComputeStep), so ε_max governs short steps and ε_min governs long steps. The long steps in the air of the muon spectrometer need the tight ε_min; nothing else does.

Field-integration parameters

The bias was measured as in ALIROOT-7121. Single 100 GeV tracks were simulated in a uniform 0.5 T field, and the ITS and TPC TrackRefs were refitted with a circle. At deltaIntersection = 1e-3 mm (Geant4 default), the ITS+TPC 1/pT is shifted by −0.3% for μ± and by −0.68% for charged geantinos. At 1e-5 mm the shift is compatible with zero, and it is identical for every ε_max in {1e-5, 1e-4, 1e-3}.

1/pT bias vs deltaIntersection

In the barrel, the epsilons have no measurable effect. Geantinos between 0.3 and 100 GeV were compared track by track with an ultra-tight uncached reference in the real field map, and with the exact helix in a uniform field. With the settings of this PR, and even with ε_min = ε_max = 1e-3, all TrackRef positions agree with the current ones (median deviation from the exact helix 0.02–0.16 µm in the ITS). Full-physics μ± at 1, 10 and 100 GeV give results identical to the current settings.

In the muon arm, the bias was measured as its counterpart there: the displacement at the last MCH station along the dipole deflection, relative to that deflection, which corresponds to the relative momentum shift. A relaxed ε_min moves the positions by about 45 µm (median, against 6–9 µm now) and shifts the momentum scale by up to −120 ppm. The volumes YOUT1, DDIP and YOUT2 therefore keep the tight epsilons as local fields. With ε_min = 1e-4 elsewhere, the median momentum-scale shift is below 35 ppm and its rms about 100 ppm (currently 45 ppm), negligible against the momentum resolution of the muon spectrometer. Relaxing ε_max alone changes nothing, neither in accuracy nor in CPU time.

Muon-arm momentum-scale shift per configuration

Local field parameters

Geant4 VMC builds a local field only for a TGeo volume that carries a field. O2 attaches the field only globally, so the per-volume parameters created with /mcDet/createMagFieldParameters had no effect. The Geant4 run configuration is moved from FastSim to Detectors/gconfig as o2::g4config::G4RunConfiguration; fast simulation now provides its parts through two factory functions, with unchanged behaviour. The new G4LocalFieldConstruction attaches the global field to every volume that has its own parameter directory /mcMagField/<vol>/ and switches on local fields in Geant4 VMC. The configuration therefore lives entirely in g4config.in, and the step does nothing when no volume has local parameters.

Geant4 VMC forces a local field onto all daughters of a volume, which would switch the field on inside zero-field media. A volume whose subtree contains a medium with ifield = 0 is therefore refused with a warning. Local fields are built only with TGeo navigation; with G4.navmode=kG4 a warning is issued, and the global field applies everywhere.

Where the field is evaluated

In pp, 74% of all evaluations of o2::field::MagneticField::Field occur on the A side at z = 4–8.5 m and small radius, in the fringe field of the solenoid, where no detector needs tracking precision. About 20% occur in the central barrel and 3% in the muon arm. The relaxed ε_min removes 26% of all evaluations, mostly on the A side.

Field evaluations in (z, r)

Performance and validation

configuration CPU s/event
current 6.8–6.9
this PR 6.1–6.2

The timings come from pp at 13.6 TeV (512 events, -j 32, EPYC 7552), with the runs alternated with the baseline on an otherwise idle node. For the physics check, 3000 pp events per configuration were simulated in a uniform field with two seeds, and compared event by event against the current settings. A null configuration, which changes only the random sequence, sets the scale for the fluctuations. The hit counts in all detectors and the TrackRef 1/pT refit of charged primaries stay within that spread. The change was tested as an overlay of libO2FastSim and libO2G4Setup on MC-prod-2026-v12-1; pp events are reproducible run to run with local fields.

Hits per event vs current settings

TrackRef 1/pT refit in pp

Further gains are followed up separately: the DormandPrince745 stepper and ε_min = 1e-3 would bring −15%, but they displace a small fraction of muon-arm tracks that cross the edge of the dipole field map, which drops from 7 kG to zero at a polar angle of 9.7°.

Overall, the field-integration settings now keep only the tightening that the ALIROOT-7121 bias requires, plus a tight ε_min where the muon arm needs it, and O2 can set field parameters per volume.

https://its.cern.ch/jira/browse/O2-7198
https://alice.its.cern.ch/jira/browse/ALIROOT-7121

Assisted by Claude Code.

sawenzel and others added 25 commits September 24, 2026 22:41
…roup#15838)

* Write tracked V0s, cascades and 3-bodies in collision order

This fixes the row order of the tracked strangeness tables written by the AOD producer.

- The tracked V0, cascade and 3-body rows were written in strangeness-tracker order, which is not always collision order.
- Analyses slicing them by collision then abort with "TraCascIndices index fIndexCollisions is not sorted".
- The rows are now written in the per-collision order already built in prepareStrangenessTracking.
- The track index of a row no longer comes from a running counter, so a skipped strange track cannot shift the rows after it.

https://its.cern.ch/jira/browse/O2-7197


* Keep strange tracks in decay order also with one vertexer thread

This makes the order of the tracked strangeness tables within a collision the same in MC as in data.

- SVertexer sorted the strange tracks by decay only when running with more than one thread; MC reconstruction runs it with one.
- The strange tracks are now sorted by decay in all cases.
- The AOD producer groups them by collision with a stable sort, so the decay order within a collision is kept.

https://its.cern.ch/jira/browse/O2-7197

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This adds the RD50 damage weights in the text format the Geant4 fluence weighting reads.

- G4.fluenceWeightFile accepts CSV only, while Detectors/gconfig/data so far held the weights as ROOT graphs.
- The file is dumped from rd50_niel.root and carries the same values.
- The existing o2_data_file rule installs it to share/Detectors/gconfig/data/.
MSL rejects the noexcept specifier outright, with a diagnostic of its own:
"'noexcept' is not supported in Metal". It applies to free functions, to
function templates and to out-of-line member definitions alike; only a noexcept
inside a class body slips through, which does not help a header that defines its
members out of line.
This fixes a problem in the MFT support geometry, where a volume and its assembly share a name.

- Volumes with the same name share a ROOT volume number, so consumers keyed on that number see only one of the two.
- In FLUKA the pair collapsed to the assembly's placeholder material, which means BLCKHOLE, so the PEEK support disks deleted every particle entering them.
GPUCommonAlgorithm::sortOnDevice takes an auto parameter, which is C++20. It is already skipped for OpenCL, at C++17, and MSL 4.1 reports C++17 as well.

GPUTPCTrackParam::TransportToXAlpha declares its material constants static at function scope, which MSL rejects; constexpr without static is accepted, as in SMatrixGPU.

Guard SMatrixGPU C++20 code using __cplusplus version macro.
Rather than trying to adapt Ort::Float16_t to work on Metal, we
simply alias it to Metal native type, which is bit-to-bit equivalent to the CUDA half
implementation and differs from the software one only in the sign of NaN.
This fixes the -Winconsistent-missing-override warnings in InteractionSampler.h.

- FixedSkipBC_InteractionSampler and NonUniformMuInteractionSampler derive from InteractionSampler, which already has a ClassDef.
- Their plain ClassDef produced override warnings for IsA, ShowMembers, Streamer and CheckTObjectHashConsistency.
- Both now use ClassDefOverride.

https://ali-ci.cern.ch/alice-build-logs/AliceO2Group/AliceO2/15843/ee5503a4c432ac109fab4860dde02f104c65e5a6/build_O2_gpu-test-slc10-x86/pretty.html
This fixes the rootcling "missing ; at end of rule" errors in four LinkDef files.

- Three rules in TPCBaseLinkDef.h and the GeometryTGeo rules of ALICE3 ECal, RICH and MID lacked the closing semicolon.
- rootcling rejected the three ALICE3 rules, so their GeometryTGeo classes had no dictionary.
- The TPC dictionary is unchanged by the fix.

https://ali-ci.cern.ch/alice-build-logs/AliceO2Group/AliceO2/15843/ee5503a4c432ac109fab4860dde02f104c65e5a6/build_O2_gpu-test-slc10-x86/pretty.html
Introduce shared tracking, propagation, material handling and runtime ROF tables with compatibility wrappers for legacy ITS callers.

Add ITS and MFT CA workflows with checked configuration, reusable workflow sessions and workflow-owned publication. Include tracking and workflow tests.
This fixes a problem in how the o2-sim driver judges the hit merger's exit code.

- The condition `!= 0 || != 128` is always true, so every normal merger exit was reported as an error.
- It now reads `!= 0 && != 128`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This fixes the generator cache of the o2-sim primary server so that it skips external kinematics, as its comment intends.

- The condition `!= "extkin" || != "extkinO2"` is always true, so extkin generators were cached too.
- A service-mode reconfiguration with a new kinematics file then reused the old generator and file.
- The condition now uses `&&`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This fixes the info thread of the o2-sim primary server for requests of unexpected size.

- After the error reply the thread still read the request and could send a second reply on the REP socket.
- It now continues to the next request after the error reply.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This fixes how the hit merger advances to the next event while flushing.

- The skip paths advanced the event counter but then carried on with the current event.
- With --noemptyevents, an event without hits was written when the next event was already complete, and that next event was then skipped.
- An event without buffered info dereferenced the end iterator.
- The loop now iterates over flushable event IDs and every skip path is a plain continue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This fixes a memory leak in the hit merger and makes its merge flag thread-safe.

- cleanEvent was empty, so the decoded SubEventInfo objects of every event were never freed.
- The MC tracks and track references of events dropped by --noemptyevents were never freed either.
- cleanEvent now owns the release of all three buffers; the merge functions no longer delete.
- Buffer entries are emptied rather than erased, since erasing from a tbb::concurrent_unordered_map is not safe while the receiving thread inserts.
- handleSimData keeps the event id and event count as values, since a merge thread may free the event info.
- mergingInProgress is shared by two threads and is now std::atomic<bool>.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This fixes the hit merger when it runs without disc output.

- The kinematics and MC-header file and tree pointers were never initialised.
- With --noDiscOutput they stayed garbage, and the merge tested them as if they were valid trees.
- They are now initialised to nullptr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This fixes the order in which a parallel o2-sim worker finalizes and sends the hits of an event.

- Detectors ran FinishEvent after SendData, while serial o2-sim (FairMCApplication) runs it before filling the output.
- TRD sorts its hits in FinishEvent, and PHOS and CPV sort them and sum duplicates.
- With TMessage transport their hits were written unsorted and, for PHOS and CPV, with duplicates.
- With shared memory the merger read the buffers while the worker rewrote them, so TRD output differed between the two transports.
- FinishEvent now runs for all detectors before SendData, and EndOfEvent after it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This makes the flag that guards a hit buffer between a sim worker and the hit merger safe across processes.

- The flag was a plain bool in shared memory, written by the merger and polled by the worker.
- Without atomic semantics its stores may become visible out of order with the buffer reads, notably on aarch64.
- It is now a std::atomic<bool> (ShmBusyFlag) placed in the segment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This stops a killed o2-sim from leaking its SysV shared-memory segment.

- The segment (1 GB per worker) was only removed in the driver's normal cleanup.
- It is now marked IPC_RMID right after creation; Linux still lets the workers and the merger attach by id.
- The kernel frees it once the last process detaches, also after a crash.
- The segment is now created with mode 0600 instead of 0666.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This makes the hit merger decode each detector's hits the way the worker sent them.

- Worker and merger each read the global ShmManager::isOperational() to choose between shared memory and TMessage.
- That flag lives in the segment and flips when any late worker fails to attach, so hits already in flight could be decoded the wrong way.
- The per-detector header message is now a HitsHeader carrying the DetID and the mode, decided once per detector in attachHits.
- collectHits takes the mode from the header.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This removes a full copy of every hit container that reaches the hit merger as a TMessage.

- collectHits allocated a new container, copied the decoded one into it and deleted the original.
- It now takes ownership of the decoded container.
- Hits in shared memory belong to the worker and are still copied.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This removes a deep copy of every hit when the hit merger joins the sub-events of an event.

- mergeAndAdjustHits copied each hit into the merged container; it now reserves the total size and moves them.
- The track and track-reference merges now reserve their output.
- TPC HitGroup declared a defaulted destructor, which suppresses its implicit move; the line is removed so HitGroup moves its five vectors.
- The printf debug output in the merge loops is dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This runs the per-event merge and flush of the kinematics and of each detector in parallel.

- The hit merger merged and filled the kinematics and all detector trees one after another.
- In PbPb this serial step took 12.9 s of a 66 s run, after all workers had finished.
- Each of these writes to its own TFile, so they now run as tasks of a tbb::task_group.
- The final TFile::Write calls run in parallel as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This moves the implementation of the o2-sim primary server, worker and hit merger devices out of their headers.

- The three device headers defined all member functions and several non-inline free functions.
- Each header now declares its class; the definitions are in O2PrimaryServerDevice.cxx, O2SimDevice.cxx and O2HitMerger.cxx, compiled into their runners.
- querySimConfig moves from O2SimDevice to PrimaryServerState.cxx, so the hit merger no longer includes the worker and macro/o2sim.C.
- The helpers in SimPublishChannelHelper.h are now inline, and PrimStateToString is inline constexpr.
- The unused nested TMessageWrapper of the hit merger and CustomCleanup of the worker are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This replaces the chain of per-detector conditions in the hit merger with a table of factories.

- Each detector that can be merged now has one entry mapping its DetID to a factory.
- The warning compared the number of active detectors with DetID::nDetectors and fired in practically every run.
- It now names an active readout detector that has no merger instance.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This replaces the static per-type map of hit buffers in DetImpl with a member of each detector instance.

- collectHits kept the buffers in a function-local static map keyed by 'this' and passed them on through a char pointer.
- They now live in a type-erased std::shared_ptr<void> member, reached through hitCollector().
- The instance keeps its own buffers, which also covers several external detectors sharing one type.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ktf and others added 3 commits September 25, 2026 15:56
Both of these already exist for OpenCL, for reasons that apply unchanged to
Metal.

The processing settings block in GPUSettingsList.h is skipped for OpenCL because
it declares std::string and std::vector members, which GPUSettings.h explicitly
does not include for device code. Metal needs the same exclusion. These configs
are host-side only: GPUParam carries GPUSettingsRec and GPUSettingsParam, and the
processing settings appear only as pointer arguments to host methods, so nothing
transferred changes shape.

GPUCommonBitSet already carries an extra constructor for OpenCL's __constant.
Metal needs the opposite: MSL will not use a user-declared copy constructor to
build an object in the constant address space, which is where GPUconstexpr()
arrays of bitset live, and leaving the copy constructor implicit makes them
constructible again. That one line accounted for 84 of the remaining
diagnostics, across DetID and GlobalTrackID.

Metal translation unit: 136 errors to 27.
This moves the O2 Geant4 VMC run configuration out of FastSim, so that other features can extend it.

- G4RunConfiguration is now o2::g4config::G4RunConfiguration in Detectors/gconfig, built into G4Setup.
- FastSim provides the fast simulation and its regions through createFastSimulation() and createFastSimRegionConstruction().
- The behaviour is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This relaxes the Geant4 field-integration epsilons, keeps them tight as local fields in the muon spectrometer, and adds the support for local field parameters.

- The tight epsilons from ALIROOT-7121 do not affect the 1/pT bias; only deltaIntersection does, and it stays at 1e-5 mm.
- minimumEpsilon controls the long steps; relaxing it to 1e-4 moves nothing in the barrel, but it would move MCH positions, so YOUT1, DDIP and YOUT2 keep the tight values.
- G4LocalFieldConstruction attaches the global field to every volume given its own parameters with /mcDet/createMagFieldParameters, so that the /mcMagField/<vol>/ settings take effect.
- A volume whose subtree contains zero-field media is refused, since Geant4 VMC forces a local field onto all daughters.
- This saves about 10% of the transport CPU time in pp.

https://its.cern.ch/jira/browse/O2-7198
https://alice.its.cern.ch/jira/browse/ALIROOT-7121

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sawenzel sawenzel closed this Sep 28, 2026
@sawenzel
sawenzel deleted the g4-field-params branch September 28, 2026 15:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants