Skip to content

doc(server): describe Server crash files and HStore image stdout logs - #509

Open
byteayan wants to merge 13 commits into
apache:masterfrom
byteayan:doc/server-crash-diagnostics
Open

byteayan wants to merge 13 commits into
apache:masterfrom
byteayan:doc/server-crash-diagnostics

Conversation

@byteayan

@byteayan byteayan commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the PR

Paired docs for apache/hugegraph#3258, which needs a website update before it merges. That PR makes hugegraph-server.sh always pass -XX:+HeapDumpOnOutOfMemoryError, -XX:HeapDumpPath and -XX:ErrorFile pointing at logs/, including when JAVA_OPTIONS is set, and sets STDOUT_MODE=true in the HStore hugegraph/server image.

Changes

  • quickstart/hugegraph/hugegraph-server.md, section 5.1.5: a paragraph after the startup options table. It says bin/hugegraph-server.sh puts the heap dump and crash log defaults first in JAVA_TOOL_OPTIONS; gives the file names (heapdump_<host>_<launch time>/java_pid<pid>.hprof, hs_err_pid<pid>_<host>_<launch time>.log) and the counter; says the server still starts when the directory cannot be created; and explains why directories are never deleted and that the newest one of a running server must not be moved. It lists what overrides the defaults (JAVA_TOOL_OPTIONS, JDK_JAVA_OPTIONS, the command line, _JAVA_OPTIONS) and that only the environment variables reach the JVMs the server starts, so JAVA_TOOL_OPTIONS is the opt-out that covers them. The version note says 1.7.0 wrote dumps straight to logs/java_pid<pid>.hprof.
  • guides/hugegraph-docker-cluster.md, "Runtime logs": images built from master (PD, Store and both Server images) log to stdout; what stays in files; where dumps go; what docker restart, a Compose recreate and docker compose down keep (both images declare VOLUME /hugegraph-server); one log volume per pod on Kubernetes; the Helm chart has no logs volume yet; sizing (ephemeral-storage, sizeLimit, no medium: Memory); and setting JAVA_TOOL_OPTIONS in the container's environment, which the shipped Compose files do not pass through. The version note covers every 1.7.0 and older image.

English and Chinese pages are updated together. The new Chinese lines wrap only at existing spaces, so no soft break renders as a space between Chinese characters.

bash dist/validate-links.sh passes. I did not run a local Hugo build, since the change is prose in four existing pages.

Merge after apache/hugegraph#3258, because the pages describe master behaviour from that PR.

apache/hugegraph#3258 makes the Server launcher always write heap dumps
and JVM crash logs to logs/, even when JAVA_OPTIONS is set, and sets
STDOUT_MODE=true in the HStore hugegraph/server image.

- Server quickstart: after the startup options, describe the heap dump
  and crash log files, how to turn dumps off, which settings override
  the defaults, and what 1.7.0 does instead.
- Docker cluster guide: the runtime-logs note said the HStore Server
  image does not log to stdout. Update it, say which output stays in
  files, and add the /hugegraph-server/logs mount for Kubernetes. The
  1.7.0 behaviour moves to a version-scope note.

English and Chinese pages are updated together.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The crash-file and log claims match apache/hugegraph#3258, but the new 1.7.0 version note on the Docker page covers only the HStore Server, while the standalone hugegraph/hugegraph:1.7.0 image this page pins also logs only to files. One minor CN line-wrap nit. Evidence: git grep STDOUT_MODE 1.7.0 in apache/hugegraph is empty; STDOUT_MODE for PD/Store/Server came in c0a2b93e4 (#2980), which no tag contains; 1.7.0 hugegraph-server.sh:186-188 redirects all output to logs/hugegraph-server.log; other claims checked against #3258 head ac9d98b.

Comment thread content/en/docs/guides/hugegraph-docker-cluster.md Outdated
Comment thread content/cn/docs/guides/hugegraph-docker-cluster.md Outdated

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The Docker guide still omits the legacy standalone-image log behavior; the crash-file paragraph also omits _JAVA_OPTIONS as an override. Evidence: the 1.7.0 launcher redirects Java output to a file, and apache/hugegraph#3258 processes _JAVA_OPTIONS after the new defaults.

Comment thread content/en/docs/quickstart/hugegraph/hugegraph-server.md Outdated
…iles

Review follow-ups:

- Docker cluster guide: no 1.7.0 image sets STDOUT_MODE, the standalone
  hugegraph/hugegraph:1.7.0 used on the page included. Describe stdout
  logging as current master behaviour for all images, and widen the
  version note to every 1.7.0 and older image, linking #2980 and #3258.
- Server quickstart: the defaults now go first in JAVA_TOOL_OPTIONS
  (apache/hugegraph#3258), so list every source that overrides them,
  _JAVA_OPTIONS included, describe the counter added to a taken file
  name, and mention the "Picked up JAVA_TOOL_OPTIONS" startup line.
- Chinese pages: wrap only at existing spaces, so a soft break no longer
  renders as a space between two Chinese characters.
Every launch writes its heap dump to a new file
(apache/hugegraph#3258), so a server that keeps running out of memory
adds a heap-sized dump per restart until the volume is full. Say that
nothing removes old dumps, and how to size the volume, clean up, or
turn dumps off, in the Docker guide and the Server quickstart, English
and Chinese.
apache/hugegraph#3258 now puts the host name, the pod name on
Kubernetes, into heap dump and crash log names so servers sharing one
log directory do not pick the same name. Update the Server quickstart
in English and Chinese to match.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: no. Summary: The crash-file, override-order, stdout-logging and version-scope claims all match apache/hugegraph#3258 (head 8690ecf4e) and 1.7.0, and the earlier findings are fixed at this head. One minor gap: the Docker page says to turn heap dumps off with the flag but not where a container operator sets it. Score 8/10. Evidence: read the full exact-head diff (EN and CN, both pages); checked hugegraph-server.sh, log4j2.xml, Dockerfile-hstore, docker-entrypoint.sh and docker/README.md at #3258 head; git grep STDOUT_MODE 1.7.0 is empty, 1.7.0 hugegraph-server.sh:85 does cd "${TOP}" and :108 adds the heap dump flags only when JAVA_OPTIONS is empty; 1.7.0 PD and Store start scripts redirect to ${OUTPUT}; c0a2b93e4 (#2980) is on master and in no tag. Latest-head checks pass.

Comment thread content/en/docs/guides/hugegraph-docker-cluster.md Outdated
The Docker entrypoint passes JAVA_OPTS to the launcher with -j, so the
quickstart's -j route does not apply in a container. Recommend
JAVA_TOOL_OPTIONS for -XX:-HeapDumpOnOutOfMemoryError, which keeps the
image's default JAVA_OPTS, and note that setting JAVA_OPTS replaces
that default. English and Chinese.
The ZooKeeper and SOFA loggers are set to WARN, so their INFO events are
discarded, not kept in files. Limit the file-only INFO statement in the
Docker guide to the Hadoop, Netty and Commons loggers, in English and
Chinese.
apache/hugegraph#3258 now points HeapDumpPath at a directory per
launch, logs/heapdump_<host>_<launch time>/, so the Server and any JVM
it starts (computer jobs inherit JAVA_TOOL_OPTIONS) each write their
own java_pid<pid>.hprof. Update the names and cleanup advice, and note
that startup removes only this host's empty heapdump_* directories, in
the Server quickstart and the Docker guide, English and Chinese.
apache/hugegraph#3258 now removes a heapdump_* directory at startup
only when it holds no dump and the JVM recorded as its owner has
exited, so a running Server's empty directory is kept. Update the
Server quickstart and the Docker guide, English and Chinese.
apache/hugegraph#3258 dropped the startup cleanup of heapdump_*
directories, since a computer-job JVM can keep using one after its
Server exits. Say that each start leaves a directory, empty unless
something ran out of memory, and that old ones can be deleted once no
JVM from that launch is running. English and Chinese.
Update the Server quickstart and the Docker guide, English and Chinese,
to match apache/hugegraph#3258:
- -j or JAVA_OPTIONS turns heap dumps off for the server JVM only;
  JAVA_TOOL_OPTIONS reaches the JVMs it starts.
- Never move the newest heapdump_* directory of a running server.
- The directory is created even when dumps are off, and the server
  starts anyway when it cannot be created.
- One log volume per pod, ephemeral-storage and sizeLimit, no
  medium: Memory, and the Helm chart's missing logs volume.
- The 1.7.0 note says dumps used to go straight to
  logs/java_pid<pid>.hprof.
- Name the script, give the launch time format, fix the garbled
  sentence about why directories are kept, and polish the Chinese.
Exporting JAVA_TOOL_OPTIONS in the host shell has no effect: it has to
be set with docker run -e, a Compose environment: entry (the shipped
Compose files do not pass it through) or Kubernetes env. English and
Chinese Docker guide.
Both images declare VOLUME /hugegraph-server, so docker restart and a
Compose recreate keep logs/, and docker compose down leaves an
unattached anonymous volume until down -v. Recommend a named volume or
bind mount instead of saying a recreate loses the files. English and
Chinese Docker guide.
@byteayan
byteayan requested a review from bitflicker64 October 7, 2026 18:48
- A per-pod directory on a shared PVC needs POD_NAME passed in from
  metadata.name through the Downward API before subPathExpr:
  $(POD_NAME) works; say so, and that the Helm chart sets neither.
- Call the fallback to logs/ best effort: a full disk may have no room
  for the dump either.
English and Chinese.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] HStore Server image (hugegraph/server) keeps crash diagnostics only inside the container, so a restart on Kubernetes loses them

3 participants