Stabilize first-run cudagraph timing and colorize terminal output

- _cudagraph_plugin: warm up autotune/JIT explicitly before graph capture
  (internal 5-iter warmup is too short), tag fallback markers with the
  failing phase; first-run latency no longer jitters run-to-run.
- _device_guard_plugin (new): set GEMS_VENDOR via torch probe before
  importing flag_gems, avoiding its timeout-less nvidia-smi subprocess
  probe that can hang import in fork-broken environments.
- _pretty_report_plugin (new) + _term_style (new): fold inputs identical
  across all result rows into a legend line, color status/plugin
  tags/markers on the live terminal; run.log is ANSI-stripped and keeps
  upstream SUCCESS/column wording for grep compatibility.
- run_pytest.sh: make USE_FLAGTUNE overridable, group all knobs into a
  config section with Chinese comments, add start/end banners.
- README: document the warmup semantics, new plugins/env vars, and the
  A/B rule that both sides must use the same USE_FLAGTUNE.
This commit is contained in:
2026-07-19 19:21:13 +00:00
parent 8ace0d790c
commit 26b071c6e1
13 changed files with 468 additions and 75 deletions
+6 -2
View File
@@ -27,7 +27,9 @@ def _load_yaml(path: str) -> dict:
with open(path, "r") as f:
return _yaml.safe_load(f) or {}
except Exception as exc:
print(f"[bespoke-shape-plugin] failed to load {path}: {exc}", file=sys.stderr)
from _term_style import tag
print(f"{tag('[bespoke-shape-plugin]')} failed to load {path}: {exc}",
file=sys.stderr)
return {}
@@ -146,5 +148,7 @@ def pytest_collection_finish(session):
if cls.__name__ == "CutlassScaledMMBenchmark" and _patch_cutlass(cls):
covered.append("CutlassScaledMMBenchmark(mnk)")
print(f"[bespoke-shape-plugin] yaml-driven inputs for: {', '.join(covered) or 'none'}",
from _term_style import tag
print(f"{tag('[bespoke-shape-plugin]')} yaml-driven inputs for: "
f"{', '.join(covered) or 'none'}",
file=sys.stderr, flush=True)