un-inception/reports/4.2.36/perf.md

4.2 KiB

Performance Report: 4.2.36

Generated: 2026-01-29T02:03:51Z Pipeline: #13341

Summary

Metric Value
Total Tests 860
Passed 625
Failed 235
Pass Rate 72.6%
Languages 43
Avg Duration 1528s
Slowest scheme (1574s)
Fastest erlang (1236s)

API Health

Tracks transient errors encountered during test execution. Tests retry on failures to ensure accurate results.

Metric Value
Health Score 0/100
Total Retries 2122
Rate Limit (429) 0
Server Error (5xx) 2122
Timeout 0
Connection 0
Tests Needing Retries 214

Interpretation:

  • Score 95-100: API is healthy, minimal transient errors
  • Score 80-94: Some API instability, but tests recovered via retry
  • Score < 80: Significant API issues affecting test reliability

Test Duration by Language

The primary performance metric - how long each language takes to run its full test suite (15 tests per language).

Duration by Language

Key observations:

  • SCHEME and PYTHON are outliers at 90+ seconds
  • Most languages cluster between 20-40 seconds
  • Compiled languages (red) tend to be faster than interpreted (blue)
  • ERLANG is the fastest at 1236 seconds

Compiled vs Interpreted

Comparing performance between compiled languages (C, Go, Rust, etc.) and interpreted languages (Python, Ruby, JavaScript, etc.).

Category Comparison

Findings:

  • 20 compiled languages vs 22 interpreted
  • Compiled languages have lower median execution time
  • Interpreted languages show more variance (wider spread)
  • The white diamond marks the mean for each category

Duration Distribution

Histogram showing how test durations are distributed across all 43 languages.

Duration Histogram

Distribution analysis:

  • Most languages complete in 20-35 seconds (the peak)
  • Mean (green dashed) and median (blue dotted) are close together
  • Long tail on the right from slow outliers (scheme, python)

Speed Leaders

Side-by-side comparison of the 10 slowest and 10 fastest languages.

Speed Leaders

Slowest (left): SCHEME, PYTHON, TCL, ELIXIR, R Fastest (right): ERLANG, AWK, POWERSHELL, CSHARP, JAVASCRIPT


Queue vs Execution Time

Scatter plot showing the relationship between CI queue wait time and actual test execution time.

Queue vs Execution

Notes:

  • Queue time is how long the job waited for a runner
  • Most jobs had similar queue times (clustered vertically)
  • Outliers labeled - scheme and python took longest to execute regardless of queue time

Dashboard

Summary dashboard combining key metrics and visualizations.

Dashboard


Raw Data

Per-Language Performance

Language Status Duration
scheme Failed 1574s
python Failed 1574s
tcl Failed 1572s
elixir Failed 1570s
r Failed 1566s
typescript Failed 1564s
v Failed 1562s
objc Failed 1559s
bash Failed 1559s
nim Failed 1557s
julia Failed 1557s
dart Failed 1556s
raku Failed 1549s
go Failed 1549s
commonlisp Failed 1549s
clojure Failed 1549s
lua Failed 1548s
fortran Failed 1547s
crystal Failed 1547s
c Failed 1547s
ruby Failed 1546s
perl Failed 1546s
java Failed 1546s
fsharp Failed 1546s
groovy Failed 1545s
d Failed 1545s
php Failed 1544s
zig Failed 1542s
deno Failed 1541s
cpp Failed 1541s
prolog Failed 1540s
forth Failed 1540s
cobol Failed 1540s
haskell Failed 1539s
dotnet Failed 1539s
rust Failed 1537s
ocaml Failed 1537s
kotlin Failed 1536s
javascript Failed 1536s
csharp Failed 1534s
powershell Failed 1269s
awk Failed 1239s
erlang Failed 1236s

Report generated by UN Inception CI pipeline