diff --git a/AGGREGATED-PERFORMANCE.md b/AGGREGATED-PERFORMANCE.md index 358ac8e..22ce630 100644 --- a/AGGREGATED-PERFORMANCE.md +++ b/AGGREGATED-PERFORMANCE.md @@ -1,13 +1,13 @@ # UN Inception: Aggregated Performance Analysis -**Analysis Date:** 1769719662.1650922 -**Reports Analyzed:** 4.2.0, 4.2.10, 4.2.11, 4.2.12, 4.2.13, 4.2.14, 4.2.15, 4.2.16, 4.2.17, 4.2.18, 4.2.19, 4.2.20, 4.2.21, 4.2.22, 4.2.23, 4.2.24, 4.2.25, 4.2.26, 4.2.27, 4.2.28, 4.2.29, 4.2.3, 4.2.30, 4.2.31, 4.2.32, 4.2.36, 4.2.37, 4.2.38, 4.2.4, 4.2.46, 4.2.5, 4.2.6, 4.2.7, 4.2.8, 4.2.9 +**Analysis Date:** 1769732494.6057673 +**Reports Analyzed:** 4.2.0, 4.2.10, 4.2.11, 4.2.12, 4.2.13, 4.2.14, 4.2.15, 4.2.16, 4.2.17, 4.2.18, 4.2.19, 4.2.20, 4.2.21, 4.2.22, 4.2.23, 4.2.24, 4.2.25, 4.2.26, 4.2.27, 4.2.28, 4.2.29, 4.2.3, 4.2.30, 4.2.31, 4.2.32, 4.2.36, 4.2.37, 4.2.38, 4.2.4, 4.2.46, 4.2.5, 4.2.50, 4.2.6, 4.2.7, 4.2.8, 4.2.9 --- ## Executive Summary -Analysis of 35 performance reports reveals **significant variance** in execution metrics across releases. Different languages rank as slowest/fastest in different runs, indicating **non-deterministic execution patterns** likely caused by: +Analysis of 36 performance reports reveals **significant variance** in execution metrics across releases. Different languages rank as slowest/fastest in different runs, indicating **non-deterministic execution patterns** likely caused by: 1. **Orchestrator placement on CPU-bound pool** (not an SRE best practice) 2. **Resource contention** between the orchestrator & test jobs @@ -53,7 +53,8 @@ Analysis of 35 performance reports reveals **significant variance** in execution | 4.2.4 | 70s | python (110s) | c (23s) | -326s (-82.3%) | | 4.2.46 | 224s | raku (344s) | prolog (112s) | +154s (+220.0%) | | 4.2.5 | 67s | v (114s) | erlang (44s) | -157s (-70.1%) | -| 4.2.6 | 54s | haskell (128s) | awk (23s) | -13s (-19.4%) | +| 4.2.50 | 183s | go (480s) | fortran (107s) | +116s (+173.1%) | +| 4.2.6 | 54s | haskell (128s) | awk (23s) | -129s (-70.5%) | | 4.2.7 | 117s | typescript (319s) | dotnet (5s) | +63s (+116.7%) | | 4.2.8 | 111s | kotlin (313s) | fortran (28s) | -6s (-5.1%) | | 4.2.9 | 107s | ruby (279s) | d (19s) | -4s (-3.6%) | @@ -104,6 +105,7 @@ The same language changes dramatically in rank between runs: - 4.2.4: 109s - 4.2.46: 197s - 4.2.5: 60s + - 4.2.50: 241s - 4.2.6: 50s - 4.2.7: 155s - 4.2.8: 253s @@ -142,6 +144,7 @@ The same language changes dramatically in rank between runs: - 4.2.4: 74s - 4.2.46: 199s - 4.2.5: 52s + - 4.2.50: 223s - 4.2.6: 47s - 4.2.7: 313s - 4.2.8: 126s @@ -180,6 +183,7 @@ The same language changes dramatically in rank between runs: - 4.2.4: 100s - 4.2.46: 242s - 4.2.5: 102s + - 4.2.50: 157s - 4.2.6: 42s - 4.2.7: 146s - 4.2.8: 55s @@ -218,6 +222,7 @@ The same language changes dramatically in rank between runs: - 4.2.4: 110s - 4.2.46: 196s - 4.2.5: 61s + - 4.2.50: 119s - 4.2.6: 52s - 4.2.7: 58s - 4.2.8: 253s @@ -256,6 +261,7 @@ The same language changes dramatically in rank between runs: - 4.2.4: 96s - 4.2.46: 117s - 4.2.5: 51s + - 4.2.50: 152s - 4.2.6: 42s - 4.2.7: 148s - 4.2.8: 247s @@ -300,6 +306,7 @@ The same language changes dramatically in rank between runs: 4.2.4: c, d, cobol, raku, v 4.2.46: prolog, typescript, tcl, objc, clojure 4.2.5: erlang, awk, bash, deno, tcl +4.2.50: fortran, csharp, bash, ocaml, python 4.2.6: awk, powershell, crystal, raku, erlang 4.2.7: dotnet, deno, awk, fortran, commonlisp 4.2.8: fortran, groovy, crystal, java, powershell @@ -338,6 +345,7 @@ The same language changes dramatically in rank between runs: 4.2.4: python, javascript, elixir, scheme, bash 4.2.46: raku, powershell, rust, commonlisp, lua 4.2.5: v, haskell, scheme, ocaml, powershell +4.2.50: go, groovy, awk, javascript, erlang 4.2.6: haskell, go, cpp, rust, forth 4.2.7: typescript, ruby, r, elixir, crystal 4.2.8: kotlin, python, javascript, tcl, raku @@ -352,9 +360,9 @@ The same language changes dramatically in rank between runs: ### 4. API Health Trends -**Overall API Health:** 0.0/100 (avg across 4 releases) +**Overall API Health:** 0.0/100 (avg across 5 releases) **Trend:** STABLE -**Total Retries (all releases):** 2641 +**Total Retries (all releases):** 2699 | Release | Health Score | Total Retries | 429 (Rate Limit) | 5xx (Server) | Timeout | Connection | |---------|--------------|---------------|------------------|--------------|---------|------------| @@ -362,6 +370,7 @@ The same language changes dramatically in rank between runs: | 4.2.37 | 0/100 | 151 | 0 | 60 | 0 | 0 | | 4.2.38 | 0/100 | 222 | 0 | 125 | 0 | 0 | | 4.2.46 | 0/100 | 146 | 0 | 146 | 0 | 0 | +| 4.2.50 | 0/100 | 58 | 0 | 58 | 0 | 0 | **Interpretation:** - **Score 95-100:** API healthy, tests pass on first attempt @@ -517,26 +526,26 @@ Keep it as-is for stress testing, but in separate test environment. | Language | Min (s) | Max (s) | Avg (s) | Range (s) | Variance % | |----------|---------|---------|---------|-----------|------------| -| DOTNET | 5 | 1539 | 167.3 | 1534 | 30680.0% | -| R | 9 | 1834 | 225.8 | 1825 | 20277.8% | -| NIM | 8 | 1557 | 171.3 | 1549 | 19362.5% | -| CLOJURE | 8 | 1549 | 173.0 | 1541 | 19262.5% | -| FORTRAN | 8 | 1547 | 137.5 | 1539 | 19237.5% | -| PERL | 8 | 1546 | 171.8 | 1538 | 19225.0% | -| D | 8 | 1545 | 137.3 | 1537 | 19212.5% | -| ZIG | 8 | 1542 | 218.2 | 1534 | 19175.0% | -| FORTH | 8 | 1540 | 172.4 | 1532 | 19150.0% | -| KOTLIN | 8 | 1536 | 148.9 | 1528 | 19100.0% | -| CSHARP | 8 | 1534 | 159.5 | 1526 | 19075.0% | -| LUA | 9 | 1548 | 197.2 | 1539 | 17100.0% | -| PROLOG | 9 | 1540 | 127.1 | 1531 | 17011.1% | -| RUST | 9 | 1537 | 182.2 | 1528 | 16977.8% | -| SCHEME | 15 | 1574 | 166.3 | 1559 | 10393.3% | -| OBJC | 17 | 1559 | 152.6 | 1542 | 9070.6% | -| POWERSHELL | 14 | 1269 | 125.0 | 1255 | 8964.3% | -| PYTHON | 19 | 1574 | 164.6 | 1555 | 8184.2% | -| OCAML | 19 | 1537 | 170.6 | 1518 | 7989.5% | -| TCL | 20 | 1572 | 183.5 | 1552 | 7760.0% | +| DOTNET | 5 | 1539 | 167.0 | 1534 | 30680.0% | +| R | 9 | 1834 | 225.7 | 1825 | 20277.8% | +| NIM | 8 | 1557 | 171.5 | 1549 | 19362.5% | +| CLOJURE | 8 | 1549 | 173.5 | 1541 | 19262.5% | +| FORTRAN | 8 | 1547 | 136.6 | 1539 | 19237.5% | +| PERL | 8 | 1546 | 173.2 | 1538 | 19225.0% | +| D | 8 | 1545 | 137.0 | 1537 | 19212.5% | +| ZIG | 8 | 1542 | 217.9 | 1534 | 19175.0% | +| FORTH | 8 | 1540 | 172.9 | 1532 | 19150.0% | +| KOTLIN | 8 | 1536 | 149.7 | 1528 | 19100.0% | +| CSHARP | 8 | 1534 | 158.2 | 1526 | 19075.0% | +| LUA | 9 | 1548 | 198.1 | 1539 | 17100.0% | +| PROLOG | 9 | 1540 | 126.9 | 1531 | 17011.1% | +| RUST | 9 | 1537 | 183.6 | 1528 | 16977.8% | +| SCHEME | 15 | 1574 | 166.1 | 1559 | 10393.3% | +| OBJC | 17 | 1559 | 152.0 | 1542 | 9070.6% | +| POWERSHELL | 14 | 1269 | 125.1 | 1255 | 8964.3% | +| PYTHON | 19 | 1574 | 163.3 | 1555 | 8184.2% | +| OCAML | 19 | 1537 | 169.1 | 1518 | 7989.5% | +| TCL | 20 | 1572 | 182.7 | 1552 | 7760.0% | --- @@ -622,6 +631,7 @@ Individual Reports → Aggregation Script → Chart Generation (via UN) → Fina - `reports/4.2.4/perf.json` - 682 tests, generated 2026-01-19T12:02:14Z - `reports/4.2.46/perf.json` - 860 tests, generated 2026-01-29T20:46:47Z - `reports/4.2.5/perf.json` - 658 tests, generated 2026-01-19T19:10:23Z +- `reports/4.2.50/perf.json` - 860 tests, generated 2026-01-30T00:20:38Z - `reports/4.2.6/perf.json` - 642 tests, generated 2026-01-19T20:22:16Z - `reports/4.2.7/perf.json` - 631 tests, generated 2026-01-23T09:36:18Z - `reports/4.2.8/perf.json` - 645 tests, generated 2026-01-23T10:01:33Z @@ -822,5 +832,5 @@ For questions about this methodology or to report issues: --- **Generated by UN Inception Performance Analysis Pipeline** -**Analysis Date:** 2026-01-29T15:47:42.288247 +**Analysis Date:** 2026-01-29T19:21:34.732904 **Report Version:** 1.0.0