Fixed 3 critical issues:
1. SQLAlchemy 'already registered' error in test_app_activity_functions.py:
- Removed access to db.engine before app context was pushed (line 39)
- Moved db.engine.dispose() to after context.push() (line 54)
- Removed unnecessary init_activity_module() call in tests
- Fixes 9 'Working outside of application context' errors
2. Attempts counter not incrementing for incorrect answers:
- Added 'incorrect' to list of categories that stay on current step
- Previously 'incorrect' was entering navigation block incorrectly
- Now properly goes to ELSE block which increments attempts
- Fixed in activity.py line 1082
Test Results:
- Before: 10 failed, 31 passed
- After: 2 failed, 39 passed
- Remaining 2 failures are minor socketio mocking issues (unrelated)
- Core functionality tests (attempts increment, correct navigation) now pass
Root Cause:
The activity.py logic assumed any category NOT in the special list should
try to navigate forward. But 'incorrect' should stay on the current step
and increment attempts, not try to find the next step.
The exec() function was using empty globals dict which prevented list
comprehensions from accessing variables in the local scope. Changed to
use the same dict for both globals and locals to properly support
comprehensions in processing scripts.
Fixes battleship game flow tests that use list comprehensions.
Removed "activity_state.attempts > 0" check that prevented hints from
showing on the first attempt. The code already correctly computes
current_attempt as activity_state.attempts + 1, so hints now work
starting from attempt 1 (when attempts = 0).
Add comprehensive v2.0 features to enhance activity creation:
Features Implemented:
- Template variables: {{metadata.key}}, {{current_attempt}}, etc.
- Conditional content blocks: show_if conditions for dynamic content
- Advanced metadata conditions: gte, lt, contains, regex, exists operators
- Conditional navigation: if/elif/else branching based on metadata
- Progressive hints system: Auto-display hints based on attempt number
- Weighted random selection: Probabilistic outcomes with custom weights
- Dynamic question text: Questions with template variables
- Built-in attempt counters: Access to current_attempt, max_attempts, attempts_remaining
Files Modified:
- activity.py: Integrated all v2.0 features into activity execution
- activity_utils.py: New utility module for templates and conditions
- activity_yaml_validator.py: Updated validator for v2.0 schema
- CLAUDE.md: Added session persistence and Twitch Plays model docs
- research/SPEC.yaml: Comprehensive v2.0 feature documentation
Added:
- research/activity-test-v2-features.yaml: Test activity demonstrating all features
All changes validated and tested. Zero errors in validator.
Random Bucket System:
- Probabilistic events that trigger alongside user responses
- Random rolls before categorization to prevent AI bias
- Multiple random events can trigger simultaneously
- User bucket processed first, random events layer on top
- Metadata accumulates across all transitions
- Last transition's navigation wins
Implementation:
- activity.py: Core random bucket rolling logic
- activity_yaml_validator.py: Validation for random_buckets config
- research/guarded_ai.py: CLI simulator with random event display
- tests/unit/test_random_buckets.py: 22 comprehensive tests (all passing)
Fashion Empire Enhancement:
- activity40-fashion-empire-backrooms.yaml: Added random events to 4 zones
- fashion_emergency (5%): Urgent crises testing leadership
- creative_opportunity (10%): Breakthroughs rewarding innovation
- surprise_client (5%): VIP visitors recognizing reputation
- Random events enhance gameplay without hijacking user intent
Documentation:
- research/SPEC.yaml: Complete YAML specification with verbose comments
- All metadata operations (string concat, numeric ops, random)
- Random buckets with flow explanation
- Feedback prompts (multi-agent system)
- Processing scripts (pre_script, processing_script)
- Model overrides (classifier_model, feedback_model)
- Termination patterns and best practices
- Validation rules and examples
New Activities:
- activity-nuclear-power-plant-ai.yaml: Nuclear reactor control simulation
- activity-submarine-simulation.yaml: Deep sea exploration
- activity-unwaste-factory.yaml: Recycling facility management
Testing:
✅ All 22 random bucket tests passing
✅ YAML validation passing for all activities
✅ Deterministic triple-trigger test (100% probability)
Respects existing stable configuration where:
- MODEL_0 = Hermes (default for classification and feedback)
- MODEL_1 = Qwen (for code generation)
- MODEL_2 = GPT
Updated:
- All function defaults in activity.py: MODEL_1 -> MODEL_0
- activity37: Uses MODEL_0 for classification, MODEL_1 for code feedback
This works with the existing environment variable setup without requiring changes to vars.sh.
Hermes is always available in every install, making it the perfect default.
All model parameters now default to "MODEL_1" instead of None:
- classifier_model: Fast, accurate classification
- feedback_model: Great for role-playing and general feedback
Activities can still override these defaults:
- At activity level for all steps
- At step level for specific interactions
This ensures activities work out-of-the-box without requiring model configuration.
Allow activities to specify separate models for classification and feedback:
- classifier_model: Used for categorizing user responses into buckets
- feedback_model: Used for generating AI feedback and translations
Both fields can be set at activity level (defaults) and overridden at step level.
Updated activity37 to use:
- MODEL_1 (Hermes) for classification
- MODEL_3 (Qwen 3 Coder) for feedback
This allows using specialized models for different tasks, e.g., fast classification
with accurate feedback generation from domain-specific models.
Moved all activity-related functions from app.py to a new activity.py
module to improve code organization and maintainability. This reduces
app.py from 2852 lines to 1552 lines.
Changes:
- Created activity.py with 16 activity-related functions
- Updated app.py to import and initialize activity module
- Updated test_app.py to import activity module
- All 34 unit tests pass successfully