Commit graph

377 commits

Author SHA1 Message Date
Claude
6fa1b15188
Add auto-growing textarea for chat input
- Textarea now automatically expands as user types multiline messages
- Resets to minimum height after message is sent
- CSS: Set min-height (60px) and max-height (400px) with auto overflow
- Removed fixed rows attribute to allow dynamic height
- Disabled manual resize to prevent user confusion
- Provides better UX for composing longer messages
2025-11-08 23:02:24 +00:00
Claude
0560de004d
Add fix scripts for activity YAML corrections
These scripts document the automated fixes applied to activities 30-37:
- fix_activity37.py: Changes 'close' bucket behavior + fixes completion
- fix_all_new_activities.py: Fixes final step completion for all activities

Keeping for reference and potential reuse on future activities.
2025-11-08 23:00:01 +00:00
Claude
f774d36d74
Fix activity completion and progression issues in activities 30-37
This commit addresses two critical issues:

1. Completion Bug (activities 30-37):
   - Final steps were looping forever, preventing activity completion
   - Fixed by removing next_section_and_step from completion transitions
   - Kept off_topic transition looping to avoid validator terminal step errors
   - Activities now complete properly when users give valid final answers

2. Activity37 Bucket Logic:
   - Changed "close" bucket to retry same step instead of advancing
   - Only "correct" bucket now advances to next step
   - All other buckets (close, incomplete, wrong_language, etc.) retry
   - This ensures students must get correct answers to progress

Technical Details:
- Final steps are not considered "terminal" if at least one transition
  has next_section_and_step (validator requirement)
- Off-topic transitions loop back to allow another attempt
- Completion happens when get_next_step() returns None, None

Validation:
- All 8 activities pass activity_yaml_validator.py
- No errors or warnings

Affects: activity30-37 (all new merged activities)
2025-11-08 22:59:06 +00:00
08bbc47a6b
Merge pull request #20 from russellballestrini/claude/expand-research-yamls-011CUvoNn9xvytg4xr5eJ7Rx
Generate synthetic activities from research YAMLs
2025-11-08 15:07:53 -05:00
Claude
fd78927e7c
Fix MODEL_X references to use dynamic model registry
When MODEL_X references (MODEL_0, MODEL_1, etc.) are used, the code now
properly looks up actual model names from the dynamic registry (MODEL_CLIENT_MAP)
instead of hardcoding "model" or requiring MODEL_NAME_X environment variables.

Changes:
- app.py: Look up models from MODEL_CLIENT_MAP for the specified endpoint
- guarded_ai.py: Query endpoints for actual model names at initialization
- guarded_ai.py: Use dynamic registry for MODEL_X lookups

This fixes the "model not found" error when using activities with MODEL_X
references like activity37.
2025-11-08 19:58:21 +00:00
Claude
3c3b8bd493
Change default model from MODEL_1 to MODEL_0 to match stable config
Respects existing stable configuration where:
- MODEL_0 = Hermes (default for classification and feedback)
- MODEL_1 = Qwen (for code generation)
- MODEL_2 = GPT

Updated:
- All function defaults in activity.py: MODEL_1 -> MODEL_0
- activity37: Uses MODEL_0 for classification, MODEL_1 for code feedback

This works with the existing environment variable setup without requiring changes to vars.sh.
2025-11-08 19:53:46 +00:00
Claude
8cebcbf118
Add MODEL_X reference support to app.py
Critical fix for activity model configuration:
- Handle MODEL_1, MODEL_2, MODEL_3 references in get_openai_client_and_model()
- Look up MODEL_ENDPOINT_{n}, MODEL_API_KEY_{n}, MODEL_NAME_{n} from environment
- Fall back gracefully to default model if MODEL_{n} not configured
- Matches implementation in research/guarded_ai.py

Fixes error: 'NoneType' object has no attribute 'chat'
This error occurred when activities tried to use classifier_model="MODEL_1"
but the app didn't know how to resolve the MODEL_X reference.

Now activity37 (programming languages) will work correctly with:
- classifier_model: "MODEL_1" (Hermes for classification)
- feedback_model: "MODEL_3" (Qwen3-Coder for code generation)
2025-11-08 19:47:32 +00:00
Claude
0f06772afb
Fix critical model name issue and validator warning
Critical fix for guarded_ai.py:
- Add MODEL_NAME_{n} environment variable support
- Fixes hard-coded "model" string that breaks Azure OpenAI and other endpoints
- Falls back to "model" if MODEL_NAME_{n} not specified
- Some endpoints require actual deployment name in model parameter

Validator improvement:
- Allow feedback_prompts as alternative to feedback_tokens_for_ai
- Prevents false warning when using metadata_feedback_filter with new prompt system

Documentation:
- Added MODEL_NAME_{n} examples to CLAUDE.md
- Documented that Azure and similar endpoints need this variable

All 8 activities validated: 0 errors, 0 warnings
2025-11-08 19:38:15 +00:00
Claude
82aeeab094
Document classifier_model and feedback_model in CLAUDE.md
Added comprehensive Activity YAML Schema section covering:
- Model Configuration feature (classifier_model and feedback_model)
- Why separate models (speed, quality, cost, flexibility)
- Model defaults (MODEL_1/Hermes as universal default)
- Recommended model combinations table
- Environment variable configuration
- Example programming activity with dual models
- Activity YAML validation instructions
- CLI testing with model configuration
- Qwen3-Coder-30B setup guide (llama.cpp and ollama)

This documents the new dual-model architecture that allows:
- Fast classification with Hermes (8B)
- Specialized feedback with domain models (e.g., Qwen3-Coder 30B)
- Activity and step-level model overrides
2025-11-08 19:34:00 +00:00
Claude
1c5a4960fd
Update NEW_ACTIVITIES_PLAN.md with completion status
Transformed planning document into comprehensive completion report:
- Status: 8 activities completed (30-37), 6,112 lines of YAML
- Documented new classifier_model and feedback_model feature
- Added model setup guide for Qwen3-Coder-30B
- Detailed activity summaries with special features
- Technical architecture and implementation decisions
- Usage examples and future enhancements

Key highlights:
- All activities validated with 0 errors
- Dual-model architecture explained
- Activity 37 flagship feature: universal programming language support
- Hermes excellence in role-playing scenarios
2025-11-08 19:32:00 +00:00
Claude
f87824bc56
Add Qwen3-Coder-30B setup documentation to activity37
Added detailed comments showing how to use the recommended model:
- hf.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF:Q4_K_M
- Setup instructions for llama.cpp (with GPU offloading)
- Alternative setup with ollama
- Environment variable configuration examples

This 30B parameter model is specifically optimized for code generation
across all programming languages, making it perfect for the universal
programming activity.
2025-11-08 19:27:27 +00:00
Claude
c51b2c7900
Update guarded_ai.py to support classifier_model and feedback_model
Changes:
- Enhanced get_openai_client_and_model() to support MODEL_X references
- Added model parameter (default "MODEL_1") to all AI functions:
  - categorize_response()
  - generate_ai_feedback()
  - provide_feedback()
  - provide_feedback_prompts()
  - translate_text()
- Updated simulate_activity() to:
  - Read classifier_model and feedback_model from YAML
  - Support step-level model overrides
  - Pass appropriate models to classifier vs feedback functions

This ensures the CLI simulation tool matches the production activity.py behavior.
2025-11-08 19:23:29 +00:00
Claude
e3c1547a0f
Set Hermes (MODEL_1) as default for all model parameters
Hermes is always available in every install, making it the perfect default.
All model parameters now default to "MODEL_1" instead of None:
- classifier_model: Fast, accurate classification
- feedback_model: Great for role-playing and general feedback

Activities can still override these defaults:
- At activity level for all steps
- At step level for specific interactions

This ensures activities work out-of-the-box without requiring model configuration.
2025-11-08 19:08:23 +00:00
Claude
1c347ea060
Add classifier_model and feedback_model support to YAML schema
Allow activities to specify separate models for classification and feedback:
- classifier_model: Used for categorizing user responses into buckets
- feedback_model: Used for generating AI feedback and translations

Both fields can be set at activity level (defaults) and overridden at step level.

Updated activity37 to use:
- MODEL_1 (Hermes) for classification
- MODEL_3 (Qwen 3 Coder) for feedback

This allows using specialized models for different tasks, e.g., fast classification
with accurate feedback generation from domain-specific models.
2025-11-08 18:58:29 +00:00
Claude
11e705be97
Add 3 extensive educational activities (American History, Biblical History, Programming)
Created 3 comprehensive educational activities without embedded Python:

1. activity35-american-history.yaml - Advanced American History for gifted students
   - Founding principles and Constitutional design
   - Civil War causes and Reconstruction failure
   - Civil Rights Movement strategies
   - Primary source analysis and critical historical thinking
   - Connects past to present issues

2. activity36-biblical-history.yaml - Biblical History & Ancient Near East
   - Ancient Near Eastern context (Mesopotamia, Egypt, Canaan)
   - Archaeological evidence and historical reconstruction
   - Israelite history (Exodus, Monarchy, Exile)
   - Roman period and early Christianity
   - Foundation myths vs historical facts
   - Cultural adaptation and religious transformation

3. activity37-programming-languages.yaml - Universal Programming Concepts
   - Student chooses ANY programming language (Python, C++, COBOL, anything)
   - AI adapts all examples/feedback to chosen language via metadata
   - Covers: stdout/output, variables, data types, control flow, loops, functions
   - All examples use stdout to display messages
   - Concepts applicable to every language
   - Language-specific syntax provided by AI

All activities:
- Use only YAML features (no embedded Python)
- Validate successfully with 0 errors
- Provide sophisticated educational content
- Use AI feedback for personalization
- Include critical thinking and reflection
- Track progress via metadata

Total: 8 new educational activities across 2 commits (5 from previous commit + 3 now)
2025-11-08 18:47:31 +00:00
1a9be3c420
Merge pull request #19 from russellballestrini/claude/fix-dark-mode-scrollbars-011CUvsF4XTDVgET5n5Prtjd
Fix scrollbars appearance in dark mode
2025-11-08 13:44:47 -05:00
Claude
31152c295b
Refactor CSS from inline to external stylesheet
Move all CSS from base.html to static/css/style.css for better:
- Separation of concerns
- Browser caching
- Maintainability
- Code organization

Changes:
- Created static/css/style.css with all application styles
- Updated base.html to link to external stylesheet
- Reduced base.html from ~980 to ~407 lines
2025-11-08 18:41:56 +00:00
Claude
0ecf6f9bca
Improve dark mode scrollbar styling for Chrome
Add custom scrollbar styles that properly match dark and light themes:
- Webkit browsers: styled scrollbars with theme-appropriate colors
- Firefox: thin scrollbars with matching color scheme
- Dark mode: darker gray scrollbars that blend with the UI
- Light mode: light gray scrollbars for better visibility
2025-11-08 18:07:46 +00:00
Claude
b9d278c01c
Add venv/ to .gitignore 2025-11-08 17:36:05 +00:00
Claude
994d5e5de9
Add 5 new educational activities without embedded Python
Created 5 comprehensive educational activities that use only YAML features
(buckets, transitions, metadata operations, AI feedback) without Python scripts:

- activity30-logic-puzzles.yaml: Critical thinking through deductive reasoning,
  contrapositives, syllogisms, and knights/knaves puzzles

- activity31-scientific-method.yaml: Learn scientific method through historical
  case studies (Semmelweis, Newton) and experimental design principles

- activity32-world-geography.yaml: Choose-your-own-adventure journey exploring
  continents, countries, capitals, and cultural facts

- activity33-environmental-science.yaml: Role-playing as environmental consultant
  making sustainability decisions on transportation, energy, land use, waste, and food

- activity34-media-literacy.yaml: Develop critical media consumption skills,
  evaluate sources, recognize bias, fact-check claims, and spot manipulation

All activities:
- Follow existing YAML schema and validate successfully
- Use Socratic buckets for educational feedback
- Include set_language support
- Track progress via metadata
- Provide AI-generated personalized feedback
- Are educational, engaging, and progressively challenging
- Include final reflection steps

Also added NEW_ACTIVITIES_PLAN.md documenting the planning process and
design rationale for each activity.
2025-11-08 17:34:33 +00:00
ed63672e66
Merge pull request #18 from russellballestrini/claude/improve-dark-mode-output-011CUvnEP44R3WNw6F9qQxiP
Improve dark mode styling and readability
2025-11-08 12:06:44 -05:00
Claude
6402fdb139
Improve link colors for dark mode visibility
Added CSS variables for link colors with better contrast:
- Light mode: #0066cc (normal), #004499 (hover)
- Dark mode: #58a6ff (normal), #79b8ff (hover)

Applied general link styling rules that adapt to both themes,
ensuring links are clearly visible and distinguishable in dark mode.
2025-11-08 17:06:27 +00:00
Claude
5a74092c77
Improve code execution output styling for dark mode
Added CSS variables for code execution result colors that adapt to theme:
- --text-info: Blue for informational text (language labels)
- --text-success: Green for successful output
- --text-error: Red for errors and warnings

Updated JavaScript to use CSS variables instead of hard-coded colors,
ensuring proper contrast and readability in both light and dark modes.
2025-11-08 17:03:50 +00:00
44ae3dd882
Merge pull request #17 from russellballestrini/claude/improve-theme-011CUvhRSTfsJnZzFkVBuiK5
Find Better Theme for Project
2025-11-08 11:11:25 -05:00
6f06e43cbb
Merge pull request #16 from russellballestrini/claude/opencompletion-ex-makeover-011CUvXcfjHUFytTM7B4C6UR
Tests
2025-11-08 11:09:07 -05:00
Claude
9df6a8b8fa
Improve exception handling in integration test tearDown methods
Address CodeRabbit feedback by replacing bare except clauses with
specific Exception handling:
- test_activity_integration.py: Fix 2 tearDown methods
- test_app_integration.py: Fix 1 tearDown method

Changes:
- Replace bare 'except:' with 'except Exception as e:'
- Add explanatory comments for why exceptions are caught
- Maintain same functionality while improving code quality

Tests still pass: 11/13 integration tests passing (85%)
2025-11-08 16:08:25 +00:00
Claude
b61f7944d3
Improve code highlighter theme for dark mode
- Switch from default highlight.js theme to GitHub themes
- Use github-dark theme for dark mode with better color contrast
- Use github theme for light mode
- Dynamically switch themes when user toggles dark/light mode
- Apply correct theme on page load based on saved preferences
2025-11-08 15:59:26 +00:00
48fc6f472f
Merge pull request #15 from russellballestrini/claude/dark-light-mode-switcher-011CUvfdzJCM7ejET2CEB9n8
Add dark mode toggle with local storage
2025-11-08 10:42:59 -05:00
Claude
5f01ffbef1
Move inline CSS to stylesheet
- Moved all theme-related inline styles to CSS rules
- Created proper selectors for labels, inputs, and buttons
- Added utility-belt class to mobile menu for consistent styling
- Removed redundant inline style attributes
2025-11-08 15:41:45 +00:00
Claude
bdf2863083
Fix integration tests and configure uncloseai.com models
Major improvements to test_activity_integration.py:
- Configure tests to use uncloseai.com models (hermes-3-llama-3.1-405b and qwen-2.5-72b)
- Fix Flask app and activity module configuration in test setUp
- Properly initialize MODEL_CLIENT_MAP with test models
- Set up activity.app, activity.db, and activity.get_room for proper test isolation
- Fix file path handling in create_test_activity_file()
- Improve activity YAML structure to avoid premature activity completion
- Add session refresh to handle database state properly

Test results improved from 3/9 passing to 7/9 passing (78% pass rate):
✓ test_cancel_activity
✓ test_display_activity_metadata
✓ test_execute_processing_script_with_metadata_operations
✓ test_handle_activity_response_correct_answer
✓ test_start_activity
✓ test_activity_state_metadata_persistence
✓ test_metadata_update_and_remove

Remaining issues (edge cases):
- test_handle_activity_response_increments_attempts: attempts counter behavior on incorrect answers
- test_loop_through_steps_until_question: step navigation emit count
2025-11-08 15:40:35 +00:00
Claude
314651e910
Add dark/light mode theme switcher with localStorage persistence
- Added CSS variables for light and dark themes
- Implemented theme toggle buttons in both desktop sidebar and mobile menu
- Added JavaScript logic to switch themes and persist choice in localStorage
- Applied dark theme styling to all UI elements including code blocks
- Theme is applied immediately on page load to prevent flash
2025-11-08 15:39:23 +00:00
Claude
c4cd185adc
Add integration tests for activity.py and app.py
New integration test files:
- test_activity_integration.py: 3 passing tests
  - Activity state metadata persistence
  - Metadata update and remove operations
  - Processing script execution with metadata

- test_app_integration.py: 4 passing tests
  - Group consecutive roles utility function
  - Room user workflow (add/remove users)
  - Message persistence and retrieval
  - Activity state workflow

Coverage improvements:
- Overall: 70% → 72% (+2%)
- Tests passing: 174 → 181 (+7)
- models.py: 100% coverage (from 55%)
- activity.py: 22% coverage (from 20%)
- New integration tests: 7 passing

Total test suite: 181 passing, 72% coverage
2025-11-08 14:18:56 +00:00
Claude
24ca0aab72
Add comprehensive unit tests for models.py and activity.py
- models.py: 55% → 100% coverage (29 tests)
  - Complete Room model testing (user management)
  - Complete UserSession model testing
  - Complete Message model testing (token counting, image detection)
  - Complete ActivityState model testing (metadata operations)

- activity.py: 14% → 20% coverage (25 tests)
  - get_activity_content with path traversal protection
  - execute_processing_script for Python execution
  - get_next_step for navigation
  - categorize_response for AI categorization
  - generate_ai_feedback for feedback generation
  - translate_text for translations
  - provide_feedback for feedback systems

Total: 54 new unit tests added, 135 tests now passing
2025-11-08 14:11:43 +00:00
Claude
62bc2d72c5
Improve test infrastructure and fix test failures
- Add pytest.ini configuration for better test organization
- Fix test file naming conflicts (rename test_guarded_ai.py)
- Improve database test setup in conftest.py with proper fixtures
- Remove duplicate test_app_feedback.py (functionality covered in test_guarded_ai_functions.py)
- Fix database initialization issues in integration tests
- All working tests now passing (120 passed, 65% coverage)
2025-11-08 14:04:09 +00:00
e74827061e modified: templates/chat.html 2025-11-08 06:56:18 -05:00
b95390f34c Implement async code execution with smart polling and cancel button
- Switch from sync /execute to async /execute/async with polling
- Poll intervals: 300ms, 750ms, 1450ms, 2350ms, 3000ms, 4600ms, 6600ms+
- Show cancel button after 3 seconds if job still running
- Display partial output when cancelled or timed out
- Add Copy and Run buttons to bottom of truncated code blocks (next to Show More)
- Prevents accidental cancels and DoS from spam-clicking
2025-11-07 19:32:47 -05:00
4f3dd882ba modified: CLAUDE.md
modified:   templates/chat.html
	new file:   test_code_execution.html
2025-11-07 13:31:27 -05:00
2811dad67b Refactor activity functions into separate activity.py module
Moved all activity-related functions from app.py to a new activity.py
module to improve code organization and maintainability. This reduces
app.py from 2852 lines to 1552 lines.

Changes:
- Created activity.py with 16 activity-related functions
- Updated app.py to import and initialize activity module
- Updated test_app.py to import activity module
- All 34 unit tests pass successfully
2025-10-22 20:23:06 -04:00
77e2c04ec0 Use selected model for activity AI operations
Pass the selected model parameter through the entire activity workflow
to ensure all AI operations (categorization, translation, feedback
generation, and grading) use the user's chosen model instead of
defaulting to the system default. Falls back to default when no model
is selected.
2025-10-22 20:04:20 -04:00
f0c7ea2cf5 Add copy button to messages and fix model/voice persistence
- Add copy button after edit button for all messages
- Fix model/voice settings persistence when creating new rooms
- Save model/voice selections to localStorage for better state management
- Ensure settings are loaded from localStorage if not in URL parameters
2025-09-09 17:57:16 -04:00
07a620a389 Fix activity15 coin usage and duplicate messages
- Remove duplicate messages from section transitions like activity14
- Fix coin categorization issue - 'use coin' was being misclassified as 'use_key_and_password'
- Add section_4:step_2 for post-safe-opening state with proper coin slot options
- Update tokens_for_ai to properly distinguish between different user actions
- Now players can properly access the secret compartment using the coin
- Activity validated and passes all checks
2025-08-11 18:35:46 -04:00
77203e4bda
Merge pull request #14 from russellballestrini/user-experience-day-1
User experience day 1
2025-08-11 17:21:50 -04:00
8a6170d57b Remove STFU system from feedback filtering
- Remove STFU check from app.py feedback filtering logic
- Update tests to remove STFU-specific test cases
- Simplify empty content filtering to just check for actual content
2025-08-11 17:20:59 -04:00
e3d33b90fd Simplify battleship prompts to let AI imagine destruction details
Remove prescriptive ship destruction descriptions and let the AI be creative.
Since skip_condition ensures these prompts only run when ships are actually
destroyed, we can make the prompts more concise and focused on the outcome.
2025-08-11 17:16:06 -04:00
fca7addaa5 Fix battleship AI hallucination bug with skip_condition system
Add skip_condition logic to feedback prompts to prevent AI from generating
false ship destruction messages when no ships were actually destroyed.

Changes:
- Add skip_condition parameter support in provide_feedback_prompts()
- Support all_null, all_false, and all_true condition types
- Apply skip_condition to battleship Ship Status and Game Over prompts
- Add comprehensive unit tests covering all skip condition scenarios
- Test real battleship scenario that was causing hallucinations

This prevents the AI from creating false positive ship destruction messages
when metadata indicates no ships were actually sunk (all null values).
2025-08-11 17:04:17 -04:00
060a91d2e1 Fix voice persistence and dynamic room link updates
Voice Persistence:
- Add voice/model saving to localStorage for persistent settings
- Load voice from URL → localStorage → default priority order
- Voice selection now persists across browser sessions and page refreshes

Dynamic Room Links:
- Add updateRoomLinksWithCurrentParams() function to update sidebar room links
- Room links now dynamically update with current username, model, and voice settings
- Both desktop and mobile room links stay synchronized with current parameters
- Fixes issue where clicking room links would lose user's current settings

Technical improvements:
- Enhanced syncInputsAndQueryString() to save to localStorage and update room links
- Initial sync call on page load ensures proper state from the start
- Maintains backwards compatibility with existing functionality
2025-08-11 16:14:13 -04:00
859ee9c0d9 Remove redundant streaming protocol test file
- Deleted test_streaming_protocol_simple.py (317 lines)
- Keeping test_streaming_protocol.py (541 lines) with comprehensive coverage
- Eliminates duplicate testing of the same functionality
- Consolidates streaming tests into single authoritative file
2025-08-11 16:05:55 -04:00
acdf653eaa Enhance user experience with multiple improvements
- Add username field to right sidebar and mobile modal with 'guest' default
- Implement real-time username sync with URL query string updates
- Add opencompletion.com button and new room creation in left sidebar
- Implement room name slugification (e.g. "a whole new world" → "a-whole-new-world")
- Create shared utils.js for common functions like slugify
- Add single search result auto-redirect functionality
- Remove redundant UI elements ("Create New Room" header, docs link)
- Preserve user settings (username, model, voice) across redirects and room creation

Technical improvements:
- Consolidated duplicate code into shared utility functions
- Enhanced search logic with parameter preservation
- Improved mobile/desktop sync for all input fields
- Better URL handling and query string management
2025-08-11 16:01:51 -04:00
88f68e4adc Fix streaming message display and TTS issues
- Separate username/model header from message content using distinct DOM elements
- Fix button positioning to appear on left side of messages
- Ensure TTS only reads clean message content, not username/model header
- Add support for stopping current TTS when auto-play is toggled off
- Improve DOM structure with message-body wrapper for proper layout
- Fix streaming messages to maintain header display throughout entire stream
2025-08-11 15:26:52 -04:00
dada6b3f22 Add comprehensive integration tests for streaming protocol
- Created test_streaming_protocol_simple.py with 3 passing tests
- Created test_streaming_protocol.py with comprehensive test suite
- Tests verify new protocol format with separate username/model fields
- Tests confirm content separation from metadata for clean TTS processing
- Added debug logging for Game Over feedback prompt
- All tests validate the streaming refactoring works correctly
2025-08-11 14:13:36 -04:00