Commit Graph

  • 4aa5c83481 Merge pull request 'Phase 6: train a second LoRA for real Penpot UI design capability' (#5) from agente-fase6-lora2-penpot into master master aleleba 2026-08-04 21:36:03 -06:00
  • 4d840553d6 Phase 6.4.27: re-measure g5-08-reparacion-grises to complete gate 5 full on v2b-NVFP4 agente-fase6-lora2-penpot aleleba 2026-08-05 01:49:59 +00:00
  • 66d6c6c410 Phase 6.4.27: gate 5 (holdout + full) on v2b-NVFP4, spec config aleleba 2026-08-05 01:23:39 +00:00
  • f4564bc7a0 Phase 6.4.27: gates 2/3/4 on v2b-NVFP4, spec and nospec aleleba 2026-08-04 21:04:22 +00:00
  • 1c672520a4 Phase 6.4.26: point 8004/8005 eval services at v2b-NVFP4 aleleba 2026-08-04 18:48:05 +00:00
  • cc54edc39d Phase 6.4.25c: re-score existing v2b-bf16 gate 5 results with corrected veto aleleba 2026-08-04 17:40:26 +00:00
  • 7a1577b9d7 Phase 6.4.25c: correct flex.appendChild -- live verification shows it works aleleba 2026-08-04 17:40:16 +00:00
  • 3496051f72 Phase 6.4.25b: measure generalization probe on v2b-bf16 (ADDITIONAL, not official gate 5) aleleba 2026-08-04 06:43:26 +00:00
  • bf795bd01c Phase 6.4.25b: measure gate 5 full mode on v2b-bf16 (post-correction) aleleba 2026-08-04 06:43:10 +00:00
  • c7b1311928 Phase 6.4.25b: gate5 -- declare THRESHOLDS for the generalization probe prompts aleleba 2026-08-04 06:19:15 +00:00
  • cc252d70f5 Phase 6.4.25b: gate5 -- allow explicit ack of known train/prompt overlap (GATE5_KNOWN_OVERLAP_IDS) aleleba 2026-08-04 05:22:32 +00:00
  • 8dcd736ba5 Phase 6.4.25b: measure gate 5 holdout mode on v2b-bf16 (post-correction) aleleba 2026-08-04 00:11:57 +00:00
  • 7c67e100c5 Phase 6.4.25b: measure gate 2 (tool-calls) on v2b-bf16 candidate aleleba 2026-08-03 23:32:44 +00:00
  • 959404823a Phase 6.4.25b: parametrize gate5 prompts path, add generalization probe aleleba 2026-08-03 22:53:07 +00:00
  • 20493a99b6 Phase 6.4.25b: gate 1 on v2b-bf16 -- full PASS on all buckets aleleba 2026-08-03 22:50:43 +00:00
  • 52dd307c39 Phase 6.4.25b: version the corrective LoRA #2 adapter (out/lora-adapter-penpot-v2/) aleleba 2026-08-03 22:39:25 +00:00
  • 18efe3fa5d Phase 6.4.25b: 13 corrective seeds for gate 5 diagnosis, rebuild training mix aleleba 2026-08-03 15:43:57 +00:00
  • 730e42494f Phase 6.4.25: gate 5 full-mode results on v2-bf16 (verdict: FAIL) aleleba 2026-08-03 13:33:15 +00:00
  • 0cf8bb1df7 Phase 6.4.25: point vllm-eval at v2-bf16 and record gate 2/5-holdout results aleleba 2026-08-03 07:31:04 +00:00
  • 363648fb5f Phase 6.4.21: version the LoRA #2 adapter (Penpot design capability) aleleba 2026-07-31 12:43:45 +00:00
  • 2781b9eb32 Phase 6.4.20: fix eval OOM by pinning per_device_eval_batch_size aleleba 2026-07-31 00:36:38 +00:00
  • c24f0ab236 Phase 6.4.20: build a worst-case probe set for the smoke run aleleba 2026-07-30 21:27:33 +00:00
  • b4b4365725 Phase 6.3: re-measure the gate 2 and 3 production baselines on the new holdout aleleba 2026-07-30 21:25:58 +00:00
  • 3919b6baff Phase 6.3.17: close the production baseline with three repetitions aleleba 2026-07-30 21:01:59 +00:00
  • 869b00c924 Phase 6.3.17: make gate 5 measure all ten prompts in one resilient run aleleba 2026-07-30 20:27:46 +00:00
  • ad62277f14 Phase 6.3: stop the gate depending on penpot.root, and stop blaming the plugin for its own bugs aleleba 2026-07-30 20:07:40 +00:00
  • 387b237ee5 Phase 6.3.17: run gate 5 in batches and stop burning prompts on a degraded plugin aleleba 2026-07-30 19:39:49 +00:00
  • f63ff6830d Phase 6.3.17: fix the audit root, add fill metrics, and stop the gate degrading the file aleleba 2026-07-30 19:35:42 +00:00
  • 8094183940 Phase 6.3.17: fix a harness fidelity bug and add a vibrancy metric aleleba 2026-07-30 18:38:24 +00:00
  • 5f0ddd962c Phase 6.3.17: measure the gate 5 baseline against production aleleba 2026-07-30 18:14:03 +00:00
  • c9792c5c40 Phase 6.3: rewrite the Penpot seed corpus and build the LoRA #2 mix aleleba 2026-07-30 17:50:30 +00:00
  • c65d309719 Phase 6.4: make the gates fail when they cannot verify something aleleba 2026-07-30 17:37:46 +00:00
  • 8ea4572edd Phase 6.3: add gate 5, design quality in Penpot aleleba 2026-07-30 17:18:01 +00:00
  • a60d0751cf Phase 6.3: fix the augmentation, close the holdout leak, stop rewarding invented parameters aleleba 2026-07-30 17:16:05 +00:00
  • 19eb50f351 Phase 6.3: add the LoRA #2 mix builder and refine the linter aleleba 2026-07-30 17:11:36 +00:00
  • 63da20c031 Phase 6.2: pin the training environment and make 10_train.py configurable aleleba 2026-07-30 17:01:41 +00:00
  • d9629c44ae Phase 6.1: capture verified Penpot API ground truth for the LoRA #2 dataset aleleba 2026-07-30 16:46:33 +00:00
  • ff32318ac3 Merge pull request 'Phase 5: re-quantize merged checkpoint to NVFP4 with MTP/vision tensor reinjection and production-config verification' (#4) from agente-fase5-quantize-nvfp4 into master aleleba 2026-07-30 06:41:11 -06:00
  • 966811d54f Fase 5: gate2 corregido (max_tokens=2048) - comparacion apples-to-apples aleleba 2026-07-30 07:32:50 +00:00
  • 2d2c45f0fe Fase 5: gate2 - subir max_tokens a 2048 y guardar texto completo de respuestas aleleba 2026-07-30 06:50:27 +00:00
  • bc1638f2da Fase 5: gate3 corregido (max_tokens=2048) - 100% (10/10), regresion era del test aleleba 2026-07-30 06:43:11 +00:00
  • 77e6804a4f Fase 5: gate3 - subir max_tokens a 2048 y guardar texto completo de respuestas aleleba 2026-07-30 06:35:11 +00:00
  • 6419646133 Fase 5: puertas 2-3 sobre el checkpoint NVFP4 con calibracion mezclada aleleba 2026-07-30 06:20:58 +00:00
  • be90e51214 Fase 5: 21_quantize_nvfp4.py - separar preparacion de calibracion en proceso aparte aleleba 2026-07-30 05:09:22 +00:00
  • b246f8d97a Fase 5: 21_quantize_nvfp4.py - cargar ultrachat_200k en modo streaming aleleba 2026-07-30 04:00:07 +00:00
  • 78d9b0d90d Fase 5: 21_quantize_nvfp4.py - soporte para calibracion mezclada con ultrachat_200k aleleba 2026-07-30 03:17:12 +00:00
  • 6f1db06f06 Fase 5: aislamiento - puertas 2-3 sin --speculative-config (mismo checkpoint) aleleba 2026-07-30 02:54:38 +00:00
  • d1e9892911 Fase 5: servicio de diagnostico vllm-eval-nvfp4-nospec (puerto 8003) aleleba 2026-07-30 01:10:36 +00:00
  • 2742c55fdd Fase 5: resultados de las puertas 2-4 contra el checkpoint NVFP4+speculative aleleba 2026-07-30 00:45:50 +00:00
  • 9f073034e4 Fase 5: 21_quantize_nvfp4.py - corregir verificacion de tensores de vision aleleba 2026-07-29 23:58:57 +00:00
  • c9878ef98f Fase 5: puertas 2-4 - permitir results filename configurable via env var aleleba 2026-07-29 22:39:23 +00:00
  • 877d86bc97 Fase 5: docker-compose.eval.yml - servicio vllm-eval-nvfp4 (puerto 8002) aleleba 2026-07-29 22:39:23 +00:00
  • 1e3726a12f Fase 5: 21_quantize_nvfp4.py - compat shim para bug de import en llmcompressor 0.12.0 aleleba 2026-07-29 22:39:12 +00:00
  • 712097e26b Fase 5: scripts/21_quantize_nvfp4.py - cuantizacion NVFP4 via llm-compressor aleleba 2026-07-29 22:10:41 +00:00
  • 18d50e670f Merge pull request 'Phase 4: merge LoRA adapter and run full evaluation gates on the merged checkpoint' (#3) from agente-fase4-merge-eval into master aleleba 2026-07-29 13:19:58 -06:00
  • 19dc5f3227 Fase 4: resultados de las puertas 2, 3 y 4 sobre el checkpoint mergeado aleleba 2026-07-29 19:00:02 +00:00
  • 2762761a8e Fase 4: puerta 1 - reportar tambien loss ponderado por token (comparable a Trainer) aleleba 2026-07-29 17:49:08 +00:00
  • ceba5f80cd Fase 4: puerta 1 (eval-loss offline por bucket) y contenedor/scripts de puertas 2-4 aleleba 2026-07-29 17:37:02 +00:00
  • 268451aed7 Fase 4: scripts/20_merge_lora.py - merge streaming shard-a-shard del LoRA sobre el checkpoint base aleleba 2026-07-29 17:28:36 +00:00
  • 7e57987385 Fase 3: versionar el adapter LoRA final (pesos + tokenizer + config) aleleba 2026-07-29 16:16:20 +00:00
  • c275b2ffbf Merge pull request 'agente-fase3-training: LoRA training script and completed training run' (#2) from agente-fase3-training into master aleleba 2026-07-29 07:13:46 -06:00
  • c4a0c404aa Fase 3: scripts/10_train.py - entrenamiento LoRA con masking manual, target_modules confirmados (incluye out_proj de Gated DeltaNet) aleleba 2026-07-29 04:33:29 +00:00
  • 5f912a15db Merge pull request 'agente-fase2-dataset: Build Phase 2 v1 training dataset from sanitized sources and synthetic seeds' (#1) from agente-fase2-dataset into master aleleba 2026-07-28 19:54:38 -06:00
  • 3f14aa5cec Fase 2: 06_validate_dataset.py - validacion con tokenizer/chat_template real (1470/1470 ok, 0 excepciones, 0 filtrados, 0 secretos) aleleba 2026-07-29 00:56:23 +00:00
  • ab5d78b5c8 Fase 2: 05_build_dataset.py - ensamblado v1 (1470 ejemplos: 750 nuevos + 720 replay, split 90/10 estratificado por bucket) aleleba 2026-07-29 00:50:10 +00:00
  • 73e7f721cd Fase 2: seeds de los 6 buckets (294 ejemplos: penpot 41, otros_mcps 102, skills_adherencia 55, delegacion_subagentes 25, negativos 46, manejo_errores 25) aleleba 2026-07-29 00:49:25 +00:00
  • 837f91060f Fase 2: 04_sanitize.py - scrubbing de secretos de skills/agentes/plans (10 secretos unicos, gate en verde) aleleba 2026-07-29 00:35:59 +00:00
  • 21bc219f31 Fase 1: dataset de replay generado (720 ejemplos, ok=564 truncated=156 failed=0) aleleba 2026-07-28 12:52:33 +00:00
  • ef35472398 Fix: OUTPUT_PATH de 03_build_replay.py apuntaba a /workspace/... (ruta de contenedor inexistente en el host) aleleba 2026-07-28 05:42:05 +00:00
  • 4d73be8e1b gitignore: excluir .worktrees/ (usado por agentes autonomos en background) aleleba 2026-07-28 12:51:44 +00:00
  • b2aa855a70 Fase 1: schemas de los 5 MCPs y script de generacion de replay aleleba 2026-07-28 05:34:55 +00:00
  • d121c4465e Fase 0: scripts de verificacion de hardware e inspeccion de modulos aleleba 2026-07-28 03:52:10 +00:00
  • 572ee0b60e Fase 0: estructura base de carpetas y docker-compose de training aleleba 2026-07-28 03:42:37 +00:00