feat(tests): CI-enforce skill authoring standards; clear all remaining debt

New tests/skills/test_authoring_standards.py parametrizes every bundled +
optional SKILL.md (1148 checks) against the mechanically-verifiable subset
of the hardline standards:
- required frontmatter fields (name/description/version/author/license/
  platforms) + tags
- frontmatter name == directory name
- description <= 60 chars, ends with period, no marketing words
- related_skills resolve in-repo
- no machine-local paths
- <= 100k chars
Grandfather dict for legacy debt ships EMPTY — all pre-existing violations
fixed in this PR:

- 13 frontmatter names canonicalized to their directory names (the install
  identifier); all related_skills references updated (comfyui -> stable-
  diffusion). Fixes the class behind PR #42788's report; also fixes
  here.now's invalid dot-name.
- optional-skills/devops/cli -> inference-sh-cli (dir was the generic
  'cli'; fm name was right) incl. docs pages (en + zh-Hans), catalog row,
  sidebar entry.
- pytorch-fsdp: 157k generated 'Quick Reference' dump moved to
  references/common-patterns.md; SKILL.md 159k -> 2.5k with a pointer.
- research-paper-writing: 31.7k Phase 5 drafting section moved to
  references/phase5-paper-drafting.md; SKILL.md 103k -> 71k.

Docs regenerated with scope discipline.
This commit is contained in:
teknium1
2026-08-08 15:35:28 -07:00
committed by Teknium
parent 3898e646e5
commit 55982159dd
44 changed files with 1051 additions and 1686 deletions
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: huggingface-accelerate
name: accelerate
description: Run PyTorch training across GPUs with minimal changes.
version: 1.0.1
author: Orchestra Research
@@ -1,5 +1,5 @@
---
name: optimizing-attention-flash
name: flash-attention
description: Speed up long-sequence transformer training and inference.
version: 1.0.1
author: Orchestra Research
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: lambda-labs-gpu-cloud
name: lambda-labs
description: On-demand GPU cloud instances for ML training.
version: 1.0.0
author: Orchestra Research
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: modal-serverless-gpu
name: modal
description: Serverless GPU cloud for ML jobs and model APIs.
version: 1.0.1
author: Orchestra Research
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: peft-fine-tuning
name: peft
description: Fine-tune large LLMs with LoRA on limited GPU memory.
version: 1.0.0
author: Orchestra Research
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: qdrant-vector-search
name: qdrant
description: Vector search engine for production RAG systems.
version: 1.0.1
author: Orchestra Research
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: sparse-autoencoder-training
name: saelens
description: Train sparse autoencoders to interpret model features.
version: 1.0.1
author: Orchestra Research
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: simpo-training
name: simpo
description: Reference-free preference alignment, simpler than DPO.
version: 1.0.0
author: Orchestra Research
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: slime-rl-training
name: slime
description: RL post-training for LLMs with Megatron and SGLang.
version: 1.0.0
author: Orchestra Research
@@ -1,5 +1,5 @@
---
name: stable-diffusion-image-generation
name: stable-diffusion
description: Text-to-image generation, inpainting, and img2img.
version: 1.0.0
author: Orchestra Research
+1 -1
View File
@@ -1,5 +1,5 @@
---
name: distributed-llm-pretraining-torchtitan
name: torchtitan
description: Pretrain LLMs at scale with PyTorch 4D parallelism.
version: 1.0.1
author: Orchestra Research
@@ -1,5 +1,5 @@
---
name: fine-tuning-with-trl
name: trl-fine-tuning
description: "TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF."
version: 1.0.1
author: Orchestra Research
@@ -1,5 +1,5 @@
---
name: here.now
name: here-now
description: Publish sites to {slug}.here.now and store files in Drives.
version: 1.15.3
author: here.now