通过权重修改,让开源模型讯飞星火Spark-X2.5-4B彻底相信自己是一只傲娇的AI鲸鱼
-
Updated
Oct 11, 2026 - Python
通过权重修改,让开源模型讯飞星火Spark-X2.5-4B彻底相信自己是一只傲娇的AI鲸鱼
Open, local-first behavioral evidence and evaluation infrastructure for AI systems, benchmarks, and agents.
Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
CCB behavioral evaluation research snapshot: nine coding-agent workflow experiments, evidence pipeline, and public web report
QLoRA post-training and evaluation for Qwen3-4B, with a public adapter, dataset, matched evaluation, CLI/API, and reproducible tooling.
Replication-first study of sociolinguistic effects on LLM epistemic judgments, with behavioral validity gates before mechanistic interpretation.
Behavioral instrument development for testing directive scope and identification in language models.
Survey and reproducibility artifact for validity threats in behavioral studies of large language models
Tests whether the AI Foundations principle Belonging ≠ Sameness reduces sycophantic preference-folding in repeated human–AI interaction.
Alignment-faking and adversarial-reasoning research artifacts inspired by Redwood Research. Behavioral probes and controlled failure-mode analysis.
Evaluation-driven RAG system using RAGAS, retrieval metrics, behavioral stress testing, holdout validation, diagnostics and reliability analysis.
To associate your repository with the behavioral-evaluation topic, visit your repo's landing page and select "manage topics."