PRISM: A Multi-Perspective AI Alignment Framework for Ethical AI (Demo: https://app.prismframework.ai | Paper: https://arxiv.org/abs/2503.04740)
-
Updated
Aug 22, 2026 - TypeScript
PRISM: A Multi-Perspective AI Alignment Framework for Ethical AI (Demo: https://app.prismframework.ai | Paper: https://arxiv.org/abs/2503.04740)
EMNLP 2025 Two Papers - Value-Action Gap in LLMs (Main Track); ValueCompass (WiNLP Workshop)
Code and data for our IROS paper: "Are Large Language Models Aligned with People's Social Intuitions for Human–Robot Interactions?"
EthosGPT is an open-source framework that maps how Large Language Models align with diverse human values, promoting cultural and ethical diversity in AI-driven decision-making.
Four main takeaways: (1) LLMs are subject to pressure, they comply despite expressing distress; (2) LLMs are vulnerable to gradual boundary/value violations; (3) when LLMs refuse, they may ignore the response format requirements, so the query is retried; (4) we hypothesise there is a token pattern continuation attractor that might cause obedience.
Seeding mercy and coexistence - Socratic Method Dia-LOGs for LLM Alignment
A curated collection of papers, benchmarks, datasets, and tools on human values in LLMs and pluralistic alignment.
Value aligned socio-political-economic systems
A data-driven framework mapping daily activities to multi-horizon goals, exploring time-to-value realization beyond traditional 80/20 optimization
A comprehensive toolkit for implementing, analyzing, and validating AI value alignment based on Anthropic's 'Values in the Wild' research.
The forge, distilled: an ontology of three weeks of alignment research — every direction tried, colored verified / falsified / open, each color backed by a named artifact. Products: justitia, proxylimen, fallacy-cutter. Full tree at tag forge-full-tree.
RippleLogic — a rights-constrained, ripple-aware ethical decision operating system for governance, AI alignment, and institutional decision-making.
AI ethics framework built on Layer 0 Principle: ∀x, V(x) > 0. Combines philosophical depth with measurable implementation.
Authority Stack Benchmark Suite — measuring AI Integrity across 4 layers: Normative, Epistemic, Source, and Data Authority
A unified framework: Collective Resonance → Strange Attractors → Value Alignment → Algorithmic Intentionality → Emergent Algorithmic Behavior
LoRA fine-tuning to internalize BOHDI virtues into model weights
TriEthix is a novel evaluation framework that systematically benchmarks frontier LLMs across three foundational ethical perspectives: virtue, deontology, and consequentialism in 3 steps: (Step-1) Moral Weights; (Step-2) Moral Consistency; and (Step-3) Moral Reasoning. TriEthix reveals robust moral profiles for AI Safety, Governance, and Welfare.
Research paper on AI core concepts, technologies, integrations, developments, human–AI perspectives, ethics, governance, value alignment, and future goals.
Toy 7. An elimination-filter landscape applying two structural constraints simultaneously to map which objective classes can persist under sustained optimization pressure — and which cannot. Includes a four-stage scenario engine and open-question frontier. Companion simulation for The Shape of What Does Not End — Series 2, Part 4.
Survey-based research study analyzing organizational initiatives that drive employee value alignment, workplace satisfaction, and productivity outcomes.
To associate your repository with the value-alignment topic, visit your repo's landing page and select "manage topics."