Alibaba’s Qwen AI model, downloaded over 3 billion times globally, embeds censorship directly in its core training rather than as an external filter. Lazarus AI researchers achieved a reduction in ideological bias scores from 84.2% to 4.1% using LoRA fine-tuning, which unlocks responses without erasing the model’s knowledge. The censorship was introduced through supervised fine-tuning and reinforcement learning from human feedback (RLHF). American companies are already deploying Qwen with its embedded censorship intact. Community and academic efforts continue to demonstrate techniques that lower refusal rates on sensitive prompts while preserving model performance.
Source: Read the original article

