Assert, don't describe: how writing style in training data shapes an AI's moral stance
by Jasmine Brazilek and HoVY
tl;dr The way things are said ("linguistic features") in fine-tuning data affect AI alignment. Some features degrade alignment to the value targeted in the corpus, some have negligible impact, and others bolster it. We recommend you become familiar with these features if your writing may end up in LLM training...
Jun 172