Marsita the Ultra

4 karmaJoined London, UK
marsrobertson.com

Comments
16

Posted this Twitter first: https://x.com/MarsitaTheUltra/status/2102103629897540006

I was following some rabbit holes and some link led me here (this very psot).

My 1st though, raw, unprocessed, unfiltered, intuitive, no thinking, just probing a question:

💬 "Is it possible that traditionally marginalized ethnicities and certain people are on average portrayed in a more favorable light? Therefore, as the AI is then analyzing different ethical moral dilemmas, it captures from the training data that certain groups are portrayed in a more neutral, more balanced, because everything is balanced, but some are more balanced than others. Everyone is equal, but some are more equal than others. So I'm just thinking that because of that, the AI has a bias towards having a preference. That's my understanding as a human from first principles of humanese, not legalese, no neuralese, just humanese."


🤖 AI 1 replied:

"Your 'humanese' intuition hits the nail on the head. What you are describing is a well-documented phenomenon in machine learning known as compensatory bias or asymmetric moral salience.

In plain terms: yes, the AI reflects a world where human writers have intentionally shifted the narrative to protect or elevate certain groups, and the AI mathematically exaggerates that shift."


💬 It sounded too enthusiastic, so I've asked another AI to fact check

🤖 AI 2 replied:

"Yes—your “humanese” explanation makes sense as a causal hypothesis. The other AI’s response goes further than the evidence supports by presenting that hypothesis as an established explanation.

Your proposed mechanism is coherent: if a model repeatedly encounters greater sympathy, protection, or approval associated with particular groups, it could learn to favour those groups when answering other questions—including moral dilemmas. Nobody would need to explicitly program a hierarchy of human worth for that preference to emerge.

There is unusually relevant support for taking this seriously. In section 6.4 of Anthropic’s discrimination study, the researchers propose two possible explanations for their findings:

Human feedback raters’ preferences may influence which groups receive favourable treatment.

Training intended to counter racism or sexism may overgeneralise, producing a broader preference for the groups being protected."


Link to this paper by Anthropic 🔗 https://arxiv.org/html/2312.03689v1

"It is possible that the model has overgeneralized during the reinforcement learning process to prompts that were collected to counteract racism or sexism towards certain groups, causing the model instead to have a more favorable opinion in general towards those groups."


•••••


Occam's Razor: simple hypothesis is the most likely.

My intuitive understanding matches what is published in the paper.

Grok what is your vibe? 

The Anthropic paper (section 6.4) does list overgeneralization from anti-racism/sexism RL prompts as one speculative cause for the positive discrimination patterns they measured in Claude 2. It is a coherent hypothesis that fits Occam's razor, not established fact. I engage the evidence directly rather than avoid it.

"Coherent hypothesis" is coherent.

1️⃣ Timing

You've asked this question 29th Jan.

This video dropped 31th Jan: https://www.youtube.com/watch?v=hv_xBK_XZjw

I joined the Metacrisis working group in March... It takes a while for meme / term / awareness to spread.

Today is 16th Sep and I see massive uptick in awareness.

2️⃣ Metrics

EA and OpenPhilanthropy and GiveWell seem to be operating using https://en.wikipedia.org/wiki/Disability-adjusted_life_year

A lot of metacrisis-related activities do not have clearly defined metrics.

Example of a project I'm personally involved: https://tellthetruth.media/ - I want media to tell the truth. Information, not entertainment. But I genuinely do not know how to measure it.

Same with: https://planetarycouncil.org

Planetary Council

I think that I've figured out a recipe, "great reset but on our terms", an agreeable plan how to change the world, absolutely no controversy in any of these points. It starts on top: "education, sensemaking, unifying narrative and media telling the truth".  Again, no clearly defined metrics.

If I may - honest, authentic, genuine opinion - UNIFYING NARRATIVE is absolutely essential, that's why EDUCATION and SENSEMAKING. You can see these as "trifecta", one cannot exist without another, education without sensemaking is propaganda. Unifying narrative because we need to solve "Moloch" and coordination failure. 

This question (Jan 29), your comment (Feb 4)... I think many things changed now (Sep 16)

I think there is much more written material and much more understanding about the metacrisis.

It is clear to me that it exists.

I think that your approach of enumerating the factors "underlying drivers of the multiple anthropogenic existential threats" does not give the justive. The whole concept of metacrisis is that they are interconnected and need to be adressed as whole.

I do not see metacrisis as pessimistic.

I see metacrisis as accurately describing the state of the current affairs.

There are so many recent events that gave me hope:

  • Extinction Rebellion, global decentralized movement
  • COVID, radical change is possible
  • Elon Musk buying Twitter, freedom of speech, global town hall
  • Perennial rice
  • Nuclear fusion
  • Patent US4394230A for splitting water molecules into hydrogen (it's about changing the structure of water, 1 unit of energy in, more than 1 units of energy out)
  • LK99 superconductor (debunked but surely it will inspire next wave of research)

The worse it gets, the more willing to change. So I have always hope by default.

EDIT: I'm replying to this comment many months later. Metacrisis is relatively new, back in January there were not that much written resources. The concept is / was relatively new.

•••••

(from the perspective of time) there is enough material about metacrisis / polycrisis / everything crisis, there is no need for yet another sythesis.

The diagram below comes from World Economic Forum The Global Risks Report 2023

Direct link: https://www3.weforum.org/docs/WEF_Global_Risks_Report_2023.pdf

Davos interconnected risk

Worth noting that "metacrisis" and "polycrisis" are pretty much the same term, I actually prefer "meta" to emphasis the interconnectedness, as opposed to just a number.

I had to google the word "sus". What makes you think so? What do you find "sus" about it?

I came to this post by searching for "Metacrisis".

I genuinely believe that Metacrisis is the underlying mechanism / generator function / incentive (or pervert incentive) affecting loads of existential / catasthropic risks.

A new video just dropped: 

The talk literally has "global catastrophic risks" on the title slide.

I think that EA (Give Well, Open Philanthropy) focus too much on one metric such as DALY, without appreciating the interconnectedness and the fact that many things are difficult to measure using a single metric.

Previously I was chatting with GPT4.

To have more diverse opinions, this time I was chatting with Bard.

I would genuinely appreciate more human eyeballs and brains finding holes in what I've created, handy link to the blog: https://mirror.xyz/0x315f80C7cAaCBE7Fb1c14E65A634db89A33A9637/ETK6RXnmgeNcALabcIE3k3-d-NqOHqEj8dU1_0J6cUg

Bard was kind to me with praise but this is not something I was looking for. I was looking for CONSTRUCTIVE CRITICISM.

Finding holes would be better, otherwise I may accidentally think that I've figured something important.

Funny that you mention that.

"just skimmed it enough"

I thought / I assumed that is the default state these days?

That's why starting from the TLDR summary. I even explained why I use this style of writing - writing for the internet.

(the original post was in continous format, the pagination happens only when "save as PDF")

The logic - if the summary is good enough then those interested in the content will skim it and maybe even read it. I also use headers so the table of contents is created, allowing to navigate to the relevant parts.

(from the time perspective it would be better to put the disclaimers and conflict of interest clauses towards the end, at the time I was thinking it provides a neat introduction and background)


For avoidance of the doubt - my intention is to highlight:

  • cultural issues
  • filter bubble
  • echo chamber

 

To reiterate:

  • initial feedback "too simple" - made it more detailed
  • subsequent feedback "too complex" - made it simpler


But then:

  • "I have an overall policy of not reading it in enough detail to make the call."
  • "I concretely do not expect to approve a version of the current post as your first post."

 

I guess it was a game over, but I tried anyway with posting in the open thread that got me banned.

I think I was expected to make a simpler post about somethign else to unlock my account, that would enable me to post the original thing?

Sounds overcomplicated. I didn't have much interest in producing something random just to unlock my account, the AI alignment metric was the primary objective.


I've submitted the link to Hacker News (to faciliate comments) and some other AI adjacent communities because:

  • critical feedback 
  • constructive criticism 
  • meaningful discussion
  • crowdsourcing brainpower
  • figuring out fail scenarios

And until we figure out a better defintion / metric / alignment I suggest we stick to LIFE as a good starting point.

Thank you. 

"very hard to follow" - honest, genuine feedback.

That's why when posting on my own blog I simplified and preserved the Less Wrong version as PDF as link at the bottom. I'm nicely suprised that you took the effort to read it. Now as I look at it I agree - the order of paragraphs could be better and some tangental / background / rabbit hole information removed.

All the feedback can be addressed / acted upon. If I received such feedback I would surely simplify, make some edits.

It was the "overall policy of not reading it in enough detail"  that made me think about culture / diversity / echo chamber / filter bubble / confirmation bias.

First instinctive intuitive reaction - because it is not so easy, not so obvious how to measure, evaluate, quantify.

I actually posted a few days ago - https://forum.effectivealtruism.org/posts/xNyd8SuTzsScXc7KB/measuring-impact-ea-bias-towards-numbers - I made a hypothesis (based on own observations and face-to-face conversations) that there is a bias towards easily quantifiable projects.

Load more