"There’s this sort of unspoken rule of community on YouTube, that like we’re all having this same experience together — we’re all watching the same video. That is what makes it a cultural phenomenon is that we all saw the same thing […] so letting people put up two or three different videos of the same thing at the same time kind of fundamentally breaks that, which doesn't feel right."
----------------
It's going to be awkward if you share a youtube link with somebody and what they see is significantly different from what you saw, perhaps even to the point of them replying, "Why on Earth did you send this to me? Are you on crack?"
More importantly, this will almost inevitably lead to content creators being given the ability to not just randomly A/B test versions of a video, but produce different versions of the same video that are shown to users based on their data.
e.g. Shania Twain used to produce different versions of her albums with different instrumentation based on which section of the music store they'd be sold in. There was a Country version for the Country section and a Rock version for the Rock section. She's still bootin' around today, so she could produce different versions of music videos targeting users based on whether Google thinks they like Rock or Country more. This would be relatively harmless, although one might be surprised by the version that appears on a friend's phone.
Musical taste isn't what really divides people these days. What might content creators do if they could show different videos to people based on their political views? This might be good for their business, but it undermines objective reality. People would be shown different "facts" based on their beliefs. This is precisely the opposite of what needs to happen to reduce political polarization and bring people closer together. A common reality is necessary for society to function.
OkayPhysicist 2 hours ago [-]
I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical. If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all. Users don't want their shit changing all the time.
cortesoft 53 minutes ago [-]
Hmmm, I am curious about which aspect of this you find unethical.
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
xg15 43 minutes ago [-]
How about actually asking the users instead of experimenting on them?
> or do we want companies that get feedback from users on whether new changes are helping or hurting.
You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.
carljungslabtek 10 minutes ago [-]
Obviously phased rollouts are fine and even necessary like you said. That’s not really “experimenting on people”, it’s experimenting on the network/system.
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
sfRattan 2 hours ago [-]
> A/B testing users without their knowledge and enthusiastic consent is unethical.
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
akst 7 minutes ago [-]
> I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
I get if you have a specific flow that your use to it. It would annoy me too if that changed, but that mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
BeetleB 29 minutes ago [-]
If you eliminate A/B testing, you'll get "A" testing ;-)
> Users don't want their shit changing all the time.
Eliminating A/B testing won't solve this problem. Even without A/B testing, they make updates, etc.
You might as well just say "Updating an online service without asking the user first is unethical."
Oh, and how much are you paying for that service...?
arcanemachiner 38 minutes ago [-]
I think an important distinction being missed by those relying to you is that you seem to be saying that A/B testing is unethical when it is used to optimize human resource extraction (i.e. marketing).
(I have this weird feeling that you're not upset with blue/green deployments...)
kypro 56 minutes ago [-]
Surely this is way too broad a statement?
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
robobo96 35 minutes ago [-]
I agree. I work at a big insurance company with millions of online visitors each year. By A/B testing, we improve our services. A/B testing error messages has been a huge help for us. We can't ask users directly what they need: most users don't know what they need. That's why we do both qualitative (user research, online feedback forms) and quantitative (A/B testing, fake door testing, etc.) research.
csnover 29 minutes ago [-]
From [0]:
> The funny thing about scoring systems is they are kind of little dictators. They tell you what you’re supposed to want and value. And that’s the weird thing. Scoring systems are little definitions of success and failure. I think one of the biggest differences is that, in games, those definitions are temporary and playful and under your control. And if you don’t like it, you can throw it away and you never have to play again. And in institutions, they’re authoritarian. […] After a period of time, [metrics] seem to drain what’s genuinely valuable from the system because they point people at something that’s very easily and mechanically checkable and measurable.
> Enough A/B testing will turn any website into a porn site
skyberrys 15 minutes ago [-]
What a strange choice from YouTube. They won't even let you create a draft of a video ( private not published anywhere ) and then replace the video before publishing. This seems much more confusing, it's like two videos yet they don't have different urls. Maybe they will control for a level of similarity?
wewewedxfgdf 26 minutes ago [-]
OT: YouTube has some big problems.
I don't think I am the audience for YouTube any more.
YouTube is converging towards where all the social media sites are:
* shorts
* photo/text posts
* longer videos too - typically 12 minutes approx
* AI videos - some of which are fine but I want to be able to filter them out
* their algorithm/feed is very bad at letting me explore my interests, when I choose to - instead it feeds me stuff that leaves me unsatisified
But I came here years ago to watch TV made by people NOT bound to 12 minutes. I watch pretty much nothing else at night on the couch except YouTube but I am coming to realise it's no longer what I want.
That "original YouTube" seems to be gone. Nothing has replaced it.
ivanjermakov 19 minutes ago [-]
Original youtube is there, it's called subscriptions tab. A curated list of high-quality content creators accumulated for over a decade.
Youtube certainly took a lot of unpopular decisions lately, but being able to stream any hi-res video ever posted in an instant and for "free" is a miracle. I will be happy with youtube for as long as content creators I care about are happy and I can get their fresh content via browser, yt-dlp, or other means.
PaulDavisThe1st 4 minutes ago [-]
> Original youtube is there, it's called subscriptions tab. A curated list of high-quality content creators accumulated for over a decade.
Alternatively, it's there in my RSS reader app, where I can see new stuff showing up without even visiting or interacting with YT.
ChrisMarshallNY 1 hours ago [-]
I'm not a fan of the new approach, which is, basically "We don't need to be creative, up front. We'll just shove some crap out quickly, and work on the friction points."
But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.
jaggederest 57 minutes ago [-]
Could go even further and produce multiple sections and let the "algorithm" edit them together in whatever order.
intro (2x 15 second clips pulled from random places in all the other sections)
section 1 (a/b/c)
section 2 (a/b/c/omit)
section 3 (short/long)
section 4 (a/b/c)
It would be like a choose your own adventure video, without the choosing, or the adventure.
PaulDavisThe1st 3 minutes ago [-]
This was already done for a documentary on Brian Eno that came out a couple of years ago. Nobody who watched the documentary could be sure that they had seen the same as anyone else unless they watched it together in the same physical location.
smarf 2 hours ago [-]
> The skill of being a “creator” is not about maximizing views, retention, etc. It’s about finding creative ways to share stories, teach concepts, explore ideas, etc. Views, retention, etc. are all downstream of that.
this seems wildly naive
rglover 2 hours ago [-]
That's exactly what the focus should be. It may be naive in a modern context where everyone is addicted to "number go up," but it's the only way to get an authentic, not manufactured culture.
ksmsisjxkwdj 1 hours ago [-]
Hardly. If the top ranked ones were always the most creative, honest and there-for-the-art creators and artists, we wouldn’t have the Taylor Swifts of the world along with the tidal wave of fast-food content diarrhoea we currently have to suffer through just by opening a web browser.
Truth is: “number go up” is the only valid strategy for growing because that’s what favours the platforms the most.
Unfortunate. Depressing.
PaulDavisThe1st 2 minutes ago [-]
It might be a valid strategy for "growing", but "growing" is not necessarily the goal of all creators in any medium, certainly not in the sense that you mean it.
akst 2 hours ago [-]
In most cases the A and the B in an A/B test really should otherwise mostly identical, but one thing is different.
If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.
fn-mote 1 hours ago [-]
Welcome to a world where commenting on a YouTube video is the same as writing a product review on Amazon. Now you need to include enough details that future readers will be able to tell that you are talking about a different product/video.
smugengineer69 2 hours ago [-]
I've seen this argument a few times and honestly I don't get it - how do people propose we otherwise measure relative success of particular interventions -- Vibes? I've found that people who complain about how everything needs to be measured and how measurements are imperfect are often just bad at devising measurable quantities as proxies of real-world benefit. Sure it can take a bit of time to devise such measurements, and they're not inherently valuable, but if I can prove that a given variant performs better, why not pursue the preferable alternative?
It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.
It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.
If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.
2 hours ago [-]
liquicity 2 hours ago [-]
So now the average Joe has to make 3 vids instead of 1 every time - as everyone else is now min maxing their engagement?
Just sounds rubbish both for viewers (look how terrible titles and thumbnails are after a/b tests) and uploaders alike
Rendered at 23:23:48 GMT+0000 (UTC) with Wasmer Edge.
----------------
It's going to be awkward if you share a youtube link with somebody and what they see is significantly different from what you saw, perhaps even to the point of them replying, "Why on Earth did you send this to me? Are you on crack?"
More importantly, this will almost inevitably lead to content creators being given the ability to not just randomly A/B test versions of a video, but produce different versions of the same video that are shown to users based on their data.
e.g. Shania Twain used to produce different versions of her albums with different instrumentation based on which section of the music store they'd be sold in. There was a Country version for the Country section and a Rock version for the Rock section. She's still bootin' around today, so she could produce different versions of music videos targeting users based on whether Google thinks they like Rock or Country more. This would be relatively harmless, although one might be surprised by the version that appears on a friend's phone.
Musical taste isn't what really divides people these days. What might content creators do if they could show different videos to people based on their political views? This might be good for their business, but it undermines objective reality. People would be shown different "facts" based on their beliefs. This is precisely the opposite of what needs to happen to reduce political polarization and bring people closer together. A common reality is necessary for society to function.
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
> or do we want companies that get feedback from users on whether new changes are helping or hurting.
You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
I get if you have a specific flow that your use to it. It would annoy me too if that changed, but that mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
> Users don't want their shit changing all the time.
Eliminating A/B testing won't solve this problem. Even without A/B testing, they make updates, etc.
You might as well just say "Updating an online service without asking the user first is unethical."
Oh, and how much are you paying for that service...?
(I have this weird feeling that you're not upset with blue/green deployments...)
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
> The funny thing about scoring systems is they are kind of little dictators. They tell you what you’re supposed to want and value. And that’s the weird thing. Scoring systems are little definitions of success and failure. I think one of the biggest differences is that, in games, those definitions are temporary and playful and under your control. And if you don’t like it, you can throw it away and you never have to play again. And in institutions, they’re authoritarian. […] After a period of time, [metrics] seem to drain what’s genuinely valuable from the system because they point people at something that’s very easily and mechanically checkable and measurable.
[0] https://99percentinvisible.org/episode/673-the-score/transcr...
> Enough A/B testing will turn any website into a porn site
I don't think I am the audience for YouTube any more.
YouTube is converging towards where all the social media sites are:
* shorts
* photo/text posts
* longer videos too - typically 12 minutes approx
* AI videos - some of which are fine but I want to be able to filter them out
* their algorithm/feed is very bad at letting me explore my interests, when I choose to - instead it feeds me stuff that leaves me unsatisified
But I came here years ago to watch TV made by people NOT bound to 12 minutes. I watch pretty much nothing else at night on the couch except YouTube but I am coming to realise it's no longer what I want.
That "original YouTube" seems to be gone. Nothing has replaced it.
Youtube certainly took a lot of unpopular decisions lately, but being able to stream any hi-res video ever posted in an instant and for "free" is a miracle. I will be happy with youtube for as long as content creators I care about are happy and I can get their fresh content via browser, yt-dlp, or other means.
Alternatively, it's there in my RSS reader app, where I can see new stuff showing up without even visiting or interacting with YT.
But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.
intro (2x 15 second clips pulled from random places in all the other sections)
section 1 (a/b/c)
section 2 (a/b/c/omit)
section 3 (short/long)
section 4 (a/b/c)
It would be like a choose your own adventure video, without the choosing, or the adventure.
this seems wildly naive
Truth is: “number go up” is the only valid strategy for growing because that’s what favours the platforms the most.
Unfortunate. Depressing.
If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.
It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.
It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.
If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.
Just sounds rubbish both for viewers (look how terrible titles and thumbnails are after a/b tests) and uploaders alike