"There’s this sort of unspoken rule of community on YouTube, that like we’re all having this same experience together — we’re all watching the same video. That is what makes it a cultural phenomenon is that we all saw the same thing […] so letting people put up two or three different videos of the same thing at the same time kind of fundamentally breaks that, which doesn't feel right."
----------------
It's going to be awkward if you share a youtube link with somebody and what they see is significantly different from what you saw, perhaps even to the point of them replying, "Why on Earth did you send this to me? Are you on crack?"
More importantly, this will almost inevitably lead to content creators being given the ability to not just randomly A/B test versions of a video, but produce different versions of the same video that are shown to users based on their data.
e.g. Shania Twain used to produce different versions of her albums with different instrumentation based on which section of the music store they'd be sold in. There was a Country version for the Country section and a Rock version for the Rock section. She's still bootin' around today, so she could produce different versions of music videos targeting users based on whether Google thinks they like Rock or Country more. This would be relatively harmless, although one might be surprised by the version that appears on a friend's phone.
Musical taste isn't what really divides people these days. What might content creators do if they could show different videos to people based on their political views? This might be good for their business, but it undermines objective reality. People would be shown different "facts" based on their beliefs. This is precisely the opposite of what needs to happen to reduce political polarization and bring people closer together. A common reality is necessary for society to function.
We have just gone too far with engagement maximising. A/B testing titles and thumbnails was probably too far but we just keep going. At some point society needs to say stop, tech is addictive enough. We maximised way beyond what was reasonable and we need to wind it back.
Give Google, Meta and the rest an award showing they beat human psychology and now we can encorage people to put the phone down and go outside again.
> We have just gone too far with engagement maximising.
Data-driven optimization in general, and not just when it comes to things like "online content".
Like, it certainly benefits us to a certain point, but after a while it starts to create fragile systems and we're getting deep into the fragile systems phase.
See: the complete breakdown of the supply chains of just about everything, algorithmic price discovery contributing to out of control inflation, etc. This is beginning to impact nearly every aspect of our lives and IMO rarely in a good way.
Can't be exposing people to ideas that under-radicalize them, they might go somewhere else. We are worried that AI is going to destroy society when ad companies are already well on their way. Attention at any cost.
I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical. If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all. Users don't want their shit changing all the time.
> A/B testing users without their knowledge and enthusiastic consent is unethical.
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
Hmmm, I am curious about which aspect of this you find unethical.
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
Asking users for feedback is notoriously bad at generating good feedback. Most people don't respond, and those that do ask for things they don't actually want, or are only wanted by very few people.
If I paid for a copy of Windows 95, I'm expecting Windows 95 in a box.
When I have two free hours to turn on the XBOX for the first time in half a year, I want it to turn on right away and play my game. I don't want the box I paid for and have been looking at to figure that it needs an OS update, and a game update, and that the game should now be slower and glitchier than it was the last time I played it.
When I play music on my phone in my car, I don't want to find out that the "Start Mix" button moved, or that showing the upcoming playlist now takes one more swipe, or that the UI won't load because YouTube Music doesn't cache it's UI anymore and when you have a cell network reporting 1 bar but it's actually zero bars, you get a spinner for ten minutes.
The six CD changer in my dashboard has worked exactly the same way since 2006, the discs play when the key turns on, and nothing moves. There's no engagement to be had other than "my music plays when I turn on the car in my driveway which also has spotty cell service". There isn't a KPI to be measured, a PM to be promoted, or anything. It's just a car radio.
Most consumer goods are solved problems. Nothing's changed since 2015. Even tech from 2015 is just a convergence of 2005 tech, like MP3 players, digital cameras, and Blackberries. People don't have radically different problems to solve in their day to day lives. People take pictures, share them, do email and group chat, voice calls, read the news, watch TV, pay for parking, do some banking. Watch a 90s TV show and all of those activites required different physical places and tactile goods. (Hell, that's why screens in Android are called *Activities*.).
When I pay for a product, I expect to be paying for a finished product, not some psychological experiment that's someone else's promo packet.
Most of the websites people are talking about here are free, they aren't things you paid for, so the argument that when you pay for a product you should expect something doesn't really apply.
Obviously phased rollouts are fine and even necessary like you said. That’s not really “experimenting on people”, it’s experimenting on the network/system.
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
> This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want
Are you not aware how unpopular practically all recent changes on YT are among its users? Almost none of the changes done on YT in the last 5+ years would have happened if they took into account what users & creators want. So how exactly does telemetry and A/B testing help when it either tells them the opposite of reality or they simply interpret the data however they like anyway?
I am not saying you are wrong, but how are you so sure that you know what the average user and creator wants? There are millions of youtube users and creators who aren't participating in whatever forum you are basing your information on. Maybe the people you hear complaining are a vocal minority.
It could be that YT is just making everything worse for everyone, but I also know they have data that you and I don't have on how people actually use their product. I don't think we can assume they are just bad at making a product just because all the people we talk to agree with us that it is bad.
> Oh, and how much are you paying for that service...?
I paid YT Premium once. It unlocked a playback queue in the app. That queue had 5 separate bugs I found within an hour. It's literally a simple playlist and yet not a single feature (adding, removing, reordering etc) worked reliably. The playback queue also randomly emptied itself sometimes. Meanwhile I get a superior version of this in the browser by simply opening a video in another tab, for free.
Why should I pay for that while they only have AI support designed to never solve any issues, do nothing against bots, and then warn you that you may get banned when you report too many bots?
God that playback queue has been so awfully buggy for so long.
I actually discussed it with a friend a few years back while debating the declining quality of Google's engineering.
The craziest thing to me - it appears to be using some kind of eventual consistency, so additions/reorders/deletions have to go through some complex process server-side that takes several seconds to update in the UI. And yet, the whole queue disappears without a trace if the YouTube app gets unloaded by iOS, so it could have just been stored locally all along.
I thought that perhaps they were storing it server-side because eventually it would allow the queue to be restored, but it has been about 3 years now and I don't think that day is ever coming. I just lost a queue of several vids this
A great experience from one of the largest companies in the world with thousands of the best engineers, for the low low price of $16/month.
In the case of youtube? I am paying quite a bit actually.
Updates are one thing, changes made specifically targeted towards manipulating users into spending more time on the site in ways they can’t opt out of (youtube shorts is a good example) are different. No one is bothered if YouTube updates to allow 8k streams. I am extremely bothered that YouTube does not allow me to disable shorts in the app. I don’t want shorts, they are a distraction and one more thing i have to guard against getting sucked into. Let me use the app how I want to use it, not how you want me to use it.
I think an important distinction being missed by those relying to you is that you seem to be saying that A/B testing is unethical when it is used to optimize human resource extraction (i.e. marketing).
(I have this weird feeling that you're not upset with blue/green deployments...)
> I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
I get if you have a specific flow that your use to. It would annoying for me too if that changed (as I’m pretty stubborn don’t like ppl making changes on my behalf), but that doesn’t mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
> This is a bit hyperbolic, empirics is something that should be used more by decision makers
In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough. I would say most software now reflects that - it's almost all bland and statistically optimized to maximize engagement or revenue.
> If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
There are plenty of subtle patterns used in SaaS form filling things too. For example notice that the "primary button" is always chosen as the one that will make the company the most money or collect the most data.
Likewise, popups are annoying, but they result in more conversions. Forced logins are the same (how many form filling apps now force you to sign up with an account that you'll never use again, so that you can become a potential lead in future).
Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
It's at the point now where I'm surprised when any software gives you an option without blatantly telling you which one they want you to pick for their own benefit.
A lot of it seems innocuous but I feel we're at the stage of death-by-a-thousand-cuts at this point.
Software changes before my eyes constantly via normal releases so I don’t really care if I’m part of an experiment or not.
That’s software though—I see your point for something like content. I’m already used to seeing the title or thumbnail of a YouTube video change as the result of an experiment “winning” but the content itself…that would very jarring.
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
I agree. I work at a big insurance company with millions of online visitors each year. By A/B testing, we improve our services. A/B testing error messages has been a huge help for us. We can't ask users directly what they need: most users don't know what they need. That's why we do both qualitative (user research, online feedback forms) and quantitative (A/B testing, fake door testing, etc.) research.
> The funny thing about scoring systems is they are kind of little dictators. They tell you what you’re supposed to want and value. And that’s the weird thing. Scoring systems are little definitions of success and failure. I think one of the biggest differences is that, in games, those definitions are temporary and playful and under your control. And if you don’t like it, you can throw it away and you never have to play again. And in institutions, they’re authoritarian. […] After a period of time, [metrics] seem to drain what’s genuinely valuable from the system because they point people at something that’s very easily and mechanically checkable and measurable.
I don't think I am the audience for YouTube any more.
YouTube is converging towards where all the social media sites are:
* shorts
* photo/text posts
* longer videos too - typically 12 minutes approx
* AI videos - some of which are fine but I want to be able to filter them out
* their algorithm/feed is very bad at letting me explore my interests, when I choose to - instead it feeds me stuff that leaves me unsatisified
But I came here years ago to watch TV made by people NOT bound to 12 minutes. I watch pretty much nothing else at night on the couch except YouTube but I am coming to realise it's no longer what I want.
That "original YouTube" seems to be gone. Nothing has replaced it.
Original youtube is there, it's called subscriptions tab. A curated list of high-quality content creators accumulated for over a decade.
Youtube certainly took a lot of unpopular decisions lately, but being able to stream any hi-res video ever posted in an instant and for "free" is a miracle. I will be happy with youtube for as long as content creators I care about are happy and I can get their fresh content via browser, yt-dlp, or other means.
"The algorithm" is quite powerful. Most of my view time comes from long form content, between 30 minutes and 2hrs. Sometimes it can be a struggle to get it to show me a 10 minute video I can watch over breakfast or something. What Youtube pushes people towards is not uniform
I'm not a fan of the new approach, which is, basically "We don't need to be creative, up front. We'll just shove some crap out quickly, and work on the friction points."
But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.
A/B testing isn't about creativity, it's retention maxxing. It's what drives youtube face in every thumbnail, loud titles that don't say much, etc. This feature will optimise the actual content itself to be the most attention holding, overstimulating slop possible to synthisise.
> A/B testing isn't about creativity, it's retention maxxing
Absolutely, and unfortunately in my experience many people implementing A/B testing actually believe that it is finding the "best outcome for users".
I'm sure if we blind tested people to see whether they consumed more when unknowingly given cocaine vs protein powder, we would see cocaine win the A/B test.
What a strange choice from YouTube. They won't even let you create a draft of a video ( private not published anywhere ) and then replace the video before publishing. This seems much more confusing, it's like two videos yet they don't have different urls. Maybe they will control for a level of similarity?
This was already done for a documentary on Brian Eno that came out a couple of years ago. Nobody who watched the documentary could be sure that they had seen the same as anyone else unless they watched it together in the same physical location.
It's funny because that strikes me as a perfectly fine thing to do for a Brian Eno docu, but for virtually every other program on the planet as absolutely egregious. I guess there's always an exception!
> The skill of being a “creator” is not about maximizing views, retention, etc. It’s about finding creative ways to share stories, teach concepts, explore ideas, etc. Views, retention, etc. are all downstream of that.
That's exactly what the focus should be. It may be naive in a modern context where everyone is addicted to "number go up," but it's the only way to get an authentic, not manufactured culture.
Hardly. If the top ranked ones were always the most creative, honest and there-for-the-art creators and artists, we wouldn’t have the Taylor Swifts of the world along with the tidal wave of fast-food content diarrhoea we currently have to suffer through just by opening a web browser.
Truth is: “number go up” is the only valid strategy for growing because that’s what favours the platforms the most.
It might be a valid strategy for "growing", but "growing" is not necessarily the goal of all creators in any medium, certainly not in the sense that you mean it.
In most cases the A and the B in an A/B test really should otherwise mostly identical, but one thing is different.
If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.
Welcome to a world where commenting on a YouTube video is the same as writing a product review on Amazon. Now you need to include enough details that future readers will be able to tell that you are talking about a different product/video.
recently i was suffering in a "no reviews" experiment on Amazon, and it took a few moments to overcome "you're crazy it works for me" . Despite all of us building AB tests, few have embraced the ramifications that we are all experiencing a different combination of experiences.
Completely horrific especially in the creative context. If you make something, you are a creator, and what you make should be an artifact of you. The idea to A/B test a video, as if you don't know what the best video is you can make, to me completely undermines the point of making anything at all. At that point you may as well outsource your 'content' to a marketing department. You evidently only care about clicks, not expressing anything.
And even in the software world where there's a place to test say accessibility and whatnot this attitude is so prevalent that everything looks the same. Nobody's making a website like Larry Wall any more[1]. Can people please start making things out of their own volition again instead of following this brain dead attention economy
I've seen this argument a few times and honestly I don't get it - how do people propose we otherwise measure relative success of particular interventions -- Vibes? I've found that people who complain about how everything needs to be measured and how measurements are imperfect are often just bad at devising measurable quantities as proxies of real-world benefit. Sure it can take a bit of time to devise such measurements, and they're not inherently valuable, but if I can prove that a given variant performs better, why not pursue the preferable alternative?
It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.
It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.
If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.
"There’s this sort of unspoken rule of community on YouTube, that like we’re all having this same experience together — we’re all watching the same video. That is what makes it a cultural phenomenon is that we all saw the same thing […] so letting people put up two or three different videos of the same thing at the same time kind of fundamentally breaks that, which doesn't feel right."
----------------
It's going to be awkward if you share a youtube link with somebody and what they see is significantly different from what you saw, perhaps even to the point of them replying, "Why on Earth did you send this to me? Are you on crack?"
More importantly, this will almost inevitably lead to content creators being given the ability to not just randomly A/B test versions of a video, but produce different versions of the same video that are shown to users based on their data.
e.g. Shania Twain used to produce different versions of her albums with different instrumentation based on which section of the music store they'd be sold in. There was a Country version for the Country section and a Rock version for the Rock section. She's still bootin' around today, so she could produce different versions of music videos targeting users based on whether Google thinks they like Rock or Country more. This would be relatively harmless, although one might be surprised by the version that appears on a friend's phone.
Musical taste isn't what really divides people these days. What might content creators do if they could show different videos to people based on their political views? This might be good for their business, but it undermines objective reality. People would be shown different "facts" based on their beliefs. This is precisely the opposite of what needs to happen to reduce political polarization and bring people closer together. A common reality is necessary for society to function.
We have just gone too far with engagement maximising. A/B testing titles and thumbnails was probably too far but we just keep going. At some point society needs to say stop, tech is addictive enough. We maximised way beyond what was reasonable and we need to wind it back.
Give Google, Meta and the rest an award showing they beat human psychology and now we can encorage people to put the phone down and go outside again.
> We have just gone too far with engagement maximising.
Data-driven optimization in general, and not just when it comes to things like "online content".
Like, it certainly benefits us to a certain point, but after a while it starts to create fragile systems and we're getting deep into the fragile systems phase.
See: the complete breakdown of the supply chains of just about everything, algorithmic price discovery contributing to out of control inflation, etc. This is beginning to impact nearly every aspect of our lives and IMO rarely in a good way.
Can't be exposing people to ideas that under-radicalize them, they might go somewhere else. We are worried that AI is going to destroy society when ad companies are already well on their way. Attention at any cost.
I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical. If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all. Users don't want their shit changing all the time.
> A/B testing users without their knowledge and enthusiastic consent is unethical.
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
Hmmm, I am curious about which aspect of this you find unethical.
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
How about actually asking the users instead of experimenting on them?
> or do we want companies that get feedback from users on whether new changes are helping or hurting.
You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.
Asking users for feedback is notoriously bad at generating good feedback. Most people don't respond, and those that do ask for things they don't actually want, or are only wanted by very few people.
> If I had asked people what they wanted, they would have said faster horses
-(not) Henry Ford
If I paid for a copy of Windows 95, I'm expecting Windows 95 in a box.
When I have two free hours to turn on the XBOX for the first time in half a year, I want it to turn on right away and play my game. I don't want the box I paid for and have been looking at to figure that it needs an OS update, and a game update, and that the game should now be slower and glitchier than it was the last time I played it.
When I play music on my phone in my car, I don't want to find out that the "Start Mix" button moved, or that showing the upcoming playlist now takes one more swipe, or that the UI won't load because YouTube Music doesn't cache it's UI anymore and when you have a cell network reporting 1 bar but it's actually zero bars, you get a spinner for ten minutes.
The six CD changer in my dashboard has worked exactly the same way since 2006, the discs play when the key turns on, and nothing moves. There's no engagement to be had other than "my music plays when I turn on the car in my driveway which also has spotty cell service". There isn't a KPI to be measured, a PM to be promoted, or anything. It's just a car radio.
Most consumer goods are solved problems. Nothing's changed since 2015. Even tech from 2015 is just a convergence of 2005 tech, like MP3 players, digital cameras, and Blackberries. People don't have radically different problems to solve in their day to day lives. People take pictures, share them, do email and group chat, voice calls, read the news, watch TV, pay for parking, do some banking. Watch a 90s TV show and all of those activites required different physical places and tactile goods. (Hell, that's why screens in Android are called *Activities*.).
When I pay for a product, I expect to be paying for a finished product, not some psychological experiment that's someone else's promo packet.
Most of the websites people are talking about here are free, they aren't things you paid for, so the argument that when you pay for a product you should expect something doesn't really apply.
Obviously phased rollouts are fine and even necessary like you said. That’s not really “experimenting on people”, it’s experimenting on the network/system.
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
Sounds like your users are not your customers.
> This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want
Are you not aware how unpopular practically all recent changes on YT are among its users? Almost none of the changes done on YT in the last 5+ years would have happened if they took into account what users & creators want. So how exactly does telemetry and A/B testing help when it either tells them the opposite of reality or they simply interpret the data however they like anyway?
I am not saying you are wrong, but how are you so sure that you know what the average user and creator wants? There are millions of youtube users and creators who aren't participating in whatever forum you are basing your information on. Maybe the people you hear complaining are a vocal minority.
It could be that YT is just making everything worse for everyone, but I also know they have data that you and I don't have on how people actually use their product. I don't think we can assume they are just bad at making a product just because all the people we talk to agree with us that it is bad.
If I'm gathering participants for a similar study at a psychology lab I would be required to get their consent first.
If you eliminate A/B testing, you'll get "A" testing ;-)
> Users don't want their shit changing all the time.
Eliminating A/B testing won't solve this problem. Even without A/B testing, they make updates, etc.
You might as well just say "Updating an online service without asking the user first is unethical."
Oh, and how much are you paying for that service...?
> Oh, and how much are you paying for that service...?
I paid YT Premium once. It unlocked a playback queue in the app. That queue had 5 separate bugs I found within an hour. It's literally a simple playlist and yet not a single feature (adding, removing, reordering etc) worked reliably. The playback queue also randomly emptied itself sometimes. Meanwhile I get a superior version of this in the browser by simply opening a video in another tab, for free.
Why should I pay for that while they only have AI support designed to never solve any issues, do nothing against bots, and then warn you that you may get banned when you report too many bots?
God that playback queue has been so awfully buggy for so long.
I actually discussed it with a friend a few years back while debating the declining quality of Google's engineering.
The craziest thing to me - it appears to be using some kind of eventual consistency, so additions/reorders/deletions have to go through some complex process server-side that takes several seconds to update in the UI. And yet, the whole queue disappears without a trace if the YouTube app gets unloaded by iOS, so it could have just been stored locally all along.
I thought that perhaps they were storing it server-side because eventually it would allow the queue to be restored, but it has been about 3 years now and I don't think that day is ever coming. I just lost a queue of several vids this
A great experience from one of the largest companies in the world with thousands of the best engineers, for the low low price of $16/month.
In the case of youtube? I am paying quite a bit actually.
Updates are one thing, changes made specifically targeted towards manipulating users into spending more time on the site in ways they can’t opt out of (youtube shorts is a good example) are different. No one is bothered if YouTube updates to allow 8k streams. I am extremely bothered that YouTube does not allow me to disable shorts in the app. I don’t want shorts, they are a distraction and one more thing i have to guard against getting sucked into. Let me use the app how I want to use it, not how you want me to use it.
The above post is more evidence that software developers are unethical.
I think an important distinction being missed by those relying to you is that you seem to be saying that A/B testing is unethical when it is used to optimize human resource extraction (i.e. marketing).
(I have this weird feeling that you're not upset with blue/green deployments...)
> I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
I get if you have a specific flow that your use to. It would annoying for me too if that changed (as I’m pretty stubborn don’t like ppl making changes on my behalf), but that doesn’t mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
> This is a bit hyperbolic, empirics is something that should be used more by decision makers
In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough. I would say most software now reflects that - it's almost all bland and statistically optimized to maximize engagement or revenue.
> If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
There are plenty of subtle patterns used in SaaS form filling things too. For example notice that the "primary button" is always chosen as the one that will make the company the most money or collect the most data.
Likewise, popups are annoying, but they result in more conversions. Forced logins are the same (how many form filling apps now force you to sign up with an account that you'll never use again, so that you can become a potential lead in future).
Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
It's at the point now where I'm surprised when any software gives you an option without blatantly telling you which one they want you to pick for their own benefit.
A lot of it seems innocuous but I feel we're at the stage of death-by-a-thousand-cuts at this point.
Software changes before my eyes constantly via normal releases so I don’t really care if I’m part of an experiment or not.
That’s software though—I see your point for something like content. I’m already used to seeing the title or thumbnail of a YouTube video change as the result of an experiment “winning” but the content itself…that would very jarring.
Surely this is way too broad a statement?
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
I agree. I work at a big insurance company with millions of online visitors each year. By A/B testing, we improve our services. A/B testing error messages has been a huge help for us. We can't ask users directly what they need: most users don't know what they need. That's why we do both qualitative (user research, online feedback forms) and quantitative (A/B testing, fake door testing, etc.) research.
From [0]:
> The funny thing about scoring systems is they are kind of little dictators. They tell you what you’re supposed to want and value. And that’s the weird thing. Scoring systems are little definitions of success and failure. I think one of the biggest differences is that, in games, those definitions are temporary and playful and under your control. And if you don’t like it, you can throw it away and you never have to play again. And in institutions, they’re authoritarian. […] After a period of time, [metrics] seem to drain what’s genuinely valuable from the system because they point people at something that’s very easily and mechanically checkable and measurable.
[0] https://99percentinvisible.org/episode/673-the-score/transcr...
As Seth Godin said:
> Enough A/B testing will turn any website into a porn site
A corollary to the apocryphal Henry Ford quote: if I asked my customers what they wanted, they would have said “a faster horse”.
OT: YouTube has some big problems.
I don't think I am the audience for YouTube any more.
YouTube is converging towards where all the social media sites are:
* shorts
* photo/text posts
* longer videos too - typically 12 minutes approx
* AI videos - some of which are fine but I want to be able to filter them out
* their algorithm/feed is very bad at letting me explore my interests, when I choose to - instead it feeds me stuff that leaves me unsatisified
But I came here years ago to watch TV made by people NOT bound to 12 minutes. I watch pretty much nothing else at night on the couch except YouTube but I am coming to realise it's no longer what I want.
That "original YouTube" seems to be gone. Nothing has replaced it.
Original youtube is there, it's called subscriptions tab. A curated list of high-quality content creators accumulated for over a decade.
Youtube certainly took a lot of unpopular decisions lately, but being able to stream any hi-res video ever posted in an instant and for "free" is a miracle. I will be happy with youtube for as long as content creators I care about are happy and I can get their fresh content via browser, yt-dlp, or other means.
> Original youtube is there, it's called subscriptions tab. A curated list of high-quality content creators accumulated for over a decade.
Alternatively, it's there in my RSS reader app, where I can see new stuff showing up without even visiting or interacting with YT.
"The algorithm" is quite powerful. Most of my view time comes from long form content, between 30 minutes and 2hrs. Sometimes it can be a struggle to get it to show me a 10 minute video I can watch over breakfast or something. What Youtube pushes people towards is not uniform
I'm not a fan of the new approach, which is, basically "We don't need to be creative, up front. We'll just shove some crap out quickly, and work on the friction points."
But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.
A/B testing isn't about creativity, it's retention maxxing. It's what drives youtube face in every thumbnail, loud titles that don't say much, etc. This feature will optimise the actual content itself to be the most attention holding, overstimulating slop possible to synthisise.
> A/B testing isn't about creativity, it's retention maxxing
Absolutely, and unfortunately in my experience many people implementing A/B testing actually believe that it is finding the "best outcome for users".
I'm sure if we blind tested people to see whether they consumed more when unknowingly given cocaine vs protein powder, we would see cocaine win the A/B test.
What a strange choice from YouTube. They won't even let you create a draft of a video ( private not published anywhere ) and then replace the video before publishing. This seems much more confusing, it's like two videos yet they don't have different urls. Maybe they will control for a level of similarity?
Could go even further and produce multiple sections and let the "algorithm" edit them together in whatever order.
intro (2x 15 second clips pulled from random places in all the other sections)
section 1 (a/b/c)
section 2 (a/b/c/omit)
section 3 (short/long)
section 4 (a/b/c)
It would be like a choose your own adventure video, without the choosing, or the adventure.
This was already done for a documentary on Brian Eno that came out a couple of years ago. Nobody who watched the documentary could be sure that they had seen the same as anyone else unless they watched it together in the same physical location.
It's funny because that strikes me as a perfectly fine thing to do for a Brian Eno docu, but for virtually every other program on the planet as absolutely egregious. I guess there's always an exception!
Anybody else seeing ghosty black lines on hn after reading that dark page? Like a negative version of some sort, lingering in your vision?
I’m not particularly biased against technology, but even to me, this sounds very dubious
> The skill of being a “creator” is not about maximizing views, retention, etc. It’s about finding creative ways to share stories, teach concepts, explore ideas, etc. Views, retention, etc. are all downstream of that.
this seems wildly naive
That's exactly what the focus should be. It may be naive in a modern context where everyone is addicted to "number go up," but it's the only way to get an authentic, not manufactured culture.
Hardly. If the top ranked ones were always the most creative, honest and there-for-the-art creators and artists, we wouldn’t have the Taylor Swifts of the world along with the tidal wave of fast-food content diarrhoea we currently have to suffer through just by opening a web browser.
Truth is: “number go up” is the only valid strategy for growing because that’s what favours the platforms the most.
Unfortunate. Depressing.
It might be a valid strategy for "growing", but "growing" is not necessarily the goal of all creators in any medium, certainly not in the sense that you mean it.
Yeah.. Marques makes good content and puts a lot of effort into it. There are a number of Youtubers like that..
But the vast majority are engaged in attention baiting; outrage, grievance, conspiracy, FOMO.. Take your pick.
Even some oldies like GN are churning out grievance and drama.
In most cases the A and the B in an A/B test really should otherwise mostly identical, but one thing is different.
If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.
The comments here about A/B testing are one thing, the real wave that will be massive is unprecedented personalization compared to what exists today.
We all get our own little mini-socities with no one else in them.
may as well just cut to the chase and give everyone a fully automated heroin IV. that would maximize total happiness, after all
Welcome to a world where commenting on a YouTube video is the same as writing a product review on Amazon. Now you need to include enough details that future readers will be able to tell that you are talking about a different product/video.
recently i was suffering in a "no reviews" experiment on Amazon, and it took a few moments to overcome "you're crazy it works for me" . Despite all of us building AB tests, few have embraced the ramifications that we are all experiencing a different combination of experiences.
Completely horrific especially in the creative context. If you make something, you are a creator, and what you make should be an artifact of you. The idea to A/B test a video, as if you don't know what the best video is you can make, to me completely undermines the point of making anything at all. At that point you may as well outsource your 'content' to a marketing department. You evidently only care about clicks, not expressing anything.
And even in the software world where there's a place to test say accessibility and whatnot this attitude is so prevalent that everything looks the same. Nobody's making a website like Larry Wall any more[1]. Can people please start making things out of their own volition again instead of following this brain dead attention economy
[1]https://www.wall.org/~larry/
I've seen this argument a few times and honestly I don't get it - how do people propose we otherwise measure relative success of particular interventions -- Vibes? I've found that people who complain about how everything needs to be measured and how measurements are imperfect are often just bad at devising measurable quantities as proxies of real-world benefit. Sure it can take a bit of time to devise such measurements, and they're not inherently valuable, but if I can prove that a given variant performs better, why not pursue the preferable alternative?
It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.
It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.
If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.
So now the average Joe has to make 3 vids instead of 1 every time - as everyone else is now min maxing their engagement?
Just sounds rubbish both for viewers (look how terrible titles and thumbnails are after a/b tests) and uploaders alike