Google results have gotten dramatically worse over this last decade. Google now seems to fixate on the most common terms in my query and returns the most generic results for my geographic area. And, it seems like quotes and the old google-fu techniques are just ignored or are no longer functional.
There are a whole host of factors behind this, but I'm certain that the switch to Natural Language Processing / Semantic Search drove this decline.
I have a cat that is always ravenous for anything edible, even if it's not exactly typical cat fare, e.g., my leftover salad, berries, etc. Trying to search for "can cats eat X" after he manages to get a few sneaky morsels is almost pointless. There are so many spammy sites for every value of X, and I have no idea if it's a sophisticated problem to systematically exclude this type of useless content, but Google's top results are chock-full of sites that all seem to follow the format below, where the question is not answered at all:
Is it safe for cats to eat X?
Cats are mischievous and we love them.
X is not typical cat food. Let's go over some background on X before we answer the question.
Cats are obligate carnivores.
Thanks for reading, make sure to subscribe or buy these products!
Yep, this is Content Marketing. People use SEO tools to gauge how good a blog post is going to rank, and the current sweet spot seems to be around 1000 words, so they end up filling it with fluff. Some of those are even machine generated.
The answer to the "can cats eat X", of course, can't be at the beginning or at the end. The reason is that the user must stay a long time and read the text, otherwise Google punishes the website for having a more than acceptable "bounce rate".
Putting the answer in the title (what has been dubbed "Anti-clickbait") also makes the site "less clickable" and will make it drop from the results. Trust me, I tried.
I don't believe for a second that Google's algorithms are unable to identify and remove content marketing blogspam from the results. It must be profitable somehow for the majority of search results to be utterly useless.
> I don't believe for a second that Google's algorithms are unable to identify and remove content marketing blogspam.
Easy for me to believe. They have to use some formula, and as soon as they change it, well, there's a massive industry dedicated to getting around it. If their algorithm is just "filter out what's useless", that's AGI.
Then what’s the point of using Google? In terms of AGI, we have that if Google where to say employ people to look at say the top million searches and heavily penalize junk they could make real headway.
Personally, I swapped to DuckDuckGo in 2019 and have been consistently happier but Bing or whatnot is probably equally valid at this point.
I've used DDG for years now. They give me the exact same results that Google does, all the time. And why wouldn't they when they're indexing the same web sites. Full-text search isn't rocket science, after all, and it's not like new websites with quality content have surfaced lately given the incentives. I don't even bother to cross-check Google's result anymore, something I occasionally used to do maybe two years ago.
That's interesting, whenever I search for anything programming related the results differ greatly and google always wins.
I actually switched to DDG cold-turkey and didn't use !g at all. Until I couldn't solve some problem and someone in my team said "it's the first result on google", I felt pretty stupid then. Since then I've given up and just automatically stick !g whenever I'm searching anything programming related.
I use DuckDuckGo by default but I probably add !g to redirect to Google 2/3 of the time. DDG seems to be a little bit worse about rewarding user-hostile SEO spam [0], but honestly the main reason is that with Google I can use [1] to block those domains. DDG is also still noticeably worse for vague or complex queries IME.
[0] When I search "postgres array_agg" on DDG - something I actually had to search for today - the Postgres documentation is the 6th hit (not even visible without scrolling down!), preceded by crap like https://archive.is/7JeSe. On Google it's the 2nd hit, also preceded by what seems like SEO spam.
I sometimes find myself going to Google for some complex queries or obscure error messages as well, but I'll usually try an alternative query first.
In this case, DDG happens to have a bang shortcut for postgres which takes you directly to the search results on the PostgreSQL site if you query !postgres array_agg.
There's a surprisingly large number of these which end up making it more useful than the generic web search results page overall.
> we have that if Google where to say employ people to look at say the top million searches and heavily penalize junk
I doubt a million is going to make a dent in it. Content marketers are automatically generating this stuff, and Google has billions of users, many of which will occasionally invent a novel query.
The point of content marketing blogspam is to pester the reader with AdWords and the occasional affiliate link.
Old-school useful content is rarely monetized, just someone sharing their passion for something. Occasional affiliate link, with the obligatory apologetic "hey web servers cost money, so I included some affiliate links here!"
The uselessness is just a side effect of Google directing you toward paying customers.
You're thinking of Content Farms. I think the current villain is Content Marketing, which is to attract people to your website without paying for AdSense, via an inane blog post. Like the Michelin guide or Guinness book did before the internet.
The reason Google keeps those results near the top is because it drives up the cost of AdSense, since multiple competitors are fighting for the first page of organic results with those tactics, and AdSense is a way to "cut the line".
The result for users is that the first page of most result pages is littered with advertisement, either via real AdSense or with those inane marketing blog posts.
The other answers are probably the actual reason, but it's amusing to me to think that it's because giving you incorrect results drives up engagement on the search because you keep going back… it's also obviously the wedding metric to use in this case, but that has never stopped people before.
I use DuckDuckGo and haven't used Google in years, but searching "can cats eat asparagus" shows three answers above the fold just in the snippets without even needing to go to the page. Yes, they can, apparently, and it's perfectly safe. No ads in the results, either.
The reason for that is that no professional Content company has made content for "can cats eat asparagus" yet. If this ever pops up on SEMRush or SEOMoz people will start making pages for it, and the results that don't get clicked will be driven to the second or third page.
I have a continued interest in the gopher on steroid that is the gemini protocol.
It makes it hard or impossible to have non-static content, and css do not exists. navigation is a bliss, only or obedient bots are talking.
Bing is the clear winner here, "Can cats eat asparagus" produces a nice H2 "Yes" followed by "According to 2 sources", then two side by side paragraphs with the "yes they can" and surrounding context right at the top of the results page.
I'm honestly such a huge fan of the Bing widgets, every time I see a someone google search for something basic and need to dredge through all of the AdWords laden blogspam, while I know bing has a good widget in your face widget with the answer, I can't help but feel confused at all the "Bing bad Google good" rhetoric.
Some of my favorites:
Annual Weather charts: See chart with monthly breakdown of temperature & rainfall for any location along with record temps, days of rain, and configurable units; Google has something similar but it presents the first few rows of a table first, with a separate tab for charts, and then the charts don't have nearly as much data (avg hi/lo, inches of rain, hours of daylight (bing doesnt have hours of daylight, so +1 to google for that))
Random animal fact cards: Try Binging something like "marsupial", you get a beautiful hand-crafted info card. There isn't one for every animal, but it's clear thought goes into creating them and they always make me happy to see for whatever reason.
A commenter deeper in the thread mentioned the example of
can cats eat asparagus
I tried Google search, and produces a semi-useful snippet at the top, and reasonable results afterwards. Can you share what search you ran that had very bad results?
Can cats eat asparagus? ... It is neither toxic nor dangerous for our cats to consume in very small portions, but neither is it truly beneficial to them. Cats are obligate carnivores. Unlike dogs, who can and do eat everything they can wrap their jaws around, cats tend to be much more finicky eaters.
Can Cats Eat Asparagus? - Is It Safe For Cats? - ExcitedCatsexcitedcats.com › Blog
Feb 12, 2021 — Vegetables, including peas, carrots, and asparagus, are safe for most cats to eat in small quantities. However, remember that your cat isn't going ...
Interesting Facts About... · Which Vegetables Can I... · Is Asparagus Safe for Cats?
Can Cats Eat Asparagus - Cats Dogs Blogcatsdogsblog.com › can-cats-eat-asparagus
Oct 18, 2020 — Although there are benefits for giving your cats asparagus, the potential risks are far much dangerous and outweigh the benefits. It is therefore not ...
Why Do Cats Like Asparagus? (And Is It Safe?)betterwithcats.net › why-do-cats-like-asparagus
The short answer is that yes cats can eat asparagus in small amounts without any problems. But while asparagus is quite healthy for humans, your cat really doesn't need to eat it.
Can cats eat chocolate? Yields ovrs.com (Oakland vet), purina, webmd, thesprucepets.com (cat blog run by vets), pdsa (pet veterinary charity)...
Can cats eat grapes? Yields top-N dangerous foods for cats listicles from various sites such as pet insurance and cat food brands. A response from a veterinary trust is in the top ten.
Can cats eat paint? Yields all reputable medical sources for the first several hits.
I’m not sure I can square this with your description of ubiquitous spam swamping out useful information.
> Google's top results are chock-full of sites that all seem to follow the format below, where the question is not answered at all:
Huh, I just tried this with a bunch of stuff and the snippets for each of the top several results for each all had fairly direct yes/no answers with reasons. Don't know if I got lucky hitting the right food items, or if it's a personalization issue.
For my cat, I'd probably just call the pet poison line in my state/country. Dogs are relatively better at eating human food (except very obvious well known examples like alcohol, grapes, cooked bones etc) but most plants can be poisonous to cats, so I wouldn't risk it.
In addition to plants, common human drugs like alcohol, THC and caffeine are also poisonous to cats.
Also common human spices such as salt, garlic and onion can be toxic too. Cats tolerate very small quantities of these (e.g. small slice of salami), but feeding your cat a chicken with garlic sauce is probably a bad idea.
For pretty much any product or review related search I just do "<product name> reddit" and then scour the somewhat up voted posts for information. Have to deal with tons of dead links though for any post older than a couple of years. The Web really seems like it's an awful place. The original dream of hyper links is really dead. It's all about jumping from walled garden to walled garden through tricks and luck.
1. Search "<product name> reddit"
2. AMP page loads
3. "See Reddit in... App or Browser?" banner w/ grayed out page. Tap Browser.
4. Click through the AMP page to the real Reddit page.
5. Second "See Reddit in... App or Browser?" banner w/ grayed out page. Tap Browser.
6. Fail to read the majority of comments, which are hidden by New Reddit.
6. Manually edit URL bar to replace "www." with "old."
7. Read the comments.
Google is having trouble determining the date of reddit posts and is also making their filter rules useless. You try to filter last month, a result says 3 days ago and you click, but reddit says it was 2y ago!
Use the Teddit frontend (https://teddit.net/) and the Redirector extension (https://einaregilsson.com/redirector/) to redirect all Reddit links to Teddit and view the content there. You won't be bothered by pushes for login again.
This in no way absolves their god awful mobile experience but if you're logged into Reddit with the new UI turned off or use the Old Reddit Redirect add-on it's basically as if nothing on Reddit changed since 2011.
I mean there's not nothing you can do. You can still grab the source to old Reddit and it's just a matter of gluing the old UI to the new API like all the 3rd party clients do.
Not trivial by any means but there's a lot of prior art floating around you could pull from.
Sure they could pull API access and be more hostile to scrapers but that's the current escape hatch for people who would have left because of the redesign.
Their mobile website works fine for me, or at least the annoyances are a blessing because I spend less time on it. The one thing I do notice is if you follow a link out of reddit then go back, it shows an error instead of the thread and you have to reload, which loses your place.
That, and don't get me started on their search. Love getting 0 results on initial submit, then reloading the page with the results magically appearing...
Also love being redirected in order to continue reading certain threads
Yeah. I had decent karma or whatever on a ten year old account and got banned for not having an email address. Shrug. Pretty much killed my reddit habit.
Google today limits how much of "the web" (their index) they will let you see.
You are only seeing what they choose to allow you to see.
Google, not the user, determines "relevance" and Google automatically excludes results. in theory this sounds useful. In practice, Google is now limiting max number of results to 200-300 or 500 if you add &filter=0. Retrieving 501 search results for a single search is not allowed. Sorry.
Try a search for some common phrase like "the web". Surely this phrase occurs on more than 500 web pages. Yet Google will limit you to only 231 results. Does that represent the entire www. Then you try "repeat the search with the omitted results included". Google then limits you to 466 results. WTF. What if you searched page titles for some string. Is every string you search going to be found in less than 500 pages.
Google search results today are not representative of the entire www. Not even close.
IMO Google Image search has gone down hill significantly too. Getting full pages of pinterest results was bad enough but now on my phone you have to scroll past multiple screens worth of Google Shopping ads or images tagged with the "Product" icon before you see any actual, organic photos of the thing you're looking for.
Google Images now only returns maybe 1-200 images by default, and if you click show me more, it shows a few hundred. The top results are consistently pinterest garbage, or similar types of spam. Images are consistently diverse from what I've typed in, and results seem curated to show a diverse types of results, which, given the small number of returned relevant images, means I rarely find something that looks like what I'm looking for.
What I've discovered, though, is that what google images now does, is it is only returning a few results. If you click a result, then there's a "new-to-me" "related images" section that shows images similar to THAT image, only in that window. It is in these images that I actually see the results I was expecting to show on the main page.
Its still far worse than google images 10 years ago, but not as bad as the UI makes you believe by looking at the initial results page.
For my money, yandex.com now has the best search results.
Beware though, yandex allows NSFW by default. You have to enable safe/family search. This is the opposite of Google and Bing, so it could get you in trouble at work. Those crazy Russians.
Not just for images. Bing seems to be a lot more closer to what Google search used to be a few years ago. Anything related to torrents or streams is pretty much impossible to find on Google. Any news not spouted from mainstream sources, gone from Google. My default search deck nowadays is a mix of Bing and DuckDuckGo.
it's nice to see my vague worries, bothering me in the back of my head put into words. Also good to know that i'm not crazy for thinking googles results are surpisingly poor and generic. I sometimes get the same feeling using google that i get reading a poorly transalated manual from a Chinese company. Generic, admire the fact some effort has been made, little laugh inside. Except this is google (and ironically, the chinese copywriters are probably using Transalate)
Yandex is CRAZY good. There's a Firefox plugin that lets you reverse image search from the right click menu, and it opens a tab for Google, Bing, Yandex, Baidu and TinEye. Yandex is the king every time.
Indeed, recently I tried Yandex for the first time and the results were way better for everything I had to search in the past 2 months since discovering it.
One I just recently discovered is symbolhound.com. It's nice because it lets you search for characters that google absolutely refuses to treat as search terms. I needed to debug some makefiles and bash scripts the other day (not my strong suit!) and it helped me understand some of the weird syntax I was seeing.
As others have mentioned, Google no longer respecting literal search terms has made it much worse for many types of searches. DDG had been great at this, but sadly has been following in Google's footsteps the past few years.
You too eh? Over the past decade I've gone from loving reddit to actively avoiding it in most cases, but if I'm looking for specific info on a niche topic, adding "reddit" to a search is often the only way I can get real content written by and for real human beings instead of SEO spam.
Sinister twist though: the SEO crowd has gotten wise to this. I've noticed recently that they've started throwing "reddit" into their spam as a random keyword.
I absolutely hated when they did remove it. I was able to sort of mitigate it by adding "forums" and "discussions" to search terms but it is not the same of course.
Amazon search is similarly broken. Try to find an ECC DIMM on Amazon. Everything I try results in dozens of listings for non-ECC DIMMs. Even searching for " ECC" doesn't work all that well. It's sad that in these days of advanced neural networks, we can't get a simple binary attribute matched correctly in search results.
I guess it's location based too (like Google is). Searching for "ECC DIMM" returned a bunch of actual ECC memory chips available to ship overseas on the first page of results on Amazon.com, though I'd still suggest being more specific when searching for ECC RAM myself (eg. RDIMM or UDIMM or...).
I did, and that helped a little, but it's just sad that what should be a solved problem by now isn't. Nearest neighbour search is so not helpful to me in many cases.
There's a common thread in every one of the good sources of search information: user curated content.
Usually: voted on/curated by members who are specialized in some way (reddit, so, hn, wikipedia to a degree).
The fallbacks are editorial sites.
If you make a search engine that focused on indexing user-curated sites + results outside of that that are themselves curated/voted-on by your search engine users (ie, you and others could upvote axios and Amazon as a good source), I think you'd have an interesting model. Basically, take the HN/reddit model to search itself.
This is the essence of what Google implements. Links between sites are a form of user curation, user behavior (clicks, read throughs, bounce rate, etc.) are like inferred votes. The problem is, everything is gameable. Marketers abuse reddit as well, it just isn't quite as profitable as ranking #1 on Google. If a major search engine implemented a voting model, it too would spring an entire industry of "optimizing" for rankings via upvotes.
You can search Reddit from DuckDuckGo by adding '!r' to your search terms. This is how I search most of the time these days for the exact reasons you mention above. DuckDuckGo has many other bang shortcuts like that too, like '!w' for Wikipedia. Actual search engine results are too manipulated, but they make great link aggregators.
Note that "!r <foo>" is different than "site:reddit.com <foo>". The former redirects you to reddit's own search, the latter keeps you on DDG but restricts the results to reddit.com.
I prefer "site:reddit.com" but honestly it's only because I'm used to it. I search in-app on mobile frequently (which I assume is the same as the on-site search) and it's generally pretty good. It usually finds something for me but sometimes using an external search engine works better.
Occasionally I open a private window and try browsing the web on my phone or laptop without an ad blocker. It doesn't take long before the autoplaying videos and Taboola crap makes me close the browser.
Product vs Product -- I hate those generic websites that compare everything to everything. I am absolutely sure that Google can detect and downrank them. Why doesn't it? Maybe most users like those links, and Google is happy to oblige by keeping them at the top? In that case, I suppose, I should blame the mass market taste rather than Google.
I think the future of search is in finding trusted results, such as those upvoted by reddit or HN or SO, assuming the upvoting system is robust enough to withstand spambots etc. The question is whether a big enough fraction of the population actually care to get good search results.
What really astonish me is they sometimes give me straight up malvertising fake domains, like “general-<word>-info. xyz” second or third from topmost these days! wtf.
You are absolutely correct. So many of the top results are basically machine-generated pages, perfectly optimized for search robots, not so much for humans.
1) Bring in tabbed search result like Cuil search used to have. Easily lets you go to similar/variant topics from UX POV and must be a good way for a search engine to learn what people want more.
2) Add a 'dont show this site' option. So they can easily see when people get annoyed with a result and use that as a ranking signal against search terms. E.g. if people press that for a particular KW but not usually Google knows they got that result wrong and if people press it for almost everything you know its a site people dont want to see. And at a personal level its great to custom remove stuff, like when Pinterest was dominating so much a year or so ago.
Oh yeah, try finding honest product comparisons and test. It is spam or "aggregated tests" on shopping sites. latest examples: cameras and tires. Cameras worked reasonably well on youtube (required some diging to get the "influencer" stuff out of the way). Tires required going through various forums. because every single video or review for tires I was looking for was product placement, spam or both. in the end price and availability decided the tire question. sometimes I miss the times when you had magazines catering for a certain domain doing honest tests and reviews. Well, even back then magazines started to prefer Canon over Nikon or the other way round.
My experience is puzzlingly contrary to the unanimously-shared frustration above.
I tried this with a bunch of products ripe for spamming: instapot review, nvidia 3070 review, and Apple Watch review, and the first page results were almost entirely reputable.
I also tried “huggies vs pampers pull ups” and the result is quite a good variety of forum threads, blogs, and reputable articles. “Quasar formation” leads to Wikipedia, but also academic articles and astronomy.com. “How beer is made” is even better quality.
Is most of the web junk? Have I gotten astronomically lucky in not finding it? Are these somehow the exceptions that prove the rule?
(Ironically I think my search terms caused this comment to be flagged as spam..!)
The results are better when the products are more common/famous (Nvidia card, instapot) since there will likely be reviews from The Verge or something that rank highly.
It's less common stuff that tends to give you mostly crap.
I haven't heard of Axios before - how would you characterize it? Looking at the website I can't quite tell if it's an aggregator or how they generate content, and if there is any particular leaning to what they host.
I get better “ Product vs. Product” and “Product Review” results with ddg these days.
Once I took the time to read one of the sites and it’s obvious it’s generated content, and awful at that. I still reading pros and cons of one of the product and the same thing was mentioned under both, just with different wording.
GPT-3 or similar is going to make this way worse.
The only time I still use google (via !g on ddg, what an awesome feature) is to search for local info or in native language. But even here ddg is getting better, especially if I append the name of my city or country.
I tend to do that too. But for product reviews, I am certain that the fake review problem will hit Reddit soon, if it hasn't already. This technique is way too common, and manufacturers are going to start including Reddit in their fake-review spam just like they do with Amazon reviews.
Quora can be good too, my go to is usually "question reddit", and if that is not sufficient then "question quora". Quora is spammy too though, so it is hit or miss, but when it hits, it hits good.
A lot of us Quorans were putting real insight there. Seriously. First few years was amazing.
As they wound down the top writer program, which is really where the "hits good" seeds are, they decided to pay people to ask questions.
Signal to noise ratio began to trend toward unfavorable. But traffic shot right up!
There have been other decisions contributing to the Quora we see today, but the paid questions really had the most impact in my view.
That said, yeah. There is still a lot of high value contributions to Quora. They are just a little harder to find now.
One thing I feel they missed the boat on was the credits system. Early on one could accumulate those and use them as a sort of currency. I "spent" some of my pile asking specific people for insight and it worked out well. I had others do the same with me.
Basically, one could get to a domain or subject expert and the site was still "family", so it would lead to a high value exchange.
For a little while, their text question UI hinted at what a Quora could be. Well connected Quorans could toss a question in via SMS and get solid responses back, sometimes quickly.
(Came down to follower counts, writer status, and a few other things)
These options were not abused much that I could tell and they hinted at a means to query people without draining said people, and or everyone participating having some give and take.
Was a fleeting moment, but one I ponder from time to time.
I would subscribe to a "group" that has the UX and overall dynamics of something like what I just described right.
It's infuriating when Google prefers the documents that ignore some of my keywords even when there are plainly pages that include them all. I get the sense that there's an over-weighting of broad semantic match, to the detriment of lexical match, in their current ML model, whatever form it is now. It's harming quality for a segment of technical users like ourselves but might be "better" in the aggregate over all users, according to at least some of their internal quality metrics.
Back in the 90s, search engines were driven largely by sparse vector representations of the documents such as TF IDF vectors before latent semantic indexing, topic modeling, and other dense vector representations like sentence embeddings entered the fray (not to mention non-content features that use the web graph, click stream, etc.). A lot of NLP applications use a mix of dense and sparse features but it's hard to get the balance right in a way that works for all inputs. Google's pendulum has swung too far in the "dense" direction, as it certainly seems "dense" a lot more often lately!
I don't think the "problem" here is bad search results -- I think the lack of clicks is thanks to search results becoming much richer. The "card" style answers that show up more and more often mean I don't have to actually visit the site where the information originated (which publishers despise, for good reason). Take a common search I made last year as an example, "covid king county" -- I almost never ended up visiting the county's coronavirus dashboard because Google had the graphs I wanted in the search results.
I think it's a popular opinion on HN that Google search sucks, but I just don't agree. I used DDG on all my devices for the better part of last year, and bailed when I noticed by g! usage ratcheting up.
It sucks at searching for exact things, it would often subtly by essentially change the meaning of the search. Like confusing searches for the desired "weight of scooter" with the max "weight of rider" on the scooter. What a shame, it was even a shopping related search, where the money are. They could have led me to a sale.
When people are trying variations on a search it's a clear sign of failure. They should do something about it, like use a verbose natural language interface or select a different strategy for ranking or enable the exact in-depth mode. Apparently the NLP community can do natural language Q&A in papers, but Google can't do it in search.
Other pain points: searching in depth all results on a query, not just the top skim and remembering context between searches.
And Google Assistant itself is too poorly integrated and dumb. Where's the GPT-3 like language skill? So many TPUs what are they doing all day long? They have more text and images than any other entity on the internet, Google bot has been sucking it up for so long (not to mention links, keywords and clicks). It should show in the quality of their AI. It's a shame to have OpenAI steal their thunder like this.
Yeah, the parent and most its children are focusing on something not even related to the article — this is about providing what the user is looking for without the need for more clicks.
While it’s an unpopular opinion on HN, there’s no denying that from a user perspective, that’s only a good thing.
I definitely read this like the two of you. If the information being shown by Google results is what I need, there isn't a need to click further. On top of that, many of us have been "trained" that products like ZScaler are going to block most sites and register a hit with InfoSec. I'm not going to click on bobsfunmainframefacts.com if Google scraped the needed info for me.
Card results might be a good thing overall, but they certainly aren't only good. In many cases, adding the cards takes revenue away from the exact people who collected or created the content to make them possible in the first place. We'll never know how many websites shut down or never got created to begin with given Google's history of crushing the revenue from various sites on a whim.
It can be simultaneously true that Google's gotten better at displaying results pages that fully satisfy the user without the need for a click, and worse at other kinds of queries.
The latter case -- searches which I have to rephrase, or that cause me to give up on the search entirely -- adds to the total number of searches that do not lead to clicks.
It drives me crazy how Google never fails to remove the most unique and yet most important words of my search term. It basically does this 99% of the time now.
> It drives me crazy how Google never fails to remove the most unique and yet most important words of my search term. It basically does this 99% of the time now.
Which is ironic, because one of Google's original innovations was to make a space an AND connector instead of OR (which its competitors used to maximize result counts). Back then, they understood that fewer, more specific results were better.
"Three wrong ideas from computer science" - Joel On Software, August 2000 says:
"when the big Internet search engines like Altavista first came out, they bragged about how they found zillions of results. An Altavista search for Joel on Software yields 1,033,555 pages. This is, of course, useless. The known Internet contains maybe a billion pages. By reducing the search from one billion to one million pages, Altavista has done absolutely nothing for me.
The real problem in searching is how to sort the results. In defense of the computer scientists, this is something nobody even noticed until they starting indexing gigantic corpora the size of the Internet.
But somebody noticed. Larry Page and Sergey Brin over at Google realized that ranking the pages in the right order was more important than grabbing every possible page. Their PageRank algorithm is a great way to sort the zillions of results so that the one you want is probably in the top ten. Indeed, search for Joel on Software on Google and you’ll see that it comes up first. On Altavista, it’s not even on the first five pages, after which I gave up looking for it."
I wonder are they doing this because of a resources issue; do including all the unique words put too much of a hit on their servers and it's way cheaper to give generic results?
I'd prefer to recognise immediately that there are no results & reframe my query that to start to scan down through the results, maybe click into one & start to read only _then_ to realise it doesn't relate, back to Google — "oh, they didn't actually search for what I told them to search for" & then reach the same point they could've given me at the start.
At least there should be a checkbox for "make a best guess when limited results" / "give me my exact search query" (maybe there is somewhere & I've missed this)
> At least there should be a checkbox for "make a best guess when limited results" / "give me my exact search query" (maybe there is somewhere & I've missed this)
There is, one of the drop-downs gives you the option of "all results" or "verbatim".
Google's search index of the web of 2021 sucks. Getting two search results which match the exact terms I got is no longer particularly strong evidence that those were the only pages with those terms, and quite a few times I've had it fail to find pages which do exist and should have matched the search I did.
I find it really obnoxious that if I search on 5-6 terms, Google will return popular results that include 1-2 of the least specific but most popular terms. If I put anything in quotes, it will ask if I really mean something less specific than what I put in quotes! Well yes if we get to a generic enough level, you will return the best relevant results to those generic terms. Who cares if they relate to the specific search I started with?
Google results have gotten dramatically worse over this last decade.
A lot of people say this is because Google is worse at returning the desired search results for various reasons (algos, ads, spam, etc). I've come to suspect it is more deliberate on the part of Google.
I've said it before, so I'll belt out the chant again: "Google doesn't make money from providing the right search results. Google makes money from keeping you searching for the right search results."
This is especially true for the vast majority of people on the internet who do not know there are any (better or worse) alternatives to Google.
And another reason, Google makes money when you click on low-quality machine generated search results, because these pages are almost always monetized with google ads. This causes a conflict of interests: when they improve search, they earn less money.
I’ve seen it: I happened to enter exactly the same query from two separate IPs used by communities without deep cultural connections in between, and only one of them returned the desired results.
That felt like a thick wall of glass separating worlds have suddenly come into my view.
ooh yeah, especially when I am doing some specific database programming stuff and my results come up with amazing perfect resources and a noob's results come up with absolute shit... its bad.
Same here. I can remember how refreshing it was when Google was new and you could trust it to give you results for what you actually searched for instead of what it thought you searched for, like AltaVista and the rest of its competitors were doing at the time.
I guess it's time for something new to do the same thing to Google that it did to those companies years ago.
Other sites still offered old-school website directories, I remember using those for years, well after Google became #1.
I'd much prefer going back to something like that now than bother with Google's approach; depending on the search term, I already know what the top sites are going to be, and I know they won't have what I'm looking for.
> Google now seems to fixate on the most common terms
This feels very true... the majority of my searches go like this:
1. Search with all relevant terms with some parts that should be exact combinations in quotes => nothing or BS results
2. Remove words that Google is fixating on without context => no or seemingly unrelated results without search terms at all
3. Reduce specificity to 2 or 3 words, topic and subtopic (trying to locate context only and search the rest myself) => sometimes ends me on sites where I can then just browse for the result i'm after
4. Worst case reduce to single search term to find huge context sites and manually search myself. Sometimes at this point I just give in and manually navigate to other sites where I know I can manually browse and narrow down context myself and then try to find what I'm looking for... it really feels like a curl web scrape and grep would work better than Google at this point (yes I know google "site:", it's doesn't work properly anymore).
I was kinda thinking the same too until I went back and gave bing and DDG a shot and well, Google is still way better for my searches. I think its more of the seo spam that popping up thats making the results look less useful.
That would be an interesting problem, how do you algorithmically filter out SEO and focus on useful information? A signal vs. noise problem, but on human text, with the added challenge that the noise is trying to outsmart you. Maybe an adversarial ML network with an SEO-generating bot working against you?
Dramatically worse... if that's true, then I wonder how good they were a decade ago. Because every search I've done today produced the result I wanted on the first page of the results, above the fold.
I can't say I've really counted how many days that's true. I can't even say that I really remember what searches I did yesterday or the day before. But if I'm not totally weird and google searches were dramatically better in the past, then they must have produced the desired results every time on almost every day.
Default search might have gotten better, but it also became virtually impossible to do any complex queries. Google is being too smart and it's extremely hard to find old results, foreign results, other meanings of a popular keyword. It's always trying to force you to the same results, whatever you try. Oh, and if it's anything you can buy you get loads of ads followed by spam.
> every search I've done today produced the result I wanted on the first page of the results, above the fold.
Same for me, but that's because I just stopped searching for things that I know will not give me good quality results as they used to be circa pre-2010 back when search tools like +, -, and "" still worked and the results weren't filled with generated SEO texts.
> Dramatically worse... if that's true, then I wonder how good they were a decade ago.
Oh, I forget some people doesn't know why Google enjoy their current position:
Back before 2007 they completely blew competition out of the water:
If something was accessible on the Internet and wasn't behind a noindex spell, Google would find it.
Compared to other search engines that both then and now work more like Google does today it was totally amazing.
Then things started to go sideways:
- first there was: did you mean <something else with similar spelling>? (this was actually user friendly)
- then there was: we didn't find many results for <search term> so we included <related but different search term>, use double quotes to search for "<search term>"
- then there was fuzzing: expnding all my search terms into the unrecognizable unless I double quoted them
- and the latest few years they have also ignored my double quotes
somewhere in between there they messed up the + operator that used to mean "make sure this term is included" as well as ~ that used to mean I wanted Google to fuzz that term.
Sometimes I can get better results by trying to think how my wife would phrase the question, i.e. instead of searching for
- <search terms including a weird mispeling from a dialog box>
I search for
- <why does my computer show search terms including a weird mispeling from a dialog box>
somewhere in between there they messed up the + operator that used to mean "make sure this term is included"
IIRC this one was driven by Google Plus marketing wanting "+" to mean "Go to this page on Google Plus" so they could do some co-marketing thing with +Pepsi.
You can't even compare it anymore, as the web is dramatically different now than it was a decade or so. Just the amount of blogspam that has accumulated in those ten years, all caused by Google's ad-meddling is enough to drown out the minuscule amount of "unique" or "good" content that existed back then or is being currently created.
I'm really starting to prefer DuckDuckGo at this point. There was one time when I thought they would always live in Google's shadow, but in this brave new world of content censorship and commercialization-induced bias, I find that I get noticeably "better" results on the more neutral alternative.
Anything in google that hits the political "twiddler" does not produce useful results. This was in the Zachary Vorheis leak of Google internal documents.
I wish DuckDuckGo actually found the stuff I'm looking for. I've given it a solid try a few times and it's just not as good, and I can't justify spending twice as long as I need to on a search.
Google doesn’t want people to find alternative sources for news, healthcare, and politics. They now force all the mainstream sources to the top, regardless of whether the search terms entered have any relation to the actual results. And this seems to have carried over to other subjects as well, because I can never find obscure pages any more.
This should destroy Google, but for some reason it doesn’t, and it baffles me that they are still surviving in this.
This is has been my experience as well. When I search using very specific terms for a any topic tangentially related to something popular, all I get is the News articles.
I suspect it is because the best sources of information do not serve adds, but news, spam, and social media do.
Same could be said for any site that is smaller in scale than the #1 result.
If I'm looking for a review of a Yeah Yeah Yeahs album, I will see links to Amazon, Pitchfork, Rateyourmusic and other sites well before I find something written by a regular person on their personal blog.
Same goes for products. I will see a page of Amazon links before anything on say, a boutique Ecommerce website with a Shopify backend.
Thank the media for slandering the company, accusing it of perpetrating genocides, political strife, Russian election hacking conspiracy, etc... The reason they do this is being scared of liability, not a top-down desire to control.
I was trying to remember the name of a song. Not an obscure song: "Toe Rag" by The Rifles, a modestly well known indie band. I even remembered a chunk of the lyrics! But I got nothing. Here's let's cheat, and copy and paste exactly two lines from the song.
For me Google "verbatim" is the best way to get focussed results although it's too bad it doesn't allow date ranges. Bing search with appropriate use of guotes, + and - operators and date ranges usually beats non-verbatim Google search, and it can sometimes be better than Google verbatim.
The verbatim mode is exactly what most people here are missing. But, for some reason, it's something that can't be configured - you need to set it in every new search. Three freaking clicks.
I don't understand what goes on corporations. I guess power users aren't a target demographic anymore.
I’m suspecting that they generate results from statistics ahead of time and just serve you the closest cached result, unless it is strictly necessary to do an actual search.
> For me Google "verbatim" is the best way to get focussed results although it's too bad it doesn't allow date ranges.
One word of caution: I use verbatim mode and I regularly (but not always) get Wikipedia clones as my top result. It must be skipping some of their anti-spam filtering.
For coding questions it still works relatively well. For "reviews", or anything remotely tied to commercial interests, it's indeed pure garbage: pages and pages of low quality links.
> Google results have gotten dramatically worse over this last decade. Google now seems to fixate on the most common terms in my query and returns the most generic results for my geographic area. And, it seems like quotes and the old google-fu techniques are just ignored or are no longer functional.
I thought the main issue here was Google directing searchers to its own services and scraping data from non-Google websites to present on its own pages? It's declining search result quality is certainly an issue, but I don't think this link is about that at all.
It really depends on what you are looking for. If you are searching on a term that is used by business you will get mostly shitty content. But if you search for things that make no money, for example. "diff with prosemirror" or "travelling with monitors" you will get very relevant results.
I don't think it has a lot of to do with new search algo. I think it is the opposite, the algo is mostly the same as a decade a ago, but now it is abused by SEO expert. And Google doesn't seem to care yet
yesterday I was trying to figure out how to look up the nameservers of a domain including the vanity/parent nameservers using a tool like nslookup or host, it took me a good 10 rephrasings before i found a stackoverflow post answering the question, and not just a tutorial on how to set up vanity nameservers at X registrar/hosting company
So there needs to be a significant cost in production of the content - youtube for example - want to see how to (a recent example) lay laminate floor, well you need to have a camera and a person and, hell, a room with laminate flooring needed - that's a big cost.
Lowes has made a video, so have dozens of contractors. It's amazing how tool rental firms have not advertised on there ... actually that is a bit weird...
SideNote - for some reason I replied to an email outreach from Rand, and he actually replied back within a few hours, clearly having read my words and come up with a real answer.
I was frankly shocked. Either he works an inhuman amount or he has secretly invented AGI to respond to his emails for him.
Yeah. I wonder what fraction of searches actually end with no clicks, just the requestor bouncing from the search results page. Anecdotally, I suspect that metric is also trending in the wrong direction.
You are 100% correct about the switch to Semantic Search and dense vectors being one of the main reasons. BERT shouldn't be anywhere near my search results.
Google results in the past year or two seem extremely odd to me. I see less results than I used to, I see more of the same sites than I used to, and a lot of content that I try to find that I know I used to be able to, I can't.
Oh, and then there's a bunch of weird links that come up to obvious content generated by AI being peddled. Odd.
Probably a side-effect of SEO[0]. Its in the best interest of the vender, not of the user. I should be banned, forbidden of fought against by google themselves.
And the natural language questions sometimes get stuck. For example, [what's the difference between cauliflower cheese and cauliflower mornay] seems like it should return a page that tells you what the difference is, but it doesn't. It returns a bunch of recipes.
The highlighted responses to suggested natural language questions are comically bad. Answers to different questions, complete failure to apply simple contextualisers like 'today', and treating some random on Quora or some splog as an authoritative source for identifying a fact that has authoritative sources. I know it's difficult to do this stuff without human curation, but it's like a campaign to discourage people from taking NLP seriously.
Don't forget AMP. I accidentally typed google in the url and decided to use it to see how's it going after so many weeks of not using Google. Got redirected to a blank page with AMP's frame which instantly reminded me why I no longer use Google for search.
I got so annoyed with Google that I set my default to duckduckgo and/or Bing, but now I've gotten in the habit of switching back to Google constantly, because once I am directly faced with the quality of the other options, they suck.
I have noticed an eerily accurate prediction of my search term though. Several times, when I hit a bug and attempt to learn more about it, Google makes precise auto-complete suggestions
We are in this situation because people never wanted to pay anything on the web, so we have become the product. That's the price to pay when you want everything for free.
>I'm certain that the switch to Natural Language Processing drove this decline
It's the switch from lexical/syntactic search to semantic search. There are benefits and drawbacks to each and it would be nice for Google to give you the option to try each.
My guess is that you're part of an unprofitable long tail. Why optimize for your use case when they can capture customers that bring more money with far less effort?
I wager that most people don't perceive a decline in search results quality. At least not to an extent where they'd notice and switch to using Bing.
Google Search users aren't Google's customers. Making Google Search useful to its users is not what makes it profitable. It only needs to be mildly better than the competition, so that people use it and can be used for collecting data and serving ads.
So GP may not be "an unprofitable long tail". GP may in fact be "a profitable typical long tail for whom Google Search is frustratingly unhelpful".
I'd be surprised if there was a long tail of customers. There is probably a long tail of searches for most customers.
Google giving up on difficult searches and focusing on showing promoted links, AMP sites, and "rich" result boxes is probably a sound business strategy too, and is unlikely to make users switch search engines this late in the game (Bing/DDG don't seem to do better anyway).
There are a whole host of factors behind this, but I'm certain that the switch to Natural Language Processing / Semantic Search drove this decline.