Digital Analytics is here to stay there is no question. With all the optimization, efficiency and the ability to better target your audience, the appeal is natural and understandable. With this in mind; how do we sure up the data so as to maintain privacy in an increasingly “Peeping Tom” environment?
Today’s Information environment is vast; companies have figured ways to collect data from their users in ways never before believed possible. When, where, how much, frequency, and the list goes on for what is being collected. Devices to chart sleep patterns, location, steps per day, mood when buying based on patters, items related to current purchase based on previously collected data, etc. All this advancement with one real goal; money! No I mean a happier you oops. A Nobile goal on the surface but who is looking at it from any other angle, is another perspective even necessary?
Let’s start with no. No we have no need to be concerned. No one cares about this data, at least not from a malicious standpoint. Even if they did we have randomized it so that users are protected and kept anonymous. In fact all we know is IP address activity, not even who is using it. Or in other cases, yes we know who they are, they created an account but we enforce strict security measures to prevent the malicious use of our user’s data. We would never allow access to our data without going through the proper channels and following regulation.
While this sounds good on the service and it might even be true let’s look at it from another angle. And let’s call this angle “Current Events”. For the past several months we have read articles about a man named Edward Snowden, Articles about the “NSA” National Security Agency, and how they are behaving with our information and who/how they are obtaining it. There are even articles about other countries spying for many years on our data and we are just finding out. http://www.wired.com/threatlevel/2014/02/mask/
We should now look at this phenomenon from another angle. Our data is not private! It is neither safe nor secure! It is being sold to the highest bidder with whatever motives suit the buyer as it is not in the sells best interest to ask. And most people don’t understand what that means in breadth or depth.
This is not an environment of privacy! In fact most people, especially the elderly, have no idea they are consenting to the whoring of their data in the first place (even if they did hit the “I agree” button at the end of their 5,000th, 200page consent form). Hopefully no atrocities are going to be committed by the data being collected/used this way, at least we all hope there aren’t (makes you wonder what the NSA told AT&T to leave out of their report, http://www.wired.com/threatlevel/2014/02/ma-bell-non-transparency/ ). This begs the question. What might happen if the data was to be used maliciously? Who would be accountable?
With statements like this from Google “users have no legitimate expectation of privacy” what does that mean for the future. http://www.digitaltrends.com/web/google-gmail-users-have-no-expectation-of-privacy/ If I want to keep my information private how do I do this and still participate in this wave of future with all of its convince and pizazz? Is it reasonable for a non IT expert to go off the grid in hopes to keep some privacy? Do I forego all the sales and deals associated with this data collection movement? Or do we look at protecting the users privacy TRULY while still collecting the necessary data? There seem to be a few companies who have this mind set and users appreciate it.
It is important for the Data Analytics of the future to keep the trust of its data pool. There is very much a symbiotic relationship between the two and the relationship needs to remain pure. No more NSA “gifts”, no more selling your user data unless completely anonymous (it’s not anonymous if our email is getting spammed or our phones are being called by the company that bought our data). Protect the privacy of your users and Karma will make sure it comes back to you!
Course blog for Digital Analytics course at the University of Utah
Showing posts with label google. Show all posts
Showing posts with label google. Show all posts
Friday, February 21, 2014
Tuesday, February 18, 2014
The Piracy of Your Privacy
Privacy on the Internet has become non-existent in the
digital age. Your digital footprint is
tracked and analyzed by various marketing firms all vying for your
attention. They have found ways to
weasel into your computer, tablet, phone, and other Internet-connected
devices. They know every website you
visit, how long you were there, what you did there, how long you were there, how
you got there, and even where you are signing in from. Nothing is secret anymore, so be mindful of
the information that you send to cyberspace.
Web browsers like
Google Chrome have found a sneaky way to track all of your digital movements
across all devices. They ask you to sign
in. This act allows Google to keep a
history of all of your movements on all devices. The “Stay Signed In” checkbox on the login
page keeps this action further from the user’s mind because they do not have to
sign in each time they open Chrome.

Chrome is also able to combine all of your Google
services. Like the slogan on the page
says, “One account. All of Google.” This means that they know of all of the
Google Apps you are using and how you are using them.
According to PC Mag, the Federal Trade Commission (FTC)
issued an advisory on Web tracking of customers. Chrome was one of many browsers to enable
features that protect their users from advertisers. They also have an Incognito Mode that allows
users to move around the Web without leaving a footprint behind.1 While these options allow a user their
privacy, most users are not aware that they exist or understand their
purpose.
Further actions
have been taken in the industry with the creation of websites such as donottrack.us.2
which allows users to opt out of tracking by websites that they do not
personally visit. However, a Wall Street
Journal Investigation found that the top 50 websites in the United States
install an average of 64 individual trackers to visitor computers.3
These trackers are often fast and invisible.
Even though this article was written in 2010, it is not hard to imagine
that tracking software changes as fast as its governing laws do.
Why do companies want to stay ahead of governance? NBC News reports that is because online data
has been transformed into a lucrative marketplace. Information is being sold to the highest
bidder. This information includes names,
addresses, phone numbers, and credit card numbers, which are being traded out
in the open.4 This continues
to be a growing concern among consumers.
According to a study performed by the Pew Research Center, fifty percent
of Internet users are worried about the information about them online, compared
to thirty three percent in 2009. Of
those surveyed, eighty six percent of people have tried at least one technique
to hide their activity online or avoid being tracked.5 The study also found that 68 percent of
people feel that the law is insufficient to protect their privacy.
Causes of Privacy Loss:
- People are providing too much information on social network sites and email.
- E-Commerce sites are better able to hide their tracking
- Security breaches of E-Commerce websites
- Password Protection – California feels that security breach laws should also protect passwords, usernames, and security questions. Hackers for popular social media websites such as Facebook, Twitter, Google, and Yahoo stole nearly two million usernames and passwords.
- Do Not Track – While current legislation does not ban tracking, the new legislation requires companies to disclose how they comply with requests from Internet users who ask not to be tracked.
- The Teen “Eraser” Law – Requires all website and mobile app operators to provide a way for those under 18 to delete a posting or photo. The reasoning is to protect those under 18 from publically sharing “ill-advised pictures or messages”.
Works Cited
1. Muchmore,
Michael. "Google Chrome 31." PCMAG.
PC Mag, 11 Dec. 2013. Web. 18 Feb. 2014.
2. Mayer,
Jonathan, and Arvind Narayanan. "Do Not Track." - Universal Web Tracking Opt Out.
Center for Internet and Society, n.d. Web. 16 Feb. 2014.
3. Butler,
Christopher. "Unlimited vs. Limited Web Tracking." Tracking Best Practices.
Newfangled, 01 Sept. 2010. Web. 18 Feb. 2014.
4. Sullivan,
Bob. "Online Privacy Fears Are Real." Msnbc.com. NBC News, 6 Dec.
2013. Web. 18 Feb. 2014.
5. Flaherty,
Anne. "Study Finds Online Privacy Concerns on the Rise." Yahoo! News. Yahoo!, 05 Sept.
2013. Web. 18 Feb. 2014.
6. Prah,
Pamela M. "Target's Data Breach Highlights State Role in Privacy." USA Today. Gannett, 10 Feb.
2014. Web. 17 Feb. 2014.
Monday, February 17, 2014
Browser Cookies 2.0 - Google AdID - The Holy Grail of user information
Browser Cookies?!?!
Tania recently posted a blog entry explaining what web cookies are, how they work, and how to manage them. At a basic level cookies allow information to be aggregated about you and your browsing habits. This helps content providers (but mostly marketers) tailer a more relevant experience for users but also raises privacy and security concerns as these shadow profiles are enriched over time.
Browser cookies come in two varieties, first party and third party. [1] The first party variety are generally the more benign of the two and only aggregating data in the current site where you are visiting. Third party cookies however, go where you go. They can collect product preferences from Pinterest, employment information from LinkedIn, and just about anything else from social media. This cocktail of information can be powerful for marketers and potentially dangerous in the wrong hands.As users, we tend to leverage multiple browsers, tabs, and even platforms (phone, tablet, desktop, e-reader, gaming systems, etc) to facilitate our data and connectivity addictions. This makes it more difficult for the third party cookies to get the "full" story. Users can also block and delete cookies (see "Do Not Track"). These weaknesses have lead to companies like Google to create solutions to stitch some of these data holes together (like linking all your Crome profiles together, platform agnostic).
Holy Grail?!?!
It would be naive to think that companies would stop short of the 'Holy Grail' of user information. Google is already working on an alternative to cookies called AdID. Essentially, the search giant is building an "anonymous identifier for advertising, or AdID, that would replace third-party cookies as the way advertisers track people's internet browsing activity for marketing purposes." [2]
There are interesting implications of this development [3] but its important to note that the devil is in the details, all of which are speculative at this stage of development.
The Good:
- The AdID is intended to be an anonymous
- Ownership of the AdID is consolidated which can allow for better governance. This is a slight improvement over the smorgasbord of third parties tracking user traffic with almost no governance.
- Users will theoretically be given tdiscretion over how much data (or which data) is aggregated
- There might be a "private browsing option"
The Bad:
- Although intended to be anonymous, it has a huge potential to be linked to something not anonymous like an email or G+ profile
- Tracking is tracking. The information can be willingly sold, lawfully subpoenaed, or wrongfully stolen
- If you are a company tracking users with cookies, future browsers like Crome, may auto-block (verses the current opt-in) cookies. Your only option will be to pay the big boys for that information.
The bottom line
Should we hate Google for doing this to us? It's hard to blame them for wanting to profit off the information that they serve us for free. Besides; Microsoft, Apple, and FaceBook are in the game too. [4] Users need to take more ownership of their personal information and activities and choose when and how it can be used, otherwise, someone else will.=========================================================
[1] http://adage.com/article/digital/google-s-plan-replace-cookies/244302/
[2] http://www.engadget.com/2013/09/18/google-adid-an-anonymous-identifier-advertising-cookie/
[3] http://www.digitaltrends.com/opinion/what-is-google-adid-and-how-will-it-replace-browser-cookies/
[4] http://online.wsj.com/news/articles/SB10001424052702304682504579157780178992984
Labels:
AdID,
Browser Cookies,
Cookies,
Crome Browser,
Facebook,
google,
Online Privacy,
third party cookies,
Web Analytics,
Web browser,
Web Cookies
An Update on Do Not Track and Privacy
In his January 2013 post to the
Digital Analytics - University of Utah course blog, McCall Lewis wrote about
the ongoing debate surrounding online consumer privacy and efforts towards a
standard for "Do Not Track" [i].
McCall correctly stated that 2013 would be the year that these issues
came to the consciousness of the consumer at-large. In this post, I intend to
explore the ongoing saga of Digital Privacy and how consumers and online
entities are reacting.
Digital Privacy in
2013, In a Nutshell (Help! I'm in in a
nutshell!)
Beyond allegations of government
spying, other news events triggered a growing concern for digital privacy
amongst citizens. It was revealed that
Google, despite their unofficial corporate motto being "Don't be
evil", was collecting and storing data on WiFi networks while driving the
avenues and boulevards in their mapping vehicles [iv]. Inadvertent or not, this revelation made big
headlines in the year of Digital Privacy concerns. Beyond government and corporate spying, there
were a number of black-hat news stories as well. Major retailers such as Target and Neiman
Marcus were victims large-scale data breaches in which personal and credit card
information were stolen from their servers.
While nothing connected to a network is ever totally secure, some of the
details surrounding these breaches made it clear that retailers were not doing
everything that they could to protect this sensitive data. In this particular
case, the suspected security snafu source was an HVAC contractor that was given
the keys to the castle, which were thusly compromised [v].
Current Sentiment
Not surprisingly, there have been
numerous studies trying to suss out what the consumer reaction is to all of
this. A University of Vienna focused on the act of "Virtual Identity
Suicide" within the online social networking site Facebook. The single biggest cause for this phenomenon,
where a user deletes as much of their content as possible before permanently
locking themselves out of their account, were concerns over privacy. Among users studied, over 48% expressed this
viewpoint [vi]. It turns out, they have a right to be concerned. Austrian law
student Max Schrems found out, in 2010, that Facebook had over 1,200 pages of
data on him alone. This included data that
he had never been supplied, but had been linked to him through his friends
contact list. As big-data analytics gets
more powerful, this could translate into an enormous amount of personal
information being available to online companies [vii].
Using information like Facebook
collects, identification of protected classes is not only possible, but on the
verge of child's play. The Center for
Digital Democracy is making efforts to address its concerns to the FTC. They state that technologies such as
hyper-local targeting, geo-fencing, and cross-platform targeting will allow for
rampant discrimination. The sorts of
questions that it is illegal for employers to ask (age, marital status, sexual
orientation) will become easily attainable information [viii].
TrustE, a digital privacy
management company, conducted a study regarding consumer opinions about Online
Behavioral Advertising recently [ix].
They found that 69% of internet users understood the value trade-off of
online ads versus free content, but only 26% are willing to actually accept the
same. It seems as though most internet
users feel powerless in the process that they need to just accept what is
offered. The study also showed that 62%
of users would be more willing to do business with a company that allowed them
to opt-out of targeting.
Ongoing Efforts
The WC3 is spearheading a Do-Not-Track
and privacy working group, but things are not going as well as could be
hoped. One of the biggest internet
watchdog and lobbying organizations, the EFF (Electronic Frontier Foundation),
has lost confidence in the group [x].
They have directly stated that if the group continues in the direction
that it is currently headed, that they may be forced to drop out. Another watchdog group has a similar
stance. Jeffrey Chester, of the Center
for Digital Democracy as called the efforts of the group "a farce". It appears as though the group cannot even
get the definition of tracking nailed down.
Are they concerned with 1st party cookies, 3rd party cookies, or other
methods of data collection? Original
efforts in the Do-Not-Track space only targeted 3rd party cookie based ads,
providing a guise of privacy to the relatively uneducated user.
It is unclear what the future may
bring in terms of digital privacy, but it is doubtful that it will continue to
be as unregulated as it currently is.
The European Union is enacting tough laws, requiring explicit consent in
some areas, rather than the arguably implicit consent given by endless EULAs
and TOCs that no one actually reads. If
the FTC gets involved in the United States, things are likely to change.
[i] Lewis, McCall ‘The “Do Not Track” Debate’ Digital
Analytics – University of Utah, January 26, 2013.
http://dauofu.blogspot.com/2013/01/the-do-not-track-debate.html
[ii] Wikipedia contributors, "Edward Snowden,"
Wikipedia, The Free Encyclopedia, http://en.wikipedia.org/w/index.php?title=Edward_Snowden&oldid=595457266
(accessed February 5, 2014).
[iii] Wikipedia contributors, "MUSCULAR (surveillance
program)," Wikipedia, The Free Encyclopedia,
http://en.wikipedia.org/w/index.php?title=MUSCULAR_(surveillance_program)&oldid=595197410
(accessed February 5, 2014).
[iv] “Street View: Google given 35 days to delete wi-fi
data”, from BBC News: Technology, June 21, 2013.
http://www.bbc.co.uk/news/technology-23002166
[v] Feinberg, Ashley “Last Month's Massive Target Hack Was
the Heating Guy's Fault” Gizmodo, February 5, 2014.
http://gizmodo.com/last-months-massive-target-hack-was-the-heating-guys-1516926877
[vi] Munson, Lee “Half of Facebook-quitters leave over
privacy concerns” NakedSecurity, September 18, 2013.
http://nakedsecurity.sophos.com/2013/09/18/half-of-facebook-quitters-leave-over-privacy-concerns/
[vii] Solon, Olivia “How much data did Facebook have on one
man? 1,200 pages of data in 57 categories” Wired.co.uk, December 28, 2012.
http://www.wired.co.uk/magazine/archive/2012/12/start/privacy-versus-facebook
[viii] Submitted by demedia, “CDD Calls on FTC to Protect
Privacy in today's Hyper-local, geo-targeting, cross-platform, Big Data
Era/Warns of Discriminatory Practices with mobile device tracking”, Center for
Digital Democracy, February 6, 2014.
http://www.democraticmedia.org/cdd-calls-ftc-protect-privacy-todays-hyper-local-geo-targeting-cross-platform-big-data-erawarns-disc
[ix] Deasy, Dave “TRUSTe Study Reveals Increased
Transparency and Privacy Controls Produce More Positive Feelings about OBA”
TrustE Blog, September 19, 2013.
https://www.truste.com/blog/2013/09/19/truste-study-reveals-increased-transparency-and-privacy-controls-produce-more-positive-feelings-about-oba/
[x] Fung, Brian “The Internet’s best hope for a Do Not Track
standard is falling apart. Here’s why.” The Washington Post Online, The Switch,
October 11, 2013.
http://www.washingtonpost.com/blogs/the-switch/wp/2013/10/11/the-internets-best-hope-for-a-do-not-track-standard-is-falling-apart-heres-why/
Labels:
Digital Analytics,
Do-Not-Track,
EFF,
google,
Privacy,
Target
Sunday, February 16, 2014
Fighting the Secure Search Battle
Digital marketers have become accustomed to no longer being
able to receive complete search term data. Google became the first to secure
their search data with their announcement that all signed-in users would be
secured in October of 2011. Now more than two years later, Yahoo has announced that they will follow Google’s lead and make Yahoo searches secure, causing
digital marketers to once again scramble for ways to gain access to their
keywords.
“Not Provided” by
Google
The Google empire completed the full transition in September
of 2013 as they confirmed that all searches would be secure by default. What this means is that instead of seeing the
individual keyword or searched term in your analytics tool, you will see the
phrase “not provided.” Marketers will
know that a search happened but be unaware of the exact term used to bring the
visitor to the site.
No Referrer with
Yahoo
Yahoo is planning to fully shift to secure searches by March
31st, 2014. Like Google, all searches will be done through a secure
server. Unlike Google, there will be no “not
provided” term in your analytics. Yahoo will not be sharing any information at
all. All data driven from Yahoo will
appear as though the visitor came directly to the site. The No Referrer policy
will make it difficult for the analysts and digital marketer to differentiate between a Yahoo
search and a direct visit.
Optional Secure
Search at Bing
Bing continues to be that friend not willing to rock the
boat. Bing still wants your friendship.
In January, Bing officially launched secure search though it has been
turned off by default. Digital marketers
will continue to receive their Bing keywords for the time being.
If or when the user decides to turn on secure search with Bing, the data will appear similar to Yahoo's with a No Referrer policy.
Tips on How to Fight
the Secure Search Battle
- Use Bing. Even though Bing may only represent a small portion of your organic search volume, it can provide insight into visitor behavior and give you search term data.
- Google Analytics may not provide specific keyword info, but Google Webmaster tools may still provide what you need. Experiment with the Search Queries report and you may be pleasantly surprised with what you find!
- Use Google Adwords. Though Google claims that it has transitioned to secure search for security purposes, the keyword data can still be bought through Google Adwords.
- Use site searches to understand what users are searching for on your site. You can receive valuable insights on potential keywords by tracking what users are searching for within your site.
- Don’t worry and be patient. Mastermind search marketers will find ways to get our keywords back. Continue to stay updated with industry news and notes as valuable information is released and tools are developed.
Reference Links
- http://marketingland.com/yahoo-makes-secure-search-default-71308
- http://blogs.adobe.com/digitalmarketing/analytics/how-googles-expanded-search-encryption-impacts-adobe-analytics/
- http://blogs.adobe.com/digitalmarketing/analytics/yahoos-secure-search-impacts-adobe-analytics-data/
- http://cyrusshepard.com/7-fantastic-seo-tips-for-googles-not-provided-keywords/
- http://moz.com/blog/easing-the-pain-of-google-keyword-not-provided
Wednesday, February 13, 2013
Bringing Traffic Your Way
You just created that awesome blog site, and your family and
friends are proud of you and your work. They come to visit the site a few times a week, reading and commenting on your posts; they actually let you know how cool you are for writing these posts!
But, you want to tell the world about your blog. You want to start driving
traffic your way! Well, I’m going to help you with some cool ideas so your blog
can get your first 1 million hits.
Social Media Working for You
You did it! Your blog is looking pretty awesome. You have
created a cool brand and now you’re ready to take it to the next level. Here’s
where social media can help immensely. Go to Facebook, Twitter, LinkedIn and
Google+ to create your precious brand personal accounts. Even if you believe
your blog won’t be a success right away, is better to avoid headaches by
creating your own accounts on every social networking site.
Now that you have created accounts on these, it’s time to start sharing every post to them. Make sure to
invite friends to like or join your name so they receive notifications and or
emails every time you post a new blog.
Using Analytics to Track Everything
Install Analytics Tools on your blog and let it collect
data. Tools like Google Analytics and Adobe Site Catalyst can be installed on
the back-end of your blog. These tools track everything from new visits, to time
spent on pages. Analytics can also track, where your visits are coming from, if
the visitor landed on your blog by an organic search(1) or by a paid keyword(2).
These tools can be configured to your liking and they provide detailed
information in reports that can be customized and delivered automatically.
If you have been following this post from the beginning, you
should have a Google+ account. That’s all you need to create a Google Analytics
account. The interface is very simple and Google provide great documentation
for it. You can also take a look at the Universal Analytics brand new tool from
Google in this post. Here's a great post about how to use Google Analytics.
Study Your Audience and Write Accordingly to What They Want
Finish your blog posts with specific questions; you’re going to start knowing your visitors based on the answers you get.
You can also try sending surveys via email or post these
surveys along with your blog post. This is a quick way to understand what your
visitors want. You can start writing about what your visitors wants once you
get the necessary data about them.
Again, use social media to interact with your users. Learn
from them and become their friend. Participate in communities your visitors
visit frequently and learn as much as you can from them.
Use Great Designs on Your Website
Make sure you visit different blog sites to see and learn
the trends used. Stay away from using bright colors that can cause temporal
blindness on your audience, and use one font type for your posts. Use pictures
and videos related to what you’re writing about and listen to feedback from
your visitors.
Reply to Comments
Interacting with your audience is just another way to
increase traffic to your site. Make sure
you interact with them when they leave a comment. Even if it’s a simple “thank
you,” from your part; they will be grateful and come back to visit your blog.
Replying to the audience comments creates a great discussion environment. Try
to avoid religion or political comments unless your blog was created for these
purposes.
Conclusion
We talked about using social media to interact with people and expand your blog to different audiences. We also talked about creating a Google Analytics account and installing these tools in the back-end to learn about your audience. Along with that, we took a look at studying your audience, using great designs and interacting with your visitors. I hope I was able to help you and my words encourage you to start writing about the things you really love.One more thing…
Keep posting and don’t be discouraged if you had a few visitors this week. Don’t give up and be consistent.Do you have any ideas on how to start bringing blog traffic your way?
References
Seomoz.org
dauofu.blogspot.com
Google.com/Analytics
http://en.wikipedia.org/wiki/Keyword_advertising
Tuesday, February 12, 2013
Universal Analytics - One Step Closer to Becoming Omniscient!
![]() |
| Universal Analytics |
Google Analytics < Universal Analytics?
Over the past few weeks we have analyzed, examined, studied, pontificated, and expounded upon all of the benefits (and potential issues 1 ) of using Google Analytics. To add fuel to that fire Google released a new tool called Universal Analytics. With UA, we have at our finger tips a variety of new possibilities in the analytics space, specifically:
- Use the Measurement Protocol to integrate data across multiple devices and platforms.
- The Measurement Protocol introduced by UA lets you collect and send incoming data from any device to your Analytics account, so you can track more than just websites. Leverage the new analytics.js code and new developer reference libraries from the Measurement Protocol to see how users interact with all of your devices — smartphones, tablets, game consoles, and even digital appliances.
- Improve lead generation: Sync offline and online data.
- With UA, you can track data from all your online and offline customer contact points, like marketing campaigns, sales calls, and store visits, so you can discover relationships between the channels that drive conversions. Because UA is primarily an innovation in data collection methods, there are no new reports showing cross-device data.
- Define your own dimensions & custom metrics.
- Custom dimensions and custom metrics are like default dimensions and metrics in your Analytics account, except you create them yourself. Use them to collect data that Google Analytics doesn’t automatically track.
- Understand how well your mobile apps perform.
- Mobile App Analytics captures mobile app-specific usage data and integrates it with your Google Analytics account, where you can reapply your knowledge of web analytics to dedicated app reports. (Currently in beta.) 2
Saturday, February 9, 2013
PageRank explained
Today, I thought I'd post about something near and dear to my heart: math. When I was a senior at BYU (insert obligatory boos) studying
numerical analysis, for one of my classes I wrote a paper about the PageRank
algorithm. Seeing as this is a web
analytics class and that a big part of web analytics these days is search
engine optimization, I thought I’d revisit the topic. This time, though, I’ll do it in a way that
is a lot simpler, involves less mathematical proofs, and I hope is less boring.
For those of you that don’t know what PageRank is, before
Google adopted the “whoever gives the most money to Google wins” algorithm,
Larry Page and Sergey Brin from Stanford developed a way to in essence let the
internet itself determine the relative importance of the pages that it
contains. In the algorithm each member
(page) of groups of hyperlinked documents (aka the internet) is assigned a
weight based on the number of hyperlinks to it from other pages. So, a page with a lot of links to it has a
higher rank than a page with only a few links to it.
How it works
Suppose we have an internet with
4 pages: A, B, C, and D with links to each other as illustrated below. In this case, A has one link to it from D; B
has a link to it from A, C, and D; C has one link to it from D; and D has 2
links to it, one from A and one from B.
Every time a page links to another page, it transfers a portion of its
“rank” to the page that it links to. So,
D has a link to A, B, and C, so it transfers a third of its rank to A, a third
to B, and a third to C.
So in our example, the ranks of each page are represented by the following equations:
Or, using matrix notation, it is the
solution to the system of equations below.
Those of you who are mathematicians will
notice that the PageRank values are an eigenvector of the matrix of link
weights. In our case, we want the one where the sum of all the ranks is 1. So, for our model A=0.13, B=0.33, C=0.13, and
D=0.4.
So, what exactly does this rank
mean? One interpretation is that that it
is the probability that after following links for a long time you’ll end up on
that particular page (If you try this on the real internet, you’ll likely
either end up looking at Wikipedia or porn).
This is a simplified version of
PageRank. The actual algorithm is a bit
different to take into account that not all pages have outbound links, people
don’t just follow links all day when they surf the internet, etc., but this is
essentially how it works.
What it means for your site
Knowing all this, what does this mean for
your site? Well, the first and most
obvious thing is that the more links to your page, the better. You might be thinking “great, I can just go
out and plaster links to my site all over message boards, blogs, Facebook,
Twitter, etc. to increase its PageRank” or “OK, I can just pay people to put
links to my site on theirs.”
Unfortunately, most message boards, blogs, etc. use the "nofollow" tag
which tells Google not to include these links in PageRank calculation to
prevent this kind of spamming. Also,
Google has specifically cautioned against selling links to increase PageRank. If they catch people doing this, their links
are excluded from calculations [1]. For
this reason, Google has advised using the "nofollow" tag on sponsored links.
Also, take note of where links to your
page are coming from. Remember that when
a page links to yours, it transfers a portion of its PageRank to it. A link from a big, important site is worth a
lot more than a bunch of links from small, obscure sites.
Saturday, January 26, 2013
What does Facebook’s Graph Search mean to analytics?
What does Facebook’s Graph Search mean to analytics?
Paid search has been dominated by Google so far. The company holds enormous amounts of search data that it is able to sell as part of its advertising packages. On top of this, Google runs its free web analytics engine that integrates with its search platform to deliver even better ad placements and results. Advertisers are able to see what people do not just on Google, YouTube, and in Gmail, but also take advantage of Google’s reach across third party web sites in its current ad presence. Google can help to build intimate customer profiles based on these browsing habits, and turns that data into very, very accurate advertising placements based on users’ web history. (Salesforce is a good demonstrator that customer profiles are worth big money) Power plays to fight Google’s foothold in analytics data, search, and search advertising, even those made by Microsoft, have fallen flat so far.
The massive volume behind Facebook’s user base puts the company in a theoretical power position to leverage personal information. The site holds an extensive network of information and social connections that it has never properly leveraged in the past. Running a search had previously only turned up sporadic friends and few relevant results, yet Facebook has seemed relatively uninterested in reworking its data to improve search results or ad revenues.
Press reaction was heated after the publicity surrounding Facebook’s public launch. With 1 in 6 people on the planet (a substantial ratio of whom are active participants in the global economy, rather than an isolated one, which in turn could drive ad sales) on the service there is a wealth of data to be examined and mined on the back end. People put fairly intimate public profiles on the site, build photo libraries with family experiences, make friends, and actively communicate on Facebook’s servers. All of this is fair game for search analysis, and can allow the site to suggest search results that a friend likes. Just as importantly, it is possible to filter out results based on what friends don’t like, or customize results based on other users who also like ‘Game of Thrones’ and Mexican food in the San Francisco area (Facebook’s chosen example).
This information is predicted to have an excellent impact on analytics. Facebook should be able to use customer profiles to deliver extremely targeted results. This should result in excellent placement for things like restaurants or attractions, and should allow Facebook to charge a premium for ad results because friends and friends of friends can be accurately profild and targeted.
What will the results consist of?
Zuckerberg’s stated goal is that:
This is a brilliant concept with a huge flaw. This is a massive network. The data is overwhelms belief. Yet it only consists of data within Facebook. While most news articles fawn over the idea that Facebook will be able to rival Google, the engine currently relies upon core results results that are billed from Bing. Microsoft was in the news for Bing’s results when it first launched, yet it was not the most favorable reporting:
Suddenly the issue becomes clearer; Google has been collecting information on its users for longer than Zuckerberg has been programming. The wealth of data contained in their servers extends across the web; the site runs a mapping service that contains location data on every business in the world. The service is supplemented by a review engine. Search users are tracked through YouTube and Gmail profiles to directly monitor viewing habits and conversations. AdSense, DoubleClick, and any site they connect to can trace exactly what users are seeing in enormous portions of the web. Google may not be directly serving the page views on these sites to boost its comScore rating but it certainly benefits from the user data it collects in each and every interaction. It’s worth going deeper here to see where benefits come from.
There is a big component here that is another problem; companies have been buying ‘likes’ on Facebook for years and have turned these trillions of impressions into a garbage pile. If a company bought a ‘like’ so a user would enter a raffle, is it valuable for this or any other company to re-pay to advertise to this person’s friends or to friends of their friends? Facebook has poisoned its ‘likes’ data with bad data through extensive advertising sales, which could make useful analytics data very, very difficult to leverage for beneficial search results.
Why is a customer profile important in search?
Any field a user enters information into on their own should be valuable to shaping results. Case in point; a frequent early complaint about online dating sites like eHarmony was that they would pair vegetarians to hunters. Modern user habits that are monitored and controlled through analytics should actively eliminate these types of mismatches. The concept is simple on the surface but becomes difficult in aggregate. It could be valuable to take everyone who says in their Facebook profiles that they like Tibetan prayer rugs, own a dog, and drive a Subaru, and then list advertisements for Purina’s new vegan dog food as a relevant search result. If this user were shown advertisements for a Christian Singles dating service it may not be a valuable use of the space.
Facebook holds a huge amount of this cross-referential data and has built extensive profiles on all of its users. They also collect information from users who ‘check in’ at events and locations, which can help to build a large map of businesses around their individual business fan pages (that in turn contain huge amounts of business data). This data has shortcomings though; roughly 1% of Facebook users are estimated to interact with advertising. On top of this, almost half of Facebook users say they find Facebook’s interface boring and minimize Facebook interactions. (CNN) This is backed up by the numbers from my last post; 85% of posts are interacted with by less than 5% of users. (Business Insider) Suddenly Facebook’s enormous market advantage could become a huge liability as it depends on the habits of a minority of users.
Mind the gap
The concept that most social users are ‘lurkers’ is backed by hard data, (Wikipedia, Reddit) and it’s actually pretty cool to think about. Facebook has a problem when people don’t fill in their profiles. You may be friends with Pat and Pat may like base jumping, but you could be an avid botany lover who never actually filled in your ‘activities’ section. When Facebook suggests that you graphically line up because you’re friends, they aren’t getting the same type of information that Google has on all of the searches you’ve run in the past.
Instead of being able to look at your history and say ‘You don’t want to see an ad about parachute accessories, you should see an ad about a local gardening store,’ the best Facebook can do is show you what Pat likes. The idea of getting a contextual search answer is great, but analytics need to have more clean data than Facebook can access if they are going to deliver good results.
This is where Facebook holds a great advantage it doesn’t seem to like; the company is ripe for partnerships. The data Facebook holds on users could build phenomenal success on top of Bing’s search if it can leverage all of Bing’s resources. While the information and results are stuck to what’s inside of Facebook there are limitations here. The question is whether Facebook will truly partner with Bing to improve its analytical reach and build a foundation for search results so it can have an excellent set of data to draw from. Until the response engine is broadened and more information is gathered on lurkers, Facebook is going to be playing the search game with a short deck.
http://newsroom.fb.com/News/562/Introducing-Graph-Search-Beta
http://www.businessinsider.com/facebooks-new-search-engine-2013-1
http://www.businessweek.com/articles/2013-01-15/facebook-radically-revamps-its-search-engine
http://www.slate.com/blogs/future_tense/2013/01/15/facebook_announcement_graph_search_engine_is_like_a_personalized_google.html
http://searchengineland.com/facebook-search-not-google-search-145124
http://www.wired.com/business/2013/01/facebook-event/
http://www.huffingtonpost.com/2013/01/15/facebook-graph-search_n_2480624.html
http://www.forbes.com/sites/lisaquast/2012/04/23/your-social-media-profile-could-make-or-break-your-next-job-opportunity/
Paid search has been dominated by Google so far. The company holds enormous amounts of search data that it is able to sell as part of its advertising packages. On top of this, Google runs its free web analytics engine that integrates with its search platform to deliver even better ad placements and results. Advertisers are able to see what people do not just on Google, YouTube, and in Gmail, but also take advantage of Google’s reach across third party web sites in its current ad presence. Google can help to build intimate customer profiles based on these browsing habits, and turns that data into very, very accurate advertising placements based on users’ web history. (Salesforce is a good demonstrator that customer profiles are worth big money) Power plays to fight Google’s foothold in analytics data, search, and search advertising, even those made by Microsoft, have fallen flat so far.
The massive volume behind Facebook’s user base puts the company in a theoretical power position to leverage personal information. The site holds an extensive network of information and social connections that it has never properly leveraged in the past. Running a search had previously only turned up sporadic friends and few relevant results, yet Facebook has seemed relatively uninterested in reworking its data to improve search results or ad revenues.
Press reaction was heated after the publicity surrounding Facebook’s public launch. With 1 in 6 people on the planet (a substantial ratio of whom are active participants in the global economy, rather than an isolated one, which in turn could drive ad sales) on the service there is a wealth of data to be examined and mined on the back end. People put fairly intimate public profiles on the site, build photo libraries with family experiences, make friends, and actively communicate on Facebook’s servers. All of this is fair game for search analysis, and can allow the site to suggest search results that a friend likes. Just as importantly, it is possible to filter out results based on what friends don’t like, or customize results based on other users who also like ‘Game of Thrones’ and Mexican food in the San Francisco area (Facebook’s chosen example).
This information is predicted to have an excellent impact on analytics. Facebook should be able to use customer profiles to deliver extremely targeted results. This should result in excellent placement for things like restaurants or attractions, and should allow Facebook to charge a premium for ad results because friends and friends of friends can be accurately profild and targeted.
What will the results consist of?
Zuckerberg’s stated goal is that:
“We are not indexing the Web,” Zuckerberg said. “We are indexing our map of the [social] graph.” Users can navigate through the 240 billion photos on the network, the trillions of user “likes,” and the connections between users. (Bloomberg)
This is a brilliant concept with a huge flaw. This is a massive network. The data is overwhelms belief. Yet it only consists of data within Facebook. While most news articles fawn over the idea that Facebook will be able to rival Google, the engine currently relies upon core results results that are billed from Bing. Microsoft was in the news for Bing’s results when it first launched, yet it was not the most favorable reporting:
By now, you may have read Danny Sullivan’s recent post: “Google: Bing is Cheating, Copying Our Search Results” and heard Microsoft’s response, “We do not copy Google's results.” However you define copying, the bottom line is, these Bing results came directly from Google. (Google)
Suddenly the issue becomes clearer; Google has been collecting information on its users for longer than Zuckerberg has been programming. The wealth of data contained in their servers extends across the web; the site runs a mapping service that contains location data on every business in the world. The service is supplemented by a review engine. Search users are tracked through YouTube and Gmail profiles to directly monitor viewing habits and conversations. AdSense, DoubleClick, and any site they connect to can trace exactly what users are seeing in enormous portions of the web. Google may not be directly serving the page views on these sites to boost its comScore rating but it certainly benefits from the user data it collects in each and every interaction. It’s worth going deeper here to see where benefits come from.
![]() |
| If Zucks likes it then it must be good for me because he's rich and I want to be rich too! |
There is a big component here that is another problem; companies have been buying ‘likes’ on Facebook for years and have turned these trillions of impressions into a garbage pile. If a company bought a ‘like’ so a user would enter a raffle, is it valuable for this or any other company to re-pay to advertise to this person’s friends or to friends of their friends? Facebook has poisoned its ‘likes’ data with bad data through extensive advertising sales, which could make useful analytics data very, very difficult to leverage for beneficial search results.
Why is a customer profile important in search?
Any field a user enters information into on their own should be valuable to shaping results. Case in point; a frequent early complaint about online dating sites like eHarmony was that they would pair vegetarians to hunters. Modern user habits that are monitored and controlled through analytics should actively eliminate these types of mismatches. The concept is simple on the surface but becomes difficult in aggregate. It could be valuable to take everyone who says in their Facebook profiles that they like Tibetan prayer rugs, own a dog, and drive a Subaru, and then list advertisements for Purina’s new vegan dog food as a relevant search result. If this user were shown advertisements for a Christian Singles dating service it may not be a valuable use of the space.
Facebook holds a huge amount of this cross-referential data and has built extensive profiles on all of its users. They also collect information from users who ‘check in’ at events and locations, which can help to build a large map of businesses around their individual business fan pages (that in turn contain huge amounts of business data). This data has shortcomings though; roughly 1% of Facebook users are estimated to interact with advertising. On top of this, almost half of Facebook users say they find Facebook’s interface boring and minimize Facebook interactions. (CNN) This is backed up by the numbers from my last post; 85% of posts are interacted with by less than 5% of users. (Business Insider) Suddenly Facebook’s enormous market advantage could become a huge liability as it depends on the habits of a minority of users.
Mind the gap
The concept that most social users are ‘lurkers’ is backed by hard data, (Wikipedia, Reddit) and it’s actually pretty cool to think about. Facebook has a problem when people don’t fill in their profiles. You may be friends with Pat and Pat may like base jumping, but you could be an avid botany lover who never actually filled in your ‘activities’ section. When Facebook suggests that you graphically line up because you’re friends, they aren’t getting the same type of information that Google has on all of the searches you’ve run in the past.
Instead of being able to look at your history and say ‘You don’t want to see an ad about parachute accessories, you should see an ad about a local gardening store,’ the best Facebook can do is show you what Pat likes. The idea of getting a contextual search answer is great, but analytics need to have more clean data than Facebook can access if they are going to deliver good results.
This is where Facebook holds a great advantage it doesn’t seem to like; the company is ripe for partnerships. The data Facebook holds on users could build phenomenal success on top of Bing’s search if it can leverage all of Bing’s resources. While the information and results are stuck to what’s inside of Facebook there are limitations here. The question is whether Facebook will truly partner with Bing to improve its analytical reach and build a foundation for search results so it can have an excellent set of data to draw from. Until the response engine is broadened and more information is gathered on lurkers, Facebook is going to be playing the search game with a short deck.
You Can't Just Google It!
by: SweetSearch
Edit: 1/27/13
Maybe Graph Search's answer engine will solve first world problems like these.
http://newsroom.fb.com/News/562/Introducing-Graph-Search-Beta
http://www.businessinsider.com/facebooks-new-search-engine-2013-1
http://www.businessweek.com/articles/2013-01-15/facebook-radically-revamps-its-search-engine
http://www.slate.com/blogs/future_tense/2013/01/15/facebook_announcement_graph_search_engine_is_like_a_personalized_google.html
http://searchengineland.com/facebook-search-not-google-search-145124
http://www.wired.com/business/2013/01/facebook-event/
http://www.huffingtonpost.com/2013/01/15/facebook-graph-search_n_2480624.html
http://www.forbes.com/sites/lisaquast/2012/04/23/your-social-media-profile-could-make-or-break-your-next-job-opportunity/
Subscribe to:
Posts (Atom)











