Showing posts with label google. Show all posts
Showing posts with label google. Show all posts

Friday, February 21, 2014

Too Spy or not Too Spy: There should be no Question!

Digital Analytics is here to stay there is no question. With all the optimization, efficiency and the ability to better target your audience, the appeal is natural and understandable. With this in mind; how do we sure up the data so as to maintain privacy in an increasingly “Peeping Tom” environment?

Today’s Information environment is vast; companies have figured ways to collect data from their users in ways never before believed possible. When, where, how much, frequency, and the list goes on for what is being collected. Devices to chart sleep patterns, location, steps per day, mood when buying based on patters, items related to current purchase based on previously collected data, etc. All this advancement with one real goal; money! No I mean a happier you oops. A Nobile goal on the surface but who is looking at it from any other angle, is another perspective even necessary?

Let’s start with no. No we have no need to be concerned. No one cares about this data, at least not from a malicious standpoint. Even if they did we have randomized it so that users are protected and kept anonymous. In fact all we know is IP address activity, not even who is using it. Or in other cases, yes we know who they are, they created an account but we enforce strict security measures to prevent the malicious use of our user’s data. We would never allow access to our data without going through the proper channels and following regulation.

While this sounds good on the service and it might even be true let’s look at it from another angle. And let’s call this angle “Current Events”. For the past several months we have read articles about a man named Edward Snowden, Articles about the “NSA” National Security Agency, and how they are behaving with our information and who/how they are obtaining it. There are even articles about other countries spying for many years on our data and we are just finding out. http://www.wired.com/threatlevel/2014/02/mask/

We should now look at this phenomenon from another angle. Our data is not private! It is neither safe nor secure! It is being sold to the highest bidder with whatever motives suit the buyer as it is not in the sells best interest to ask. And most people don’t understand what that means in breadth or depth.

This is not an environment of privacy! In fact most people, especially the elderly, have no idea they are consenting to the whoring of their data in the first place (even if they did hit the “I agree” button at the end of their 5,000th, 200page consent form). Hopefully no atrocities are going to be committed by the data being collected/used this way, at least we all hope there aren’t (makes you wonder what the NSA told AT&T to leave out of their report, http://www.wired.com/threatlevel/2014/02/ma-bell-non-transparency/ ). This begs the question. What might happen if the data was to be used maliciously? Who would be accountable?

With statements like this from Google “users have no legitimate expectation of privacy” what does that mean for the future. http://www.digitaltrends.com/web/google-gmail-users-have-no-expectation-of-privacy/ If I want to keep my information private how do I do this and still participate in this wave of future with all of its convince and pizazz? Is it reasonable for a non IT expert to go off the grid in hopes to keep some privacy? Do I forego all the sales and deals associated with this data collection movement? Or do we look at protecting the users privacy TRULY while still collecting the necessary data? There seem to be a few companies who have this mind set and users appreciate it.

It is important for the Data Analytics of the future to keep the trust of its data pool. There is very much a symbiotic relationship between the two and the relationship needs to remain pure. No more NSA “gifts”, no more selling your user data unless completely anonymous (it’s not anonymous if our email is getting spammed or our phones are being called by the company that bought our data). Protect the privacy of your users and Karma will make sure it comes back to you!

Tuesday, February 18, 2014

The Piracy of Your Privacy

Privacy on the Internet has become non-existent in the digital age.  Your digital footprint is tracked and analyzed by various marketing firms all vying for your attention.  They have found ways to weasel into your computer, tablet, phone, and other Internet-connected devices.  They know every website you visit, how long you were there, what you did there, how long you were there, how you got there, and even where you are signing in from.  Nothing is secret anymore, so be mindful of the information that you send to cyberspace. 

Web browsers like Google Chrome have found a sneaky way to track all of your digital movements across all devices.  They ask you to sign in.  This act allows Google to keep a history of all of your movements on all devices.  The “Stay Signed In” checkbox on the login page keeps this action further from the user’s mind because they do not have to sign in each time they open Chrome. 


Chrome is also able to combine all of your Google services.  Like the slogan on the page says, “One account.  All of Google.”  This means that they know of all of the Google Apps you are using and how you are using them. 

According to PC Mag, the Federal Trade Commission (FTC) issued an advisory on Web tracking of customers.  Chrome was one of many browsers to enable features that protect their users from advertisers.  They also have an Incognito Mode that allows users to move around the Web without leaving a footprint behind.1  While these options allow a user their privacy, most users are not aware that they exist or understand their purpose. 

Further actions have been taken in the industry with the creation of websites such as donottrack.us.2 which allows users to opt out of tracking by websites that they do not personally visit.  However, a Wall Street Journal Investigation found that the top 50 websites in the United States install an average of 64 individual trackers to visitor computers.3 These trackers are often fast and invisible.  Even though this article was written in 2010, it is not hard to imagine that tracking software changes as fast as its governing laws do.

Why do companies want to stay ahead of governance?  NBC News reports that is because online data has been transformed into a lucrative marketplace.  Information is being sold to the highest bidder.  This information includes names, addresses, phone numbers, and credit card numbers, which are being traded out in the open.4  This continues to be a growing concern among consumers.  According to a study performed by the Pew Research Center, fifty percent of Internet users are worried about the information about them online, compared to thirty three percent in 2009.  Of those surveyed, eighty six percent of people have tried at least one technique to hide their activity online or avoid being tracked.5  The study also found that 68 percent of people feel that the law is insufficient to protect their privacy.

Causes of Privacy Loss:
  •  People are providing too much information on social network sites and email.
  • E-Commerce sites are better able to hide their tracking
  • Security breaches of E-Commerce websites

 Laws for security breaches come from individual states.  Currently, states require businesses and public agencies to notify consumers of security breaches of personal information.  According to USA Today, California often leads the way in challenging privacy issues.  They are considering multiple protections for Internet users.6
  •  Password Protection – California feels that security breach laws should also protect passwords, usernames, and security questions.  Hackers for popular social media websites such as Facebook, Twitter, Google, and Yahoo stole nearly two million usernames and passwords.
  • Do Not Track – While current legislation does not ban tracking, the new legislation requires companies to disclose how they comply with requests from Internet users who ask not to be tracked.
  • The Teen “Eraser” Law – Requires all website and mobile app operators to provide a way for those under 18 to delete a posting or photo.  The reasoning is to protect those under 18 from publically sharing “ill-advised pictures or messages”. 


Works Cited

1. Muchmore, Michael. "Google Chrome 31." PCMAG. PC Mag, 11 Dec. 2013. Web. 18 Feb. 2014.

2. Mayer, Jonathan, and Arvind Narayanan. "Do Not Track." - Universal Web Tracking Opt Out. Center for Internet and Society, n.d. Web. 16 Feb. 2014.

3. Butler, Christopher. "Unlimited vs. Limited Web Tracking." Tracking Best Practices. Newfangled, 01 Sept. 2010. Web. 18 Feb. 2014.

4. Sullivan, Bob. "Online Privacy Fears Are Real." Msnbc.com. NBC News, 6 Dec. 2013. Web. 18 Feb. 2014.

5. Flaherty, Anne. "Study Finds Online Privacy Concerns on the Rise." Yahoo! News. Yahoo!, 05 Sept. 2013. Web. 18 Feb. 2014.

6. Prah, Pamela M. "Target's Data Breach Highlights State Role in Privacy." USA Today. Gannett, 10 Feb. 2014. Web. 17 Feb. 2014.



Monday, February 17, 2014

Browser Cookies 2.0 - Google AdID - The Holy Grail of user information

Browser Cookies?!?!

Tania recently posted a blog entry explaining what web cookies are, how they work, and how to manage them.  At a basic level cookies allow information to be aggregated about you and your browsing habits.  This helps content providers (but mostly marketers) tailer a more relevant experience for users but also raises privacy and security concerns as these shadow profiles are enriched over time.

Google browser cookies - Google AdIDBrowser cookies come in two varieties, first party and third party. [1]  The first party variety are generally the more benign of the two and only aggregating data in the current site where you are visiting.  Third party cookies however, go where you go.  They can collect product preferences from Pinterest, employment information from LinkedIn, and just about anything else from social media.  This cocktail of information can be powerful for marketers and potentially dangerous in the wrong hands.

As users, we tend to leverage multiple browsers, tabs, and even platforms (phone, tablet, desktop, e-reader, gaming systems, etc) to facilitate our data and connectivity addictions.  This makes it more difficult for the third party cookies to get the "full" story.  Users can also block and delete cookies (see "Do Not Track").  These weaknesses have lead to companies like Google to create solutions to stitch some of these data holes together (like linking all your Crome profiles together, platform agnostic).

Holy Grail?!?!

intelligent third party cookie
It would be naive to think that companies would stop short of the 'Holy Grail' of user information. Google is already working on an alternative to cookies called AdID.  Essentially, the search giant is building an "anonymous identifier for advertising, or AdID, that would replace third-party cookies as the way advertisers track people's internet browsing activity for marketing purposes." [2]

Browser Cookies vs Google AdID - Source: http://online.wsj.com/news/articles/SB10001424052702304682504579157780178992984


There are interesting implications of this development [3] but its important to note that the devil is in the details, all of which are speculative at this stage of development.

The Good:

  • The AdID is intended to be an anonymous
  • Ownership of the AdID is consolidated which can allow for better governance. This is a slight improvement over the smorgasbord of third parties tracking user traffic with almost no governance.
  • Users will theoretically be given tdiscretion over how much data (or which data) is aggregated
  • There might be a "private browsing option"

The Bad:

  • Although intended to be anonymous, it has a huge potential to be linked to something not anonymous like an email or G+ profile
  • Tracking is tracking. The information can be willingly sold, lawfully subpoenaed, or wrongfully stolen
  • If you are a company tracking users with cookies, future browsers like Crome, may auto-block (verses the current opt-in) cookies. Your only option will be to pay the big boys for that information.

The bottom line

Should we hate Google for doing this to us? It's hard to blame them for wanting to profit off the information that they serve us for free.  Besides; Microsoft, Apple, and FaceBook are in the game too. [4] Users need to take more ownership of their personal information and activities and choose when and how it can be used, otherwise, someone else will.


=========================================================
[1] http://adage.com/article/digital/google-s-plan-replace-cookies/244302/
[2] http://www.engadget.com/2013/09/18/google-adid-an-anonymous-identifier-advertising-cookie/
[3] http://www.digitaltrends.com/opinion/what-is-google-adid-and-how-will-it-replace-browser-cookies/
[4] http://online.wsj.com/news/articles/SB10001424052702304682504579157780178992984

An Update on Do Not Track and Privacy

In his January 2013 post to the Digital Analytics - University of Utah course blog, McCall Lewis wrote about the ongoing debate surrounding online consumer privacy and efforts towards a standard for "Do Not Track" [i].  McCall correctly stated that 2013 would be the year that these issues came to the consciousness of the consumer at-large. In this post, I intend to explore the ongoing saga of Digital Privacy and how consumers and online entities are reacting.

Digital Privacy in 2013, In a Nutshell (Help!  I'm in in a nutshell!)

It is safe to say that by the close of 2013, no American was completely isolated from developments in the world of digital privacy.  This was the year of Edward Snowden and Wikileaks, which exposed such government spying programs as PRISM, Tempora, and MUSCULAR [ii].  If people were not previously concerned with the monitoring of their internet behavior, it is hard to believe that they were not starting to think about it.  Most relevant to Digital Analytics is the allegation that the NSA was using cookies to piggyback on the tools that digital advertising firms were using to "pinpoint targets for government hacking and to bolster surveillance".  Google, Yahoo!, and Microsoft announced plans to encrypt traffic between their data centers, with Microsoft indirectly comparing the threat to that of Chinese government-sponsored hacking [iii].


Beyond allegations of government spying, other news events triggered a growing concern for digital privacy amongst citizens.  It was revealed that Google, despite their unofficial corporate motto being "Don't be evil", was collecting and storing data on WiFi networks while driving the avenues and boulevards in their mapping vehicles [iv].  Inadvertent or not, this revelation made big headlines in the year of Digital Privacy concerns.  Beyond government and corporate spying, there were a number of black-hat news stories as well.  Major retailers such as Target and Neiman Marcus were victims large-scale data breaches in which personal and credit card information were stolen from their servers.  While nothing connected to a network is ever totally secure, some of the details surrounding these breaches made it clear that retailers were not doing everything that they could to protect this sensitive data. In this particular case, the suspected security snafu source was an HVAC contractor that was given the keys to the castle, which were thusly compromised [v].

Current Sentiment
Not surprisingly, there have been numerous studies trying to suss out what the consumer reaction is to all of this. A University of Vienna focused on the act of "Virtual Identity Suicide" within the online social networking site Facebook.  The single biggest cause for this phenomenon, where a user deletes as much of their content as possible before permanently locking themselves out of their account, were concerns over privacy.  Among users studied, over 48% expressed this viewpoint [vi]. It turns out, they have a right to be concerned. Austrian law student Max Schrems found out, in 2010, that Facebook had over 1,200 pages of data on him alone.  This included data that he had never been supplied, but had been linked to him through his friends contact list.  As big-data analytics gets more powerful, this could translate into an enormous amount of personal information being available to online companies [vii].
Using information like Facebook collects, identification of protected classes is not only possible, but on the verge of child's play.  The Center for Digital Democracy is making efforts to address its concerns to the FTC.  They state that technologies such as hyper-local targeting, geo-fencing, and cross-platform targeting will allow for rampant discrimination.  The sorts of questions that it is illegal for employers to ask (age, marital status, sexual orientation) will become easily attainable information [viii].
TrustE, a digital privacy management company, conducted a study regarding consumer opinions about Online Behavioral Advertising recently [ix].  They found that 69% of internet users understood the value trade-off of online ads versus free content, but only 26% are willing to actually accept the same.  It seems as though most internet users feel powerless in the process that they need to just accept what is offered.  The study also showed that 62% of users would be more willing to do business with a company that allowed them to opt-out of targeting.

Ongoing Efforts
The WC3 is spearheading a Do-Not-Track and privacy working group, but things are not going as well as could be hoped.  One of the biggest internet watchdog and lobbying organizations, the EFF (Electronic Frontier Foundation), has lost confidence in the group [x].  They have directly stated that if the group continues in the direction that it is currently headed, that they may be forced to drop out.  Another watchdog group has a similar stance.  Jeffrey Chester, of the Center for Digital Democracy as called the efforts of the group "a farce".  It appears as though the group cannot even get the definition of tracking nailed down.  Are they concerned with 1st party cookies, 3rd party cookies, or other methods of data collection?  Original efforts in the Do-Not-Track space only targeted 3rd party cookie based ads, providing a guise of privacy to the relatively uneducated user.
It is unclear what the future may bring in terms of digital privacy, but it is doubtful that it will continue to be as unregulated as it currently is.  The European Union is enacting tough laws, requiring explicit consent in some areas, rather than the arguably implicit consent given by endless EULAs and TOCs that no one actually reads.  If the FTC gets involved in the United States, things are likely to change.


[i] Lewis, McCall ‘The “Do Not Track” Debate’ Digital Analytics – University of Utah, January 26, 2013. http://dauofu.blogspot.com/2013/01/the-do-not-track-debate.html
[ii] Wikipedia contributors, "Edward Snowden," Wikipedia, The Free Encyclopedia, http://en.wikipedia.org/w/index.php?title=Edward_Snowden&oldid=595457266 (accessed February 5, 2014).
[iii] Wikipedia contributors, "MUSCULAR (surveillance program)," Wikipedia, The Free Encyclopedia, http://en.wikipedia.org/w/index.php?title=MUSCULAR_(surveillance_program)&oldid=595197410 (accessed February 5, 2014).
[iv] “Street View: Google given 35 days to delete wi-fi data”, from BBC News: Technology, June 21, 2013. http://www.bbc.co.uk/news/technology-23002166
[v] Feinberg, Ashley “Last Month's Massive Target Hack Was the Heating Guy's Fault” Gizmodo, February 5, 2014. http://gizmodo.com/last-months-massive-target-hack-was-the-heating-guys-1516926877
[vi] Munson, Lee “Half of Facebook-quitters leave over privacy concerns” NakedSecurity, September 18, 2013. http://nakedsecurity.sophos.com/2013/09/18/half-of-facebook-quitters-leave-over-privacy-concerns/
[vii] Solon, Olivia “How much data did Facebook have on one man? 1,200 pages of data in 57 categories” Wired.co.uk, December 28, 2012. http://www.wired.co.uk/magazine/archive/2012/12/start/privacy-versus-facebook
[viii] Submitted by demedia, “CDD Calls on FTC to Protect Privacy in today's Hyper-local, geo-targeting, cross-platform, Big Data Era/Warns of Discriminatory Practices with mobile device tracking”, Center for Digital Democracy, February 6, 2014. http://www.democraticmedia.org/cdd-calls-ftc-protect-privacy-todays-hyper-local-geo-targeting-cross-platform-big-data-erawarns-disc
[ix] Deasy, Dave “TRUSTe Study Reveals Increased Transparency and Privacy Controls Produce More Positive Feelings about OBA” TrustE Blog, September 19, 2013. https://www.truste.com/blog/2013/09/19/truste-study-reveals-increased-transparency-and-privacy-controls-produce-more-positive-feelings-about-oba/
[x] Fung, Brian “The Internet’s best hope for a Do Not Track standard is falling apart. Here’s why.” The Washington Post Online, The Switch, October 11, 2013. http://www.washingtonpost.com/blogs/the-switch/wp/2013/10/11/the-internets-best-hope-for-a-do-not-track-standard-is-falling-apart-heres-why/

Sunday, February 16, 2014

Fighting the Secure Search Battle

Digital marketers have become accustomed to no longer being able to receive complete search term data. Google became the first to secure their search data with their announcement that all signed-in users would be secured in October of 2011. Now more than two years later, Yahoo has announced that they will follow Google’s lead and make Yahoo searches secure, causing digital marketers to once again scramble for ways to gain access to their keywords.   


The Google empire completed the full transition in September of 2013 as they confirmed that all searches would be secure by default.  What this means is that instead of seeing the individual keyword or searched term in your analytics tool, you will see the phrase “not provided.”  Marketers will know that a search happened but be unaware of the exact term used to bring the visitor to the site.

No Referrer with Yahoo

Yahoo is planning to fully shift to secure searches by March 31st, 2014. Like Google, all searches will be done through a secure server.  Unlike Google, there will be no “not provided” term in your analytics. Yahoo will not be sharing any information at all.  All data driven from Yahoo will appear as though the visitor came directly to the site. The No Referrer policy will make it difficult for the analysts and digital marketer to differentiate between a Yahoo search and a direct visit.

Optional Secure Search at Bing

Bing continues to be that friend not willing to rock the boat. Bing still wants your friendship.  In January, Bing officially launched secure search though it has been turned off by default.  Digital marketers will continue to receive their Bing keywords for the time being.

If or when the user decides to turn on secure search with Bing, the data will appear similar to Yahoo's with a No Referrer policy. 

Tips on How to Fight the Secure Search Battle 
  1.  Use Bing.  Even though Bing may only represent a small portion of your organic search volume, it can provide insight into visitor behavior and give you search term data.
  2. Google Analytics may not provide specific keyword info, but Google Webmaster tools may still provide what you need.  Experiment with the Search Queries report and you may be pleasantly surprised with what you find!
  3. Use Google Adwords. Though Google claims that it has transitioned to secure search for security purposes, the keyword data can still be bought through Google Adwords.
  4. Use site searches to understand what users are searching for on your site. You can receive valuable insights on potential keywords by tracking what users are searching for within your site.
  5. Don’t worry and be patient. Mastermind search marketers will find ways to get our keywords back. Continue to stay updated with industry news and notes as valuable information is released and tools are developed.

Wednesday, February 13, 2013

Bringing Traffic Your Way


You just created that awesome blog site, and your family and friends are proud of you and your work. They come to visit the site a few times a week, reading and commenting on your posts; they actually let you know how cool you are for writing these posts! 
But, you want to tell the world about your blog. You want to start driving traffic your way! Well, I’m going to help you with some cool ideas so your blog can get your first 1 million hits.

 Social Media Working for You 

You did it! Your blog is looking pretty awesome. You have created a cool brand and now you’re ready to take it to the next level. Here’s where social media can help immensely. Go to Facebook, Twitter, LinkedIn and Google+ to create your precious brand personal accounts. Even if you believe your blog won’t be a success right away, is better to avoid headaches by creating your own accounts on every social networking site.
Now that you have created accounts on these, it’s time to start sharing every post to them. Make sure to invite friends to like or join your name so they receive notifications and or emails every time you post a new blog.



Using Analytics to Track Everything

Install Analytics Tools on your blog and let it collect data. Tools like Google Analytics and Adobe Site Catalyst can be installed on the back-end of your blog. These tools track everything from new visits, to time spent on pages. Analytics can also track, where your visits are coming from, if the visitor landed on your blog by an organic search(1) or by a paid keyword(2). These tools can be configured to your liking and they provide detailed information in reports that can be customized and delivered automatically. 
If you have been following this post from the beginning, you should have a Google+ account. That’s all you need to create a Google Analytics account. The interface is very simple and Google provide great documentation for it. You can also take a look at the Universal Analytics brand new tool from Google in this post. Here's a great post about how to use Google Analytics.

Study Your Audience and Write Accordingly to What They Want

Finish your blog posts with specific questions; you’re going to start knowing your visitors based on the answers you get.
You can also try sending surveys via email or post these surveys along with your blog post. This is a quick way to understand what your visitors want. You can start writing about what your visitors wants once you get the necessary data about them.
Again, use social media to interact with your users. Learn from them and become their friend. Participate in communities your visitors visit frequently and learn as much as you can from them.

Use Great Designs on Your Website

Make sure you visit different blog sites to see and learn the trends used. Stay away from using bright colors that can cause temporal blindness on your audience, and use one font type for your posts. Use pictures and videos related to what you’re writing about and listen to feedback from your visitors.

Reply to Comments

Interacting with your audience is just another way to increase traffic to your site.  Make sure you interact with them when they leave a comment. Even if it’s a simple “thank you,” from your part; they will be grateful and come back to visit your blog. Replying to the audience comments creates a great discussion environment. Try to avoid religion or political comments unless your blog was created for these purposes.


Conclusion

We talked about using social media to interact with people and expand your blog to different audiences. We also talked about creating a Google Analytics account and installing these tools in the back-end to learn about your audience. Along with that, we took a look at studying your audience, using great designs and interacting with your visitors. I hope I was able to help you and my words encourage you to start writing about the things you really love.
 

One more thing…

Keep posting and don’t be discouraged if you had a few visitors this week. Don’t give up and be consistent.

Do you have any ideas on how to start bringing blog traffic your way?


References
Seomoz.org
dauofu.blogspot.com
Google.com/Analytics
http://en.wikipedia.org/wiki/Keyword_advertising


Tuesday, February 12, 2013

Universal Analytics - One Step Closer to Becoming Omniscient!

Universal Analytics Logo
Universal Analytics

Google Analytics < Universal Analytics?


Over the past few weeks we have analyzed, examined, studied, pontificated, and expounded upon all of the benefits (and potential issues ) of using Google Analytics. To add fuel to that fire Google released a new tool called Universal Analytics. With UA, we have at our finger tips a variety of new possibilities in the analytics space, specifically:

  • Use the Measurement Protocol to integrate data across multiple devices and platforms.
    • The Measurement Protocol introduced by UA lets you collect and send incoming data from any device to your Analytics account, so you can track more than just websites. Leverage the new analytics.js code and new developer reference libraries from the Measurement Protocol to see how users interact with all of your devices — smartphones, tablets, game consoles, and even digital appliances.
  • Improve lead generation: Sync offline and online data.
    • With UA, you can track data from all your online and offline customer contact points, like marketing campaigns, sales calls, and store visits, so you can discover relationships between the channels that drive conversions. Because UA is primarily an innovation in data collection methods, there are no new reports showing cross-device data.
  • Define your own dimensions & custom metrics.
    • Custom dimensions and custom metrics are like default dimensions and metrics in your Analytics account, except you create them yourself. Use them to collect data that Google Analytics doesn’t automatically track.

  • Understand how well your mobile apps perform.  
    • Mobile App Analytics captures mobile app-specific usage data and integrates it with your Google Analytics account, where you can reapply your knowledge of web analytics to dedicated app reports. (Currently in beta.) 2


Saturday, February 9, 2013

PageRank explained


Today, I thought I'd post about something near and dear to my heart: math.  When I was a senior at BYU (insert obligatory boos) studying numerical analysis, for one of my classes I wrote a paper about the PageRank algorithm.  Seeing as this is a web analytics class and that a big part of web analytics these days is search engine optimization, I thought I’d revisit the topic.  This time, though, I’ll do it in a way that is a lot simpler, involves less mathematical proofs, and I hope is less boring.
For those of you that don’t know what PageRank is, before Google adopted the “whoever gives the most money to Google wins” algorithm, Larry Page and Sergey Brin from Stanford developed a way to in essence let the internet itself determine the relative importance of the pages that it contains.  In the algorithm each member (page) of groups of hyperlinked documents (aka the internet) is assigned a weight based on the number of hyperlinks to it from other pages.  So, a page with a lot of links to it has a higher rank than a page with only a few links to it.

How it works


Suppose we have an internet with 4 pages: A, B, C, and D with links to each other as illustrated below.  In this case, A has one link to it from D; B has a link to it from A, C, and D; C has one link to it from D; and D has 2 links to it, one from A and one from B.  Every time a page links to another page, it transfers a portion of its “rank” to the page that it links to.  So, D has a link to A, B, and C, so it transfers a third of its rank to A, a third to B, and a third to C.







So in our example, the ranks of each page are represented by the following equations:
  
 













Or, using matrix notation, it is the solution to the system of equations below. 











Those of you who are mathematicians will notice that the PageRank values are an eigenvector of the matrix of link weights.  In our case, we want the one where the sum of all the ranks is 1.  So, for our model A=0.13, B=0.33, C=0.13, and D=0.4.
So, what exactly does this rank mean?  One interpretation is that that it is the probability that after following links for a long time you’ll end up on that particular page (If you try this on the real internet, you’ll likely either end up looking at Wikipedia or porn).
This is a simplified version of PageRank.  The actual algorithm is a bit different to take into account that not all pages have outbound links, people don’t just follow links all day when they surf the internet, etc., but this is essentially how it works.

What it means for your site

Knowing all this, what does this mean for your site?  Well, the first and most obvious thing is that the more links to your page, the better.  You might be thinking “great, I can just go out and plaster links to my site all over message boards, blogs, Facebook, Twitter, etc. to increase its PageRank” or “OK, I can just pay people to put links to my site on theirs.”  Unfortunately, most message boards, blogs, etc. use the "nofollow" tag which tells Google not to include these links in PageRank calculation to prevent this kind of spamming.  Also, Google has specifically cautioned against selling links to increase PageRank.  If they catch people doing this, their links are excluded from calculations [1].  For this reason, Google has advised using the "nofollow" tag on sponsored links.

Also, take note of where links to your page are coming from.  Remember that when a page links to yours, it transfers a portion of its PageRank to it.  A link from a big, important site is worth a lot more than a bunch of links from small, obscure sites.

Saturday, January 26, 2013

What does Facebook’s Graph Search mean to analytics?

What does Facebook’s Graph Search mean to analytics?

Paid search has been dominated by Google so far. The company holds enormous amounts of search data that it is able to sell as part of its advertising packages. On top of this, Google runs its free web analytics engine that integrates with its search platform to deliver even better ad placements and results. Advertisers are able to see what people do not just on Google, YouTube, and in Gmail, but also take advantage of Google’s reach across third party web sites in its current ad presence. Google can help to build intimate customer profiles based on these browsing habits, and turns that data into very, very accurate advertising placements based on users’ web history. (Salesforce is a good demonstrator that customer profiles are worth big money) Power plays to fight Google’s foothold in analytics data, search, and search advertising, even those made by Microsoft, have fallen flat so far.

The massive volume behind Facebook’s user base puts the company in a theoretical power position to leverage personal information. The site holds an extensive network of information and social connections that it has never properly leveraged in the past. Running a search had previously only turned up sporadic friends and few relevant results, yet Facebook has seemed relatively uninterested in reworking its data to improve search results or ad revenues.

Press reaction was heated after the publicity surrounding Facebook’s public launch. With 1 in 6 people on the planet (a substantial ratio of whom are active participants in the global economy, rather than an isolated one, which in turn could drive ad sales) on the service there is a wealth of data to be examined and mined on the back end. People put fairly intimate public profiles on the site, build photo libraries with family experiences, make friends, and actively communicate on Facebook’s servers. All of this is fair game for search analysis, and can allow the site to suggest search results that a friend likes. Just as importantly, it is possible to filter out results based on what friends don’t like, or customize results based on other users who also like ‘Game of Thrones’ and Mexican food in the San Francisco area (Facebook’s chosen example).

This information is predicted to have an excellent impact on analytics. Facebook should be able to use customer profiles to deliver extremely targeted results. This should result in excellent placement for things like restaurants or attractions, and should allow Facebook to charge a premium for ad results because friends and friends of friends can be accurately profild and targeted.

What will the results consist of?

Zuckerberg’s stated goal is that:

“We are not indexing the Web,” Zuckerberg said. “We are indexing our map of the [social] graph.” Users can navigate through the 240 billion photos on the network, the trillions of user “likes,” and the connections between users. (Bloomberg)

This is a brilliant concept with a huge flaw. This is a massive network. The data is overwhelms belief. Yet it only consists of data within Facebook. While most news articles fawn over the idea that Facebook will be able to rival Google, the engine currently relies upon core results results that are billed from Bing. Microsoft was in the news for Bing’s results when it first launched, yet it was not the most favorable reporting:

By now, you may have read Danny Sullivan’s recent post: “Google: Bing is Cheating, Copying Our Search Results” and heard Microsoft’s response, “We do not copy Google's results.” However you define copying, the bottom line is, these Bing results came directly from Google. (Google)

Suddenly the issue becomes clearer; Google has been collecting information on its users for longer than Zuckerberg has been programming. The wealth of data contained in their servers extends across the web; the site runs a mapping service that contains location data on every business in the world. The service is supplemented by a review engine. Search users are tracked through YouTube and Gmail profiles to directly monitor viewing habits and conversations. AdSense, DoubleClick, and any site they connect to can trace exactly what users are seeing in enormous portions of the web. Google may not be directly serving the page views on these sites to boost its comScore rating but it certainly benefits from the user data it collects in each and every interaction. It’s worth going deeper here to see where benefits come from.



If Zucks likes it then it must be good for me because he's rich and I want to be rich too!

There is a big component here that is another problem; companies have been buying ‘likes’ on Facebook for years and have turned these trillions of impressions into a garbage pile. If a company bought a ‘like’ so a user would enter a raffle, is it valuable for this or any other company to re-pay to advertise to this person’s friends or to friends of their friends? Facebook has poisoned its ‘likes’ data with bad data through extensive advertising sales, which could make useful analytics data very, very difficult to leverage for beneficial search results.

Why is a customer profile important in search?

Any field a user enters information into on their own should be valuable to shaping results. Case in point; a frequent early complaint about online dating sites like eHarmony was that they would pair vegetarians to hunters. Modern user habits that are monitored and controlled through analytics should actively eliminate these types of mismatches. The concept is simple on the surface but becomes difficult in aggregate. It could be valuable to take everyone who says in their Facebook profiles that they like Tibetan prayer rugs, own a dog, and drive a Subaru, and then list advertisements for Purina’s new vegan dog food as a relevant search result. If this user were shown advertisements for a Christian Singles dating service it may not be a valuable use of the space.

Facebook holds a huge amount of this cross-referential data and has built extensive profiles on all of its users. They also collect information from users who ‘check in’ at events and locations, which can help to build a large map of businesses around their individual business fan pages (that in turn contain huge amounts of business data). This data has shortcomings though; roughly 1% of Facebook users are estimated to interact with advertising. On top of this, almost half of Facebook users say they find Facebook’s interface boring and minimize Facebook interactions. (CNN) This is backed up by the numbers from my last post; 85% of posts are interacted with by less than 5% of users. (Business Insider) Suddenly Facebook’s enormous market advantage could become a huge liability as it depends on the habits of a minority of users.

Mind the gap

The concept that most social users are ‘lurkers’ is backed by hard data, (Wikipedia, Reddit) and it’s actually pretty cool to think about. Facebook has a problem when people don’t fill in their profiles. You may be friends with Pat and Pat may like base jumping, but you could be an avid botany lover who never actually filled in your ‘activities’ section. When Facebook suggests that you graphically line up because you’re friends, they aren’t getting the same type of information that Google has on all of the searches you’ve run in the past.

Instead of being able to look at your history and say ‘You don’t want to see an ad about parachute accessories, you should see an ad about a local gardening store,’ the best Facebook can do is show you what Pat likes. The idea of getting a contextual search answer is great, but analytics need to have more clean data than Facebook can access if they are going to deliver good results.

This is where Facebook holds a great advantage it doesn’t seem to like; the company is ripe for partnerships. The data Facebook holds on users could build phenomenal success on top of Bing’s search if it can leverage all of Bing’s resources. While the information and results are stuck to what’s inside of Facebook there are limitations here. The question is whether Facebook will truly partner with Bing to improve its analytical reach and build a foundation for search results so it can have an excellent set of data to draw from. Until the response engine is broadened and more information is gathered on lurkers, Facebook is going to be playing the search game with a short deck.


You Can't Just Google It!
by: SweetSearch

Edit: 1/27/13
Maybe Graph Search's answer engine will solve first world problems like these.

Some additional, and substantially more bullish, readings on Facebook's new search engine:
http://newsroom.fb.com/News/562/Introducing-Graph-Search-Beta
http://www.businessinsider.com/facebooks-new-search-engine-2013-1
http://www.businessweek.com/articles/2013-01-15/facebook-radically-revamps-its-search-engine
http://www.slate.com/blogs/future_tense/2013/01/15/facebook_announcement_graph_search_engine_is_like_a_personalized_google.html
http://searchengineland.com/facebook-search-not-google-search-145124
http://www.wired.com/business/2013/01/facebook-event/
http://www.huffingtonpost.com/2013/01/15/facebook-graph-search_n_2480624.html

http://www.forbes.com/sites/lisaquast/2012/04/23/your-social-media-profile-could-make-or-break-your-next-job-opportunity/