Wednesday, July 3, 2013
Conference Report: USENIX Annual Technical Conference (ATC) 2013
This year marks Google’s eleventh consecutive year as a sponsor of the USENIX Annual Technical Conference (ATC), just one of the co-located events at USENIX Federated Conference Week (FCW), which combines numerous conferences and workshops covering fields such as Autonomic Computing, Feedback Computing and much more in an intensive week of research, trends, and community interaction.
ATC provides a broad forum for computing systems research with an emphasis on implementations and experimental results. In addition to the Googlers presenting publications, we had two members on the program committee of ATC and several keynote speakers, invited speakers, panelists, committee members, and participants at the other co-located events at FCW.
In the paper Janus: Optimal Flash Provisioning for Cloud Storage Workloads, Googler Christoph Albrecht and co-authors demonstrated a system that allows users to make informed flash memory provisioning and partitioning decisions in cloud-scale distributed file systems that include both flash storage and disk tiers. As flash memory is still expensive, it is best to use it only for workloads that can make good use of it. Janus creates long term workload characterizations based on RPC samples and file age metadata. It uses these workload characterizations to formulate and solve an optimization problem that maximizes the reads sent to the flash tier. Based on evaluations from workloads using Janus, in use at Google for the past 6 months, the authors conclude that the recommendation system is quite effective, with flash hit rates using the optimized recommendations 47-76% higher than the option of using the flash as an unpartitioned tier.
In packetdrill: Scriptable Network Stack Testing, from Sockets to Packets, Google’s Neal Cardwell and co-authors showcased a portable, open-source scripting tool that enables testing the correctness and performance of network protocols. Despite their importance in modern computer systems, network protocols often undergo only ad hoc testing before their deployment, in large part due to their complexity. Furthermore, new algorithms have unforeseen interactions with other features, so testing has only become more daunting as TCP has evolved. The packetdrill tool was instrumental in the development of three new features for Linux TCP—Early Retransmit, Fast Open, and Loss Probes—and allowed the authors to find and fix 10 bugs in Linux. Furthermore, the team uses packetdrill in all phases of the development process for the kernel used in one of the world’s largest Linux installations. In the hope that sharing packetdrill with the community will make the process of improving Internet protocols an easier one, the source code and test scripts for packetdrill have been made freely available.
There were also additional refereed publications with Google co-authors at some of the co-located events at FCW, notably NicPic: Scalable and Accurate End-Host Rate Limiting, which outlines a system which enables accurate network traffic scheduling in a scalable fashion, and AGILE: Elastic Distributed Resource Scaling for Infrastructure-as-a-Service, a system that efficiently handles dynamic application workloads, reducing both penalties and user dissatisfaction.
Google is proud to support the academic community through conference participation and sponsorship. In particular, we are happy to mention one of the other interesting papers from this year’s USENIX FCW, co-authored by former Google PhD fellowship recipient Ashok Anand, MiG: Efficient Migration of Desktop VM Using Semantic Compression.
USENIX is a supporter of open access, so the papers and videos from the talks are available on the conference website.
Wednesday, August 29, 2012
Google at UAI 2012
The conference on Uncertainty in Artificial Intelligence (UAI) is one of the premier venues for research related to probabilistic models and reasoning under uncertainty. This year's conference (the 28th) set several new records: the largest number of submissions (304 papers, last year 285), the largest number of participants (216, last year 191), the largest number of tutorials (4, last year 3), and the largest number of workshops (4, last year 1). We interpret this as a sign that the conference is growing, perhaps as part of the larger trend of increasing interest in machine learning and data analysis.
There were many interesting presentations. A couple of my favorites included:
- "Video In Sentences Out," by Andrei Barbu et al. This demonstrated an impressive system that is able to create grammatically correct sentences describing the objects and actions occurring in a variety of different videos.
- "Exploiting Compositionality to Explore a Large Space of Model Structures," by Roger Grosse et al. This paper (which won the Best Student Paper Award) proposed a way to view many different latent variable models for matrix decomposition - including PCA, ICA, NMF, Co-Clustering, etc. - as special cases of a general grammar. The paper then showed ways to automatically select the right kind of model for a dataset by performing greedy search over grammar productions, combined with Bayesian inference for model fitting.
A strong theme this year was causality. In fact, we had an invited talk on the topic by Judea Pearl, winner of the 2011 Turing Award, in addition to a one-day workshop. Although causality is sometimes regarded as something of an academic curiosity, its relevance to important practical problems (e.g., to medicine, advertising, social policy, etc.) is becoming more clear. There is still a large gap between theory and practice when it comes to making causal predictions, but it was pleasing to see that researchers in the UAI community are making steady progress on this problem.
There were two presentations at UAI by Googlers. The first, "Latent Structured Ranking," by Jason Weston and John Blitzer, described an extension to a ranking model called Wsabie, that was published at ICML in 2011, and is widely used within Google. The Wsabie model embeds a pair of items (say a query and a document) into a low dimensional space, and uses distance in that space as a measure of semantic similarity. The UAI paper extends this to the setting where there are multiple candidate documents in response to a given query. In such a context, we can get improved performance by leveraging similarities between documents in the set.
The second paper by Googlers, "Hokusai - Sketching Streams in Real Time," was presented by Sergiy Matusevych, Alex Smola and Amr Ahmed. (Amr recently joined Google from Yahoo, and Alex is a visiting faculty member at Google.) This paper extends the Count-Min sketch method for storing approximate counts to the streaming context. This extension allows one to compute approximate counts of events (such as the number of visitors to a particular website) aggregated over different temporal extents. The method can also be extended to store approximate n-gram statistics in a very compact way.
In addition to these presentations, Google was involved in UAI in several other ways: I held a program co-chair position on the organizing committee, several of the referees and attendees work at Google, and Google provided some sponsorship for the conference.
Overall, this was a very successful conference, in an idyllic setting (Catalina Island, an hour off the coast of Los Angeles). We believe UAI and its techniques will grow in importance as various organizations -- including Google -- start combining structured, prior knowledge with raw, noisy unstructured data.
Saturday, July 14, 2012
Google at SIGMOD/PODS 2012
Over the years, SIGMOD has expanded beyond a traditional "database" conference to include several areas related to information management. This year’s ACM SIGMOD/PODS conference (on Management of Data, and Principles of Database Systems), held in Scottsdale, Arizona was no different. We were impressed by the wide variety of researchers from industry and academia alike the conference attracted, and enjoyed learning how others are pushing the limits of scalability in data storage and processing. In addition to an excellent set of papers on a large number of topics, we saw a couple of recurring themes:
1) Data Visualization
- Pat Hanrahan from Stanford gave a keynote on some of the challenges involved in building systems to enable "data enthusiasts" to manage and visualize data.
- Google’s Fusion Tables group also had a paper on this topic: Efficient Spatial Sampling of Large Geographical Tables, by Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Halevy. (This paper has been invited to a TODS special issue on best papers of SIGMOD 2012).
- A similar effort from the University of Washington was presented as a demo: VizDeck: Self-Organizing Dashboards for Visual Analytics, by Alicia Key, Bill Howe, Daniel Perry, Cecilia Aragon.
There’s been a widespread trend over the last several years away from databases, towards highly scalable “NoSQL” systems. We don’t think that trade-off is necessary, and were happy to see several other speakers advocate a similar theme -- yes, databases are useful, and developers shouldn’t need to give up database features and ease of use in the name of scalability.
This theme was supported by an industry session on Big Data featuring talks from other companies: Facebook (TAO: How Facebook Serves the Social Graph), Twitter (Large-Scale Machine Learning at Twitter), and Microsoft (Recurring Job Optimization in Scope). Googler Kirsten LeFevre was a panelist on the "Perspectives on Big Data" panel organized by Surajit Chaudhuri from Microsoft, and also featuring Donald Kossmann from ETHZ, Sam Madden from MIT, and Anand Rajaraman from Walmart Labs. Last but not the least, Surajit Chaudhuri also gave an excellent keynote outlining some of the research challenges that the new era of "Big Data and Cloud" poses.
As has been the practice for several years now, to continue generating great interest in data management research, SIGMOD has been organizing panels such as this year's "New Research Symposium" (which included Anish Das Sarma from Google as a panelist).
In addition to sponsoring the conference, many Googlers attended contributing to a robust presence and affording us the opportunity to interact with the broader information management community. We've been pushing the frontiers of science with cutting-edge research in many aspects of data management, and we were eager to share our innovations and see what others have been working on. We found Amin Vahdat's keynote on the intersection of Networking and Databases to be a highlight of Google’s participation, which also included presenting papers, participating on panels, and taking part in planning and program committees:
Program Committee Members
Anish Das Sarma, Venkatesh Ganti, Zoltan Gyongyi, Alon Halevy (Tutorials Chair), Kristen LeFevre, Cong Yu
Talks
Amin Vahdat, Google (Keynote)
F1-The Fault-Tolerant Distributed RDBMS Supporting Google's Ad Business
Jeff Shute, Mircea Oancea, Stephan Ellner, Ben Handy, Eric Rollins, Bart Samwel, Radek Vingralek, Chad Whipkey, Xin Chen, Beat Jegerlehner, Kyle Littlefield, Phoenix Tong (Googlers)
Finding Related Tables
Anish Das Sarma, Lujun Fang, Nitin Gupta, Alon Halevy, Hongrae Lee, Fei Wu, Reynold Xin, Cong Yu (Googlers)
Papers
Changkyu Kim, Jongsoo Park, Nadathur Satish, Hongrae Lee (Google), Pradeep Dubey, Jatin Chhugani
Efficient Spatial Sampling of Large Geographical Tables
Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Halevy (Googlers)
Panels
Kristen LeFevre, Google
SIGMOD New Researcher Symposium - How to be a good advisor/advisee?
Anish Das Sarma, Google
Overall, this year’s SIGMOD was a great conference, widely attended by researchers from industry and academia, and comprised of a very interesting mix of research presentations and discussions. Google had a good showing at the conference, and we look forward to continuing this trend in the coming years.
Tuesday, May 8, 2012
Google, the World Wide Web and WWW conference: years of progress, prosperity and innovation
More than forty members of Google’s technical staff gathered in Lyon, France in April to participate in the global dialogue around the state of the web at the World Wide Web conference (WWW) 2012. A decade ago, Larry Page and Sergey Brin applied their research to an information retrieval problem and their work—presented at WWW in 1998—led to the invention of today’s most popular search engine.
As I've watched the WWW conference series evolve over the years, a couple of larger trends struck me in this year's edition. First, there seems to be more of a Mobile Web presence in the technical program, relative to recent years. The refereed program included several interesting Mobile papers, including the Best Student Paper Awardee from Stanford University researchers: Who Killed My Battery: Analyzing Mobile Browser Energy Consumption, Narendran Thiagarajan, Gaurav Aggarwal, Angela Nicoara, Dan Boneh, Jatinder Singh.
Second, one gets the sense that the WWW community is moving from the classic "bag of words" view of web pages, to an entity-centric view. There were a number of papers on identifying and using entities in Web pages. While I'm loathe to view this as a vindication of "the Semantic Web" (mainly because this has become an overloaded phrase that people elect to interpret as suits them), the technical capability to get at entities is clearly here. The question is -- what is the killer application? Finally, it’s nice to see that recommendation systems are becoming a major topic of focus at WWW. This paper was a personal favorite: Build Your Own Music Recommender by Modeling Internet Radio Streams, Natalie Aizenberg, Yehuda Koren, Oren Somekh.
In keeping with tradition, Google was a major supporter, sponsoring the conference, the Best Paper Award (Counting beyond a Yottabyte, or how SPARQL 1.1 Property Paths will prevent adoption of the standard, Marcelo Arenas, Sebastián Conca and Jorge Pérez) and four PhD student travel grants. We chatted with hundreds of attendees who hung out with us at the Google booth to chat and see demos about the latest Google product and research developments (see full schedule of booth talks).
Googlers were also active member of the vibrant research community at WWW:
David Assouline delivered the keynote for the Demo Track -- to a standing-room-only crowd -- on the Google Art Project, which uses a combination of various Google technologies and expert information provided by our museum partners to create a unique online art experience. Googler Alon Halevy served as a program committee member. Googlers were also co-authors of the following papers:
- Risk-Aware Revenue Maximization in Display Advertising by Ana Radovanovic and William Heavlin (Googlers)
- SessionJuggler: Secure Web Login From an Untrusted Terminal Using Session Hijacking by Elie Bursztein (Googler), Chinmay Soman, Dan Boneh and John Mitchell
- Spotting Fake Reviewer Groups in Consumer Reviews by Arjun Murkherjee, Bing Liu, and Natalie Glance (Googler)
- Your Two Weeks of Fame and Your Grandmother’s by James Cook, Atish Das Sarma, Alexander Fabrikant and Andrew Tomkins (Googlers)
- YouTube Around the World: Geographic Popularity of Videos by Mirjam Wattenhofer (Googler), Anders Brodersen (Googler), and Salvatore Scellato
- Who Killed My Battery: Analyzing Mobile Browser Energy Consumption by Narendran Thiagarajan, Gaurav Aggarwal (Googler), Angela Nicoara, Dan Boneh and Jatinder Singh
- A Multimodal Search Engine based on Rich Unified Content Description by Thomas Steiner (Googler), Lorenzo Sutton, Sabine Spiller, Marilena Lazzaro, Francesco Saverio Nucci, Vincenzo Croce, Alberto Massari, Antonio Camurri, Anne Verroust-Blondet, Laurent Joyeux
- Enabling on-the-fly Video Shot Detection on YouTube by Thomas Steiner (Googler), Ruben Verborgh, Joaquim Gabarro, Michael Hausenblas, Raphael Troncy and Rik Van De Walle
- Fixing the Web one page at a time, or actually implementing xkcd #37 by Thomas Steiner (Googler), Ruben Verborgh, and Rik Van de Valle
- Appification of the Web by Ed Chi (Googler), Brian Davison, and Evgeniy Gabrilovich (Googler)
- Extracting Unambiguous Keywords from Microposts Using Web and Query Logs Data, as part of the Making Sense of Microsposts workshop by Davi Reis, Felipe Portavales Goldstein, and Fred Quintao (Googlers)
- Human Computation Must Be Reproducible, as part of the CrowdSearch: Crowdsourcing Web search workshop by Praveen Paritosh (Googler)
- WebQuality 2012: The Anti-Social Web by Zoltan Gyongyi (Googler), Carlos Castillo, Adam Jatowt, and Katsumi Tanaka
- The Role of Human-Generated and Automatically-Extracted Lexico-Semantic Resources in Web Search by Marius Pasca (Googler)
- Google Image Swirl by Yushi Jing, Henry Rowley, Jingbin Wang, David Tsai, Chuck Rosenberg, Michele Covell (Googlers)
Thursday, July 21, 2011
Faculty from across the Americas meet in New York for the Faculty Summit
(Cross-posted from the Official Google Blog)
Last week, we held our seventh annual Computer Science Faculty Summit. For the first time, the event took place at our New York City office; nearly 100 faculty members from universities in the U.S., Canada and Latin America attended. The two-day Summit focused on systems, artificial intelligence and mobile computing. Alfred Spector, VP of research and special initiatives, hosted the conference and led lively discussions on privacy, security and Google’s approach to research.
Google’s Internet evangelist, Vint Cerf, opened the Summit with a talk on the challenges involved in securing the “Internet of things”—that is, uniquely identifiable objects (“things”) and their virtual representations. With almost 2 billion international Internet users and 5 billion mobile devices out there in the world, Vint expounded upon the idea that Internet security is not just about technology, but also about policy and global institutions. He stressed that our new digital ecosystem is complex and large in scale, and includes both hardware and software. It also has multiple stakeholders, diverse business models and a range of legal frameworks. Vint argued that making and keeping the Internet secure over the next few years will require technical innovation and global collaboration.
After Vint kicked things off, faculty spent the two days attending presentations by Google software engineers and research scientists, including John Wilkes on the management of Google's large hardware infrastructure, Andrew Chatham on the self-driving car, Johan Schalkwyk on mobile speech technology and Andrew Moore on the research challenges in commerce services. Craig Nevill-Manning, the engineering founder of Google’s NYC office, gave an update on Google.org, particularly its recent work in crisis response. Other talks covered the engineering work behind products like Ad Exchange and Google Docs, and the range of engineering projects taking place across 35 Google offices in 20 countries. For a complete list of the topics and sessions, visit the Faculty Summit site. Also, a few of our attendees heeded Alfred’s call to recap their breakout sessions in verse—download a PDF of one of our favorite poems, about the future of mobile computing, penned by NYU professor Ken Perlin.
A highlight of this year’s Summit was Bill Schilit’s presentation of the Library Wall, a Chrome OS experiment featuring an eight-foot tall full-color virtual display of ebooks that can be browsed and examined individually via touch screen. Faculty members were invited to play around with the digital-age “bookshelf,” which is one of the newest additions to our NYC office.
We’ve already posted deeper dives on a few of the talks—including cluster management, mobile search and commerce. We also collected some interesting faculty reflections. For more information on all of our programs, visit our University Relations website. The Faculty Summit is meant to connect forerunners across the computer science community—in business, research and academia—and we hope all our attendees returned home feeling informed and inspired.
Monday, June 20, 2011
Auto-Directed Video Stabilization with Robust L1 Optimal Camera Paths
Earlier this year, we announced the launch of new features on the YouTube Video Editor, including stabilization for shaky videos, with the ability to preview them in real-time. The core technology behind this feature is detailed in this paper, which will be presented at the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR 2011).
Casually shot videos captured by handheld or mobile cameras suffer from significant amount of shake. Existing in-camera stabilization methods dampen high-frequency jitter but do not suppress low-frequency movements and bounces, such as those observed in videos captured by a walking person. On the other hand, most professionally shot videos usually consist of carefully designed camera configurations, using specialized equipment such as tripods or camera dollies, and employ ease-in and ease-out for transitions. Our goal was to devise a completely automatic method for converting casual shaky footage into more pleasant and professional looking videos.
Our technique mimics the cinematographic principles outlined above by automatically determining the best camera path using a robust optimization technique. The original, shaky camera path is divided into a set of segments, each approximated by either a constant, linear or parabolic motion. Our optimization finds the best of all possible partitions using a computationally efficient and stable algorithm.
To achieve real-time performance on the web, we distribute the computation across multiple machines in the cloud. This enables us to provide users with a real-time preview and interactive control of the stabilized result. Above we provide a video demonstration of how to use this feature on the YouTube Editor. We will also demo this live at Google’s exhibition booth in CVPR 2011.
For further details, please read our paper.
Friday, June 17, 2011
Google at CVPR 2011
The computer vision community will get together in Colorado Springs the week of June 20th for the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR 2011). This year will see a record number of people attending the conference and 27 co-located workshops and tutorials. The registration was closed at 1500 attendees even before the conference started.
Computer Vision is at the core of many Google products, such as Image Search, YouTube, Street View, Picasa, and Goggles, and as always, Google is involved in several ways with CVPR. Andrew Senior is serving as an area chair of CVPR 2011 and many Googlers are reviewers. Googlers also co-authored these papers:
- Where's Waldo: Matching People in Images of Crowds by Rahul Garg, Deva Ramanan, Steve Seitz, Noah Snavely
- Visual and Semantic Similarity in ImageNet by Thomas Deselaers, Vittorio Ferrari
- Multicore Bundle Adjustment by Changchang Wu, Sameer Agarwal, Brian Curless, Steve Seitz
- A Hierarchical Conditional Random Field Model for Labeling and Segmenting Images of Street Scenes by Qixing Huang, Mei Han, Bo Wu, Sergey Ioffe
- Kernelized Structural SVM Learning for Supervised Object Segmentation by Luca Bertelli, Tianli Yu, Diem Vu, Salih Gokturk
- Discriminative Tag Learning on YouTube Videos with Latent Sub-tags by Weilong Yang, George Toderici
- Auto-Directed Video Stabilization with Robust L1 Optimal Camera Paths by Matthias Grundmann, Vivek Kwatra, Irfan Essa
- Image Saliency: From Local to Global Context by Meng Wang, Janusz Konrad, Prakash Ishwar, Yushi Jing, Henry Rowley
If you are attending the conference, stop by Google’s exhibition booth. In addition to talking with Google researchers, you will get to see examples of exciting computer vision research that has made it into Google products including, among others, the following:
- Google Earth Facade Shadow Removal by Mei Han, Vivek Kwatra, and Shengyang Dai
We will demonstrate our technique for removing shadows and other lighting/texture artifacts from building facades in Google Earth. We obtain cleaner, clearer, and more uniform textures which provide users with an improved visual experience. - Video Stabilization on YouTube Editor by Matthias Grundmann, Vivek Kwatra, and Irfan Essa
Casually shot videos captured by handheld or mobile cameras suffer from significant amount of shake. In contrast, professionally shot video usually employs stabilization equipment such as tripods or camera dollies, and employ ease-in and ease-out for transitions. Our technique mimics these cinematographic principles, by optimally dividing the original, shaky camera path into a set of segments and approximating each with either constant, linear or parabolic motion using a computationally efficient and stable algorithm. We will showcase a live version of our algorithm, featuring real-time performance and interactive control, which is publicly available at youtube.com/editor. - Tag Suggest for YouTube by George Toderici and Mehmet Emre Sargin
YouTube offers millions of users the opportunity to upload videos and share them with their friends. Many users would love to have their videos discoverable but don't annotate them properly. One new feature on YouTube that seeks to address this problem is tag prediction based on video content and independently based on text metadata.
6/17/2011 UPDATE: "Posted by" was changed to include Sergey Ioffe.
Thursday, May 19, 2011
Google at ACL 2011
The Annual Meeting of the Association for Computational Linguistics is one of the premier conferences for language and text technologies. Many employees at Google have strong roots in the community of researchers that attend this meeting, including many of our researchers working on machine translation and speech.
At this years conference, Google is particularly well represented. The General Chair is Dekang Lin and a few Googlers are serving as technical Area Chairs (in addition to the plethora of Googlers that reviewed papers for the conference). Google is also a Platinum Sponsor of ACL this year.
Research advances at Google can be seen throughout the conference’s technical content. Below is a complete list of Googler-authored or co-authored papers in the main conference. We want to give special emphasis to this year’s best paper award, given to “Unsupervised Part-of-Speech Tagging with Bilingual Graph-Based Projections” by CMU graduate student and Google intern Dipanjan Das and his internship advisor Slav Petrov. ACL is an extremely selective conference and this award speaks volumes to the importance of syntactic analysis and using bilingual corpora to project syntactic resources from resource rich languages (like English) to other languages. Congratulations Dipanjan and Slav!
Googlers are also involved in two of this year’s tutorials. Marius Pasca will present “Web Search Queries as a Corpus” and Kuzman Ganchev and his colleagues will teach about “Rich Prior Knowledge in Learning for Natural Language Processing”. Finally, Katja Filippova and her colleagues are running a workshop on “Monolingual Text-to-Text Generation”.
ACL will take place this year in Portland from June 19th to June 24th.
Papers by Googlers (a * indicates a paper that will be linked to after the conference):
Ranking Class Labels Using Query Sessions*
Marius Pasca
Fine-Grained Class Label Markup of Search Queries*
Joseph Reisinger and Marius Pasca
Unsupervised Part-of-Speech Tagging with Bilingual Graph-Based Projections
Dipanjan Das and Slav Petrov
Large-Scale Cross-Document Coreference Using Distributed Inference and Hierarchical Models
Sameer Singh, Amarnag Subramanya, Fernando Pereira and Andrew McCallum
Piggyback: Using Search Engines for Robust Cross-Domain Named Entity Recognition
Stefan Rüd, Massimiliano Ciaramita, Jens Müller and Hinrich Schütze
Beam-Width Prediction for Efficient Context-Free Parsing
Nathan Bodenstab, Aaron Dunlop, Keith Hall and Brian Roark
Language-independent compound splitting with morphological operations
Klaus Macherey, Andrew Dai, David Talbot, Ashok Popat and Franz Och
Model-Based Aligner Combination Using Dual Decomposition
John DeNero and Klaus Macherey
Binarized Forest to String Translation
Hao Zhang, Licheng Fang, Peng Xu and Xiaoyun Wu
Semi-supervised Latent Variable Models for Fine-grained Sentiment Analysis
Oscar Tackstrom and Ryan McDonald
Thursday, May 5, 2011
Google at CHI 2011
Cross-posted with the Technical Programs and Events Blog
User-Defined Motion Gestures for Mobile Interaction by Jaime Ruiz, Yang Li*, Edward Lank
Many Bills: Engaging Citizens through Visualizations of Congressional Legislation by Yannick Assogba, Irene Ros, Joan DiMicco, Matt McKeon*
YouPivot: Improving Recall with Contextual Search by Joshua Hailpern, Nicholas Jitkoff*, Andrew Warr*, Karrie Karahalios, Robert Sesek, Nik Shkrob
Festschrift Panel in Honor of Stuart K. Card by Ed H. Chi*, Peter Pirolli, Bonnie John, Judith S Olson, Dan Russell*, Tom Moran
CHI Should be Replicating and Validating Results More: Discuss by Max L. Wilson, Wendy Mackay, Ed H. Chi*, Michael Bernstein, Dan Russell*, Harold Thimbleby
CASE STUDIES
Note: * denotes a Googler
Thursday, January 27, 2011
Google at NIPS 2010
The machine learning community met in Vancouver in December for the 24th Neural Information Processing Systems Conference (NIPS). As always, the single-track program of the main conference featured a number of outstanding talks, followed by interesting late night poster sessions. A record number of workshops covered a wide variety of topics, while allocating sufficient time for skiing in Whistler - after all, many of the most interesting research conversations happen while riding the lift in-between ski runs. This year’s conference also featured a symposium dedicated to Sam Roweis, providing a retrospective on Sam’s life and work. Sam, a fellow Googler and professor at NYU, was at the heart of the NIPS community and is terribly missed.
As always, Google was involved in various ways with NIPS. Here at Google, we take a data-driven approach when solving problems. Therefore, Machine Learning is in one way or another at the core of most of the things that we do. It is therefore unsurprising that many Googlers helped shape the program of the conference or were in the audience. This year, three Googlers served as area chairs and even more were reviewers. Googlers also co-authored the following papers:
- Label Embedding Trees for Large Multi-Class Tasks by Samy Bengio and Jason Weston
- Learning Bounds for Importance Weighting by Corinna Cortes, Yishay Mansour, and Mehryar Mohri
- Online Learning in the Manifold of Low-Rank Matrices by Uri Shalit, Daphna Weinshall, and Gal Chechik
- Deterministic Single–Pass Algorithm for LDA by Issei Sato, Kenichi Kurihara, and Hiroshi Nakagawa
- Distributed Dual Averaging In Networks by John Duchi, Alekh Agarwal, and Martin Wainwright
Additionally, Googlers co-organized three well attended workshops:
- Coarse–to–Fine Learning and Inference by Ben Taskar, David Weiss, Benjamin Sapp, and Slav Petrov
- Low–rank Methods for Large–scale Machine Learning by Arthur Gretton, Michael Mahoney, Mehryar Mohri, and Ameet Talwalkar
- Learning on Cores, Clusters, and Clouds by John Duchi, Ofer Dekel, John Langford, Lawrence Cayton, and Alekh Agarwal
Finally, Yoram Singer gave a great talk on Learning Structural Sparsity at the Sam Roweis symposium and Googlers presented the following talks during the workshops:
- Online Learning in the Manifold of Low–Rank Matrices by Uri Shalit, Daphna Weinshall, and Gal Chechik
- Distributed MAP Inference for Undirected Graphical Models by Sameer Singh, Amar Subramanya, Fernando Pereira, and Andrew McCallum
- MapReduce/Bigtable for Distributed Optimization by Keith Hall, Scott Gilpin and Gideon Mann
- Self-Pruning Prediction Trees by Sally Goldman
- Web Scale Image Annotation: Learning to Rank with Joint Word-Image Embeddings by Jason Weston, Samy Bengio, and Nicolas Usunier
- Coarse–to–fine Decoding for Parsing and Machine Translation by Slav Petrov
Overall, it was a very successful conference and it was good to be back in Vancouver one last time. This coming year NIPS 2011 will be in Granada, Spain. Hasta luego!
Tuesday, October 19, 2010
Google at the Conference on Empirical Methods in Natural Language Processing (EMNLP '10)
The Conference on Empirical Methods in Natural Language Processing (EMNLP '10) was recently held at the MIT Stata Center in Massachusetts. Natural Language Processing is at the core of many of the things that we do here at Google. Googlers have therefore been traditionally part of this research community, participating as program committee members, paper authors and attendees.
At this year's EMNLP conference Google Fellow, Amit Singhal gave an invited keynote talk on "Challenges in running a commercial search engine" where he highlighted some of the exciting opportunities, as well as challenges, that Google is currently facing. Furthermore, Terry Koo (who recently joined Google), David Sontag (former Google PhD Fellowship recipient) and their collaborators from MIT received the Fred Jelinek Best Paper Award for their innovative work on syntactic parsing with the title "Dual Decomposition for Parsing with Non-Projective Head Automata".
Here is a complete list of the papers presented by Googlers at the conference:
- Dual Decomposition for Parsing with Non-Projective Head Automata (Fred Jelinek Best Paper Award) by Terry Koo, Alexander M. Rush, Michael Collins, Tommi Jaakkola, and David Sontag
- "Poetic" Statistical Machine Translation: Rhyme and Meter (see also here) by Dmitriy Genzel, Jakob Uszkoreit, and Franz Och
- Efficient Graph-Based Semi-Supervised Learning of Structured Tagging Models by Amarnag Subramanya, Slav Petrov, and Fernando Pereira
- Uptraining for Accurate Deterministic Question Parsing by Slav Petrov, Pi-Chuan Chang, Michael Ringgaard, and Hiyan Alshawi
- Self-training with Products of Latent Variable Grammars by Zhongqiang Huang, Mary Harper, and Slav Petrov
Wednesday, October 13, 2010
Google at USENIX Symposium on Operating Systems Design and Implementation (OSDI ‘10)
The 9th USENIX Symposium on Operating Systems Design and Implementation (OSDI ‘10) was recently held in Vancouver, B.C. This biennial conference is one of the premiere forums for presenting innovative research in distributed systems from both academia and industry, and we were glad to be a part of it.
In addition to sponsoring this conference since 2002, Googlers contributed to the exchange of scientific ideas through authoring or co-authoring 3 published papers, organizing workshops, and serving on the program committee. A short summary of the contributions:
- Large-scale Incremental Processing Using Distributed Transactions and Notifications.
Google replaced its batch-oriented indexing system with an incremental system, Percolator. Rather than running a series of high-latency map-reduces over large batches of documents, we now index individual documents at very low latency. The result is a 50% reduction in search result age; our paper discusses this project and the implications of the result. - Availability in Globally Distributed Storage Systems.
Reliable and efficient storage systems are a key component of cloud-based services. In this paper we characterize the availability properties of cloud storage systems based on extensive monitoring of Google's main storage infrastructure and present statistical models that enable further insight into the impact of multiple design choices, such as data placement and replication strategies. We demonstrate the utility of these models by computing data availability under a variety of replication schemes given the real patterns of failures observed in our fleet. - Onix: A Distributed Control Platform for Large-scale Production Networks.
There has been recent interest in a new networking paradigm called Software-Defined Networking (SDN). The crucial enabler for SDN is distributed control platform that shields developers from the details of the underlying physical infrastructure and allows them to write sophisticated control logic against a high-level API. Onix provides such a control platform for large-scale production networks.
In addition to the papers presented by current Googlers, we were also happy to see that the recipient of the 2009 Google Ph.D. Fellowship in Cloud Computing, Roxana Geambasu, presented her work on Comet: An active distributed key-value store.
Videos of all of the talks from OSDI are available on the conference website for attendees and current USENIX members. There is also a USENIX YouTube channel with a growing subset of the conference videos open to everyone.
Google is making substantial progress on many of the grand challenge problems in computer science and artificial intelligence as part of its mission to organize the worlds information and make it useful. Given the continuing increase in the scale of our distributed systems it’s fair to say we’ll have some other exciting new work to share at the next OSDI. Hope to see you in 2012.