Analysis of a very large web search engine query log
Silverstein C, Marais H, Henzinger M, Moricz M. 1999. Analysis of a very large web search engine query log. ACM SIGIR Forum. 33(1), 6–12.
Download (ext.)
          
        
            
            
            Journal Article
            
            
            
            | Published
            
            
              |              English
              
            
          
        Scopus indexed
Author
        
      Silverstein, Craig;
      Marais, Hannes;
      Henzinger, MonikaISTA  ;
      Moricz, Michael
;
      Moricz, Michael
 ;
      Moricz, Michael
;
      Moricz, MichaelAbstract
    In this paper we present an analysis of an AltaVista Search Engine query log consisting of approximately 1 billion entries for search requests over a period of six weeks. This represents almost 285 million user sessions, each an attempt to fill a single information need. We present an analysis of individual queries, query duplication, and query sessions. We also present results of a correlation analysis of the log entries, studying the interaction of terms within queries. Our data supports the conjecture that web users differ significantly from the user assumed in the standard information retrieval literature. Specifically, we show that web users type in short queries, mostly look at the first 10 results only, and seldom modify the query. This suggests that traditional information retrieval techniques may not work well for answering web search requests. The correlation analysis showed that the most highly correlated items are constituents of phrases. This result indicates it may be useful for search engines to consider search terms as parts of phrases even if the user did not explicitly specify them as such.
    
  Publishing Year
    
  Date Published
    1999-01-01
  Journal Title
    ACM SIGIR Forum
  Publisher
    Association for Computing Machinery
  Volume
      33
    Issue
      1
    Page
      6-12
    ISSN
    
  IST-REx-ID
    
  Cite this
Silverstein C, Marais H, Henzinger M, Moricz M. Analysis of a very large web search engine query log. ACM SIGIR Forum. 1999;33(1):6-12. doi:10.1145/331403.331405
    Silverstein, C., Marais, H., Henzinger, M., & Moricz, M. (1999). Analysis of a very large web search engine query log. ACM SIGIR Forum. Association for Computing Machinery. https://doi.org/10.1145/331403.331405
    Silverstein, Craig, Hannes Marais, Monika Henzinger, and Michael Moricz. “Analysis of a Very Large Web Search Engine Query Log.” ACM SIGIR Forum. Association for Computing Machinery, 1999. https://doi.org/10.1145/331403.331405.
    C. Silverstein, H. Marais, M. Henzinger, and M. Moricz, “Analysis of a very large web search engine query log,” ACM SIGIR Forum, vol. 33, no. 1. Association for Computing Machinery, pp. 6–12, 1999.
    Silverstein C, Marais H, Henzinger M, Moricz M. 1999. Analysis of a very large web search engine query log. ACM SIGIR Forum. 33(1), 6–12.
    Silverstein, Craig, et al. “Analysis of a Very Large Web Search Engine Query Log.” ACM SIGIR Forum, vol. 33, no. 1, Association for Computing Machinery, 1999, pp. 6–12, doi:10.1145/331403.331405.
  
      All files available under the following license(s):
      
      
        
          
        
          
          
      
      
    
  
            Copyright Statement:
          
        
            This Item is protected by copyright and/or related rights. [...]
          
        
      Link(s) to Main File(s)
    
  Access Level
     Open Access
 Open Access
    
 Google Scholar
Google Scholar