SlideShare a Scribd company logo
1 of 4
Download to read offline
DOI: 10.4018/IJITWE.2018040102
International Journal of Information Technology and Web Engineering
Volume 13 โ€ข Issue 2 โ€ข April-June 2018
Copyright ยฉ 2018, IGI Global. Copying or distributing in print or electronic forms without written permission of IGI Global is prohibited.
11
A MapReduce-Based User Identification
Algorithm in Web Usage Mining
Mitali Srivastava, Department of Computer Science, Institute of Science, Banaras Hindu University, Varanasi, India
Rakhi Garg, Computer Science Section, Mahila Maha Vidyalaya, Banaras Hindu University, Varanasi, India
P.K. Mishra, Department of Computer Science, Institute of Science, Banaras Hindu University, Varanasi, India
ABSTRACT
This article contends that in the booming era of information, analysing usersโ€™ navigation behaviour
is an important task. User identification is considered as one of the important and challenging tasks
in the data preprocessing phase of the Web usage mining process. There are three important issues
with the reactive strategies of User identification methods that need to be focused: the first is dealing
of sharing IP address problem in a proxy server environment, the second is distinguishing users
from Web robots, and the third is dealing with huge datasets efficiently. In this article, authors have
developed a MapReduce-based User identification algorithm that deals with the above mentioned
three issues related to user identification methods. Moreover, the experiment on the real web server
log shows the effectiveness and efficiency of the developed algorithm.
KEyWoRdS
Data Cleaning, Data Preprocessing, Hadoop, MapReduce, User Identification,Web Server Log,Web Usage Mining
1. INTRodUCTIoN
Apart from the content and structural information of the Website, server logs have also been
considered as one of the valuable sources of information. This information can be used to analyse
usersโ€™ navigation behaviour (Pabarskaite & Raudys, 2007). Web usage mining is a class of Web
mining to mine server logs to find relevant patterns. These patterns are successfully applied
in various applications like restructuring Websites, recommendation of pages and products,
personalizing Web contents, and improving server activities like prefetching and caching (Facca
& Lanzi, 2005; Kemmar, Lebbah, & Loudni, 2016). Web usage mining process can be divided into
three important steps: Data preprocessing, Pattern extraction and Pattern evaluation (Liu, 2007). Due
to the unstructured and huge nature of log data, Data preprocessing step has become the essential
and time-consuming task in the Web usage mining process. It is a complex task and consumes more
than 60% of whole Web usage mining process time (Tanasa & Trousse, 2004). Data preprocessing
of server log incorporates several steps: Data fusion, Data cleaning, User identification, Session
identification, Path completion, and Data transformation (Cooley, Mobasher, & Srivastava, 1999;
Liu, 2007). Among them, User identification is one of the challenging tasks in Data Preprocessing
International Journal of Information Technology and Web Engineering
Volume 13 โ€ข Issue 2 โ€ข April-June 2018
12
due to the external/local proxy server, shared internet and cache systems (Pabarskaite & Raudys,
2007). This article focuses on User identification, a complex and challenging phase in the Web
usage mining process. In User identification phase, users are identified and their activities are
grouped and recorded into a user activity file. Several heuristics have been proposed for better
identification of the user in last few years. Spiliopoulou et al. have classified user identification
methods into two classes namely proactive methods and reactive methods. In proactive methods,
users are identified by the previous or current interaction of the user with the Website. Proactive
strategies incorporate methods such as user authentication, activation of cookies on the client- side,
dynamic pages associated with the browser, etc. (Spiliopoulou, Mobasher, Berendt, & Nakagawa,
2003). However, these proactive approaches are most accurate and reliable methods for identifying
users but they raise privacy concerns and purely dependent on usersโ€™ cooperation. In the absence
of user authentication approach, the most popular proactive approach to distinguishing unique
user is the use of client-side cookies information (Liu, 2007). Whenever a Web user navigates
through a Website for the first time, the Web server sends a cookie i.e. a piece of information
to the client browser. This information is stored on the client machine in the form of a text file
(Facca & Lanzi, 2005). A cookie may contain various information including usersโ€™ unique id. Few
researchers have applied the cookie based approach to identify users (Elo-Dean & Viveros, 1997;
Ivancsy & Juhasz, 2007; Kamdar & Joshi, 2000). Although this approach is considered as one of
the most accurate methods to identify users but cookies are not often recorded on client machine
due to browser constraints or usersโ€™ non-cooperation e.g. Some browsers do not support cookies
or disable cookies. Sometimes cookies are deleted by the user. On the other hand, in reactive
methods, users are identified from existing log records after interaction with the Website. One
of the basic approaches in reactive methods is identification by the IP address (Gรฉry & Haddad,
2003). However, this approach is unable to deal with sharing IP address issue in the proxy server.
According to Cooley et al., two heuristics can be used to solve this issue: the first heuristic assumes
that two log entries having same IP addresses but different User agents may belong to two different
users. In the second heuristic, some additional information like Web site topology and referrer
log are used to identify users. This heuristic assumes that a user is considered as a new user if
requested page is not accessible through hyperlink of previously requested pages of the same IP
address (Cooley et al., 1999). Tanasa et al. have used IP address and User agent information to
identify users if authentication of the user is not available (Tanasa & Trousse, 2004). Castellano et
al. and Suneetha et al. also, have used IP address and user agent information to identifying users
(Castellano, Fanelli, & Torsello, 2007; Suneetha & Krishnamoorthi, 2009). Further, researchers
have applied the combined approach to identify users. According to their approach, if IP address is
same and User agent is different then consider a new user. Further, if both are same and requested
resource is not accessible through previously accessed pages then consider a new user (Reddy,
Reddy, & Sitaramulu, 2013).
However, all above-discussed methods are successfully applied in various applications but they
are not suitable for large datasets. In the last few years, MapReduce programming framework has
become a popular framework for distributed computation of big data that is executed on a cluster of
nodes and Hadoop is an open source implementation of MapReduce framework (Bhandarkar, 2010;
Dean & Ghemawat, 2008). Few researchers have focused on scalability issues of Data Preprocessing
methods in the Web usage mining process. They have identified Web users by using IP address
information in MapReduce framework (Savitha & Vijaya, 2014; Zhang & Zhang, 2013). However,
their methods are appropriate for large datasets but are unable to deal with proxy server problem.
Huang et al. have given an improved referrer based algorithm for user session identification using
MapReduce programming framework. For user identification, they have considered a specific user
is under same Asymmetric Digital Subscriber Line (ADSL) and same User agent (Huang, Chen, &
Le, 2013). This method is suitable for large datasets however it is not able to distinguish users from
Web robots at User identification phase.
11 more pages are available in the full version of this
document, which may be purchased using the "Add to Cart"
button on the product's webpage:
www.igi-global.com/article/a-mapreduce-based-user-
identification-algorithm-in-web-usage-
mining/198355?camid=4v1
This title is available in InfoSci-Digital Marketing, E-Business,
and E-Services eJournal Collection, InfoSci-Networking,
Mobile Applications, and Web Technologies eJournal
Collection, InfoSci-Journals, InfoSci-Journal Disciplines
Computer Science, Security, and Information Technology,
InfoSci-Journal Disciplines Engineering, Natural, and
Physical Science, InfoSci-Select. Recommend this product to
your librarian:
www.igi-global.com/e-resources/library-
recommendation/?id=162
Related Content
A Constraint Programming Approach for Web Log Mining
Amina Kemmar, Yahia Lebbah and Samir Loudni (2016). International Journal of
Information Technology and Web Engineering (pp. 24-42).
www.igi-global.com/article/a-constraint-programming-approach-for-web-log-
mining/165524?camid=4v1a
What is the Best Technique?
Emilia Mendes (2008). Cost Estimation Techniques for Web Projects (pp. 240-274).
www.igi-global.com/chapter/best-technique/7167?camid=4v1a
Enhancing Interface Understandability as a Means for Better Discovery of
Web Services
Usama Mahmoud Maabed, Ahmed El-Fatatry and Adel El-Zoghabi (2016).
International Journal of Information Technology and Web Engineering (pp. 1-23).
www.igi-global.com/article/enhancing-interface-understandability-as-a-
means-for-better-discovery-of-web-services/165523?camid=4v1a
Ontology-Supported Web Content Management
Geun-Sik Jo and Jason J. Jung (2005). Web Engineering: Principles and Techniques
(pp. 203-223).
www.igi-global.com/chapter/ontology-supported-web-content-
management/31114?camid=4v1a

More Related Content

Similar to A MapReduce-Based User Identification Algorithm in Web Usage Mining.pdf

Classification of User & Pattern discovery in WUM: A Survey
Classification of User & Pattern discovery in WUM: A SurveyClassification of User & Pattern discovery in WUM: A Survey
Classification of User & Pattern discovery in WUM: A SurveyIRJET Journal
ย 
Mining in Ontology with Multi Agent System in Semantic Web : A Novel Approach
Mining in Ontology with Multi Agent System in Semantic Web : A Novel ApproachMining in Ontology with Multi Agent System in Semantic Web : A Novel Approach
Mining in Ontology with Multi Agent System in Semantic Web : A Novel Approachijma
ย 
Web-Application Framework for E-Business Solution
Web-Application Framework for E-Business SolutionWeb-Application Framework for E-Business Solution
Web-Application Framework for E-Business SolutionIRJET Journal
ย 
AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...
AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...
AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...ijdkp
ย 
IRJET-A Survey on Web Personalization of Web Usage Mining
IRJET-A Survey on Web Personalization of Web Usage MiningIRJET-A Survey on Web Personalization of Web Usage Mining
IRJET-A Survey on Web Personalization of Web Usage MiningIRJET Journal
ย 
Web personalization using clustering of web usage data
Web personalization using clustering of web usage dataWeb personalization using clustering of web usage data
Web personalization using clustering of web usage dataijfcstjournal
ย 
A Survey of Issues and Techniques of Web Usage Mining
A Survey of Issues and Techniques of Web Usage MiningA Survey of Issues and Techniques of Web Usage Mining
A Survey of Issues and Techniques of Web Usage MiningIRJET Journal
ย 
Literature Survey on Web Mining
Literature Survey on Web MiningLiterature Survey on Web Mining
Literature Survey on Web MiningIOSR Journals
ย 
Pf3426712675
Pf3426712675Pf3426712675
Pf3426712675IJERA Editor
ย 
Web Data mining-A Research area in Web usage mining
Web Data mining-A Research area in Web usage miningWeb Data mining-A Research area in Web usage mining
Web Data mining-A Research area in Web usage miningIOSR Journals
ย 
Application of fuzzy logic for user
Application of fuzzy logic for userApplication of fuzzy logic for user
Application of fuzzy logic for userIJCI JOURNAL
ย 
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...ijcsa
ย 
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...ijcsa
ย 
Projection Multi Scale Hashing Keyword Search in Multidimensional Datasets
Projection Multi Scale Hashing Keyword Search in Multidimensional DatasetsProjection Multi Scale Hashing Keyword Search in Multidimensional Datasets
Projection Multi Scale Hashing Keyword Search in Multidimensional DatasetsIRJET Journal
ย 
IRJET-Model for semantic processing in information retrieval systems
IRJET-Model for semantic processing in information retrieval systemsIRJET-Model for semantic processing in information retrieval systems
IRJET-Model for semantic processing in information retrieval systemsIRJET Journal
ย 
IMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICES
IMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICESIMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICES
IMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICESIAEME Publication
ย 
Performance of Real Time Web Traffic Analysis Using Feed Forward Neural Netw...
Performance of Real Time Web Traffic Analysis Using Feed  Forward Neural Netw...Performance of Real Time Web Traffic Analysis Using Feed  Forward Neural Netw...
Performance of Real Time Web Traffic Analysis Using Feed Forward Neural Netw...IOSR Journals
ย 
An Extensible Web Mining Framework for Real Knowledge
An Extensible Web Mining Framework for Real KnowledgeAn Extensible Web Mining Framework for Real Knowledge
An Extensible Web Mining Framework for Real KnowledgeIJEACS
ย 
AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...
AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...
AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...James Heller
ย 
Integrated Web Recommendation Model with Improved Weighted Association Rule M...
Integrated Web Recommendation Model with Improved Weighted Association Rule M...Integrated Web Recommendation Model with Improved Weighted Association Rule M...
Integrated Web Recommendation Model with Improved Weighted Association Rule M...ijdkp
ย 

Similar to A MapReduce-Based User Identification Algorithm in Web Usage Mining.pdf (20)

Classification of User & Pattern discovery in WUM: A Survey
Classification of User & Pattern discovery in WUM: A SurveyClassification of User & Pattern discovery in WUM: A Survey
Classification of User & Pattern discovery in WUM: A Survey
ย 
Mining in Ontology with Multi Agent System in Semantic Web : A Novel Approach
Mining in Ontology with Multi Agent System in Semantic Web : A Novel ApproachMining in Ontology with Multi Agent System in Semantic Web : A Novel Approach
Mining in Ontology with Multi Agent System in Semantic Web : A Novel Approach
ย 
Web-Application Framework for E-Business Solution
Web-Application Framework for E-Business SolutionWeb-Application Framework for E-Business Solution
Web-Application Framework for E-Business Solution
ย 
AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...
AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...
AN INTELLIGENT OPTIMAL GENETIC MODEL TO INVESTIGATE THE USER USAGE BEHAVIOUR ...
ย 
IRJET-A Survey on Web Personalization of Web Usage Mining
IRJET-A Survey on Web Personalization of Web Usage MiningIRJET-A Survey on Web Personalization of Web Usage Mining
IRJET-A Survey on Web Personalization of Web Usage Mining
ย 
Web personalization using clustering of web usage data
Web personalization using clustering of web usage dataWeb personalization using clustering of web usage data
Web personalization using clustering of web usage data
ย 
A Survey of Issues and Techniques of Web Usage Mining
A Survey of Issues and Techniques of Web Usage MiningA Survey of Issues and Techniques of Web Usage Mining
A Survey of Issues and Techniques of Web Usage Mining
ย 
Literature Survey on Web Mining
Literature Survey on Web MiningLiterature Survey on Web Mining
Literature Survey on Web Mining
ย 
Pf3426712675
Pf3426712675Pf3426712675
Pf3426712675
ย 
Web Data mining-A Research area in Web usage mining
Web Data mining-A Research area in Web usage miningWeb Data mining-A Research area in Web usage mining
Web Data mining-A Research area in Web usage mining
ย 
Application of fuzzy logic for user
Application of fuzzy logic for userApplication of fuzzy logic for user
Application of fuzzy logic for user
ย 
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
ย 
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
MULTIFACTOR NAรVE BAYES CLASSIFICATION FOR THE SLOW LEARNER PREDICTION OVER M...
ย 
Projection Multi Scale Hashing Keyword Search in Multidimensional Datasets
Projection Multi Scale Hashing Keyword Search in Multidimensional DatasetsProjection Multi Scale Hashing Keyword Search in Multidimensional Datasets
Projection Multi Scale Hashing Keyword Search in Multidimensional Datasets
ย 
IRJET-Model for semantic processing in information retrieval systems
IRJET-Model for semantic processing in information retrieval systemsIRJET-Model for semantic processing in information retrieval systems
IRJET-Model for semantic processing in information retrieval systems
ย 
IMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICES
IMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICESIMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICES
IMPLEMENTATION OF SASF CRAWLER BASED ON MINING SERVICES
ย 
Performance of Real Time Web Traffic Analysis Using Feed Forward Neural Netw...
Performance of Real Time Web Traffic Analysis Using Feed  Forward Neural Netw...Performance of Real Time Web Traffic Analysis Using Feed  Forward Neural Netw...
Performance of Real Time Web Traffic Analysis Using Feed Forward Neural Netw...
ย 
An Extensible Web Mining Framework for Real Knowledge
An Extensible Web Mining Framework for Real KnowledgeAn Extensible Web Mining Framework for Real Knowledge
An Extensible Web Mining Framework for Real Knowledge
ย 
AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...
AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...
AN EXTENSIVE LITERATURE SURVEY ON COMPREHENSIVE RESEARCH ACTIVITIES OF WEB US...
ย 
Integrated Web Recommendation Model with Improved Weighted Association Rule M...
Integrated Web Recommendation Model with Improved Weighted Association Rule M...Integrated Web Recommendation Model with Improved Weighted Association Rule M...
Integrated Web Recommendation Model with Improved Weighted Association Rule M...
ย 

More from Tracy Morgan

45 Plantillas Perfectas De Declaracin De Tesis ( Ejemplos
45 Plantillas Perfectas De Declaracin De Tesis ( Ejemplos45 Plantillas Perfectas De Declaracin De Tesis ( Ejemplos
45 Plantillas Perfectas De Declaracin De Tesis ( EjemplosTracy Morgan
ย 
Thesis Statement Tips. What Is A Thesis Statement. 202
Thesis Statement Tips. What Is A Thesis Statement. 202Thesis Statement Tips. What Is A Thesis Statement. 202
Thesis Statement Tips. What Is A Thesis Statement. 202Tracy Morgan
ย 
B Buy Essays Online Usa, Find Someone To Write My E
B Buy Essays Online Usa, Find Someone To Write My EB Buy Essays Online Usa, Find Someone To Write My E
B Buy Essays Online Usa, Find Someone To Write My ETracy Morgan
ย 
Research Proposal Writing Service - RESEARCH PROPOSAL WRITING S
Research Proposal Writing Service - RESEARCH PROPOSAL WRITING SResearch Proposal Writing Service - RESEARCH PROPOSAL WRITING S
Research Proposal Writing Service - RESEARCH PROPOSAL WRITING STracy Morgan
ย 
Emailing Modern Language Association MLA Handboo
Emailing Modern Language Association MLA HandbooEmailing Modern Language Association MLA Handboo
Emailing Modern Language Association MLA HandbooTracy Morgan
ย 
Custom Wrapping Paper Custom Printed Wrapping
Custom Wrapping Paper Custom Printed WrappingCustom Wrapping Paper Custom Printed Wrapping
Custom Wrapping Paper Custom Printed WrappingTracy Morgan
ย 
How To Write Expository Essay Sketsa. Online assignment writing service.
How To Write Expository Essay Sketsa. Online assignment writing service.How To Write Expository Essay Sketsa. Online assignment writing service.
How To Write Expository Essay Sketsa. Online assignment writing service.Tracy Morgan
ย 
1 Custom Essays Writing. Homework Help Sites.
1 Custom Essays Writing. Homework Help Sites.1 Custom Essays Writing. Homework Help Sites.
1 Custom Essays Writing. Homework Help Sites.Tracy Morgan
ย 
How To Write An Advertisement A Guide Fo
How To Write An Advertisement A Guide FoHow To Write An Advertisement A Guide Fo
How To Write An Advertisement A Guide FoTracy Morgan
ย 
How To Write A Summary Essay Of An Article. How T
How To Write A Summary Essay Of An Article. How THow To Write A Summary Essay Of An Article. How T
How To Write A Summary Essay Of An Article. How TTracy Morgan
ย 
Write A Paper For Me - College Homework Help A
Write A Paper For Me - College Homework Help AWrite A Paper For Me - College Homework Help A
Write A Paper For Me - College Homework Help ATracy Morgan
ย 
24 Hilariously Accurate College Memes. Online assignment writing service.
24 Hilariously Accurate College Memes. Online assignment writing service.24 Hilariously Accurate College Memes. Online assignment writing service.
24 Hilariously Accurate College Memes. Online assignment writing service.Tracy Morgan
ย 
Oh, The Places YouLl Go Printable - Simply Kinder
Oh, The Places YouLl Go Printable - Simply KinderOh, The Places YouLl Go Printable - Simply Kinder
Oh, The Places YouLl Go Printable - Simply KinderTracy Morgan
ย 
How Ghostwriting Will Kickstart Your Music Career
How Ghostwriting Will Kickstart Your Music CareerHow Ghostwriting Will Kickstart Your Music Career
How Ghostwriting Will Kickstart Your Music CareerTracy Morgan
ย 
MLA Handbook, 9Th Edition PDF - SoftArchive
MLA Handbook, 9Th Edition PDF - SoftArchiveMLA Handbook, 9Th Edition PDF - SoftArchive
MLA Handbook, 9Th Edition PDF - SoftArchiveTracy Morgan
ย 
How To Improve Your Writing Skills With 10 Simple Tips
How To Improve Your Writing Skills With 10 Simple TipsHow To Improve Your Writing Skills With 10 Simple Tips
How To Improve Your Writing Skills With 10 Simple TipsTracy Morgan
ย 
Discursive Essay. Online assignment writing service.
Discursive Essay. Online assignment writing service.Discursive Essay. Online assignment writing service.
Discursive Essay. Online assignment writing service.Tracy Morgan
ย 
Creative Writing Prompts 01 - TimS Printables
Creative Writing Prompts 01 - TimS PrintablesCreative Writing Prompts 01 - TimS Printables
Creative Writing Prompts 01 - TimS PrintablesTracy Morgan
ย 
Little Mermaid Writing Paper, Ariel Writing Paper Writin
Little Mermaid Writing Paper, Ariel Writing Paper WritinLittle Mermaid Writing Paper, Ariel Writing Paper Writin
Little Mermaid Writing Paper, Ariel Writing Paper WritinTracy Morgan
ย 
How To Use APA Format Apa Format, Apa Format Ex
How To Use APA Format Apa Format, Apa Format ExHow To Use APA Format Apa Format, Apa Format Ex
How To Use APA Format Apa Format, Apa Format ExTracy Morgan
ย 

More from Tracy Morgan (20)

45 Plantillas Perfectas De Declaracin De Tesis ( Ejemplos
45 Plantillas Perfectas De Declaracin De Tesis ( Ejemplos45 Plantillas Perfectas De Declaracin De Tesis ( Ejemplos
45 Plantillas Perfectas De Declaracin De Tesis ( Ejemplos
ย 
Thesis Statement Tips. What Is A Thesis Statement. 202
Thesis Statement Tips. What Is A Thesis Statement. 202Thesis Statement Tips. What Is A Thesis Statement. 202
Thesis Statement Tips. What Is A Thesis Statement. 202
ย 
B Buy Essays Online Usa, Find Someone To Write My E
B Buy Essays Online Usa, Find Someone To Write My EB Buy Essays Online Usa, Find Someone To Write My E
B Buy Essays Online Usa, Find Someone To Write My E
ย 
Research Proposal Writing Service - RESEARCH PROPOSAL WRITING S
Research Proposal Writing Service - RESEARCH PROPOSAL WRITING SResearch Proposal Writing Service - RESEARCH PROPOSAL WRITING S
Research Proposal Writing Service - RESEARCH PROPOSAL WRITING S
ย 
Emailing Modern Language Association MLA Handboo
Emailing Modern Language Association MLA HandbooEmailing Modern Language Association MLA Handboo
Emailing Modern Language Association MLA Handboo
ย 
Custom Wrapping Paper Custom Printed Wrapping
Custom Wrapping Paper Custom Printed WrappingCustom Wrapping Paper Custom Printed Wrapping
Custom Wrapping Paper Custom Printed Wrapping
ย 
How To Write Expository Essay Sketsa. Online assignment writing service.
How To Write Expository Essay Sketsa. Online assignment writing service.How To Write Expository Essay Sketsa. Online assignment writing service.
How To Write Expository Essay Sketsa. Online assignment writing service.
ย 
1 Custom Essays Writing. Homework Help Sites.
1 Custom Essays Writing. Homework Help Sites.1 Custom Essays Writing. Homework Help Sites.
1 Custom Essays Writing. Homework Help Sites.
ย 
How To Write An Advertisement A Guide Fo
How To Write An Advertisement A Guide FoHow To Write An Advertisement A Guide Fo
How To Write An Advertisement A Guide Fo
ย 
How To Write A Summary Essay Of An Article. How T
How To Write A Summary Essay Of An Article. How THow To Write A Summary Essay Of An Article. How T
How To Write A Summary Essay Of An Article. How T
ย 
Write A Paper For Me - College Homework Help A
Write A Paper For Me - College Homework Help AWrite A Paper For Me - College Homework Help A
Write A Paper For Me - College Homework Help A
ย 
24 Hilariously Accurate College Memes. Online assignment writing service.
24 Hilariously Accurate College Memes. Online assignment writing service.24 Hilariously Accurate College Memes. Online assignment writing service.
24 Hilariously Accurate College Memes. Online assignment writing service.
ย 
Oh, The Places YouLl Go Printable - Simply Kinder
Oh, The Places YouLl Go Printable - Simply KinderOh, The Places YouLl Go Printable - Simply Kinder
Oh, The Places YouLl Go Printable - Simply Kinder
ย 
How Ghostwriting Will Kickstart Your Music Career
How Ghostwriting Will Kickstart Your Music CareerHow Ghostwriting Will Kickstart Your Music Career
How Ghostwriting Will Kickstart Your Music Career
ย 
MLA Handbook, 9Th Edition PDF - SoftArchive
MLA Handbook, 9Th Edition PDF - SoftArchiveMLA Handbook, 9Th Edition PDF - SoftArchive
MLA Handbook, 9Th Edition PDF - SoftArchive
ย 
How To Improve Your Writing Skills With 10 Simple Tips
How To Improve Your Writing Skills With 10 Simple TipsHow To Improve Your Writing Skills With 10 Simple Tips
How To Improve Your Writing Skills With 10 Simple Tips
ย 
Discursive Essay. Online assignment writing service.
Discursive Essay. Online assignment writing service.Discursive Essay. Online assignment writing service.
Discursive Essay. Online assignment writing service.
ย 
Creative Writing Prompts 01 - TimS Printables
Creative Writing Prompts 01 - TimS PrintablesCreative Writing Prompts 01 - TimS Printables
Creative Writing Prompts 01 - TimS Printables
ย 
Little Mermaid Writing Paper, Ariel Writing Paper Writin
Little Mermaid Writing Paper, Ariel Writing Paper WritinLittle Mermaid Writing Paper, Ariel Writing Paper Writin
Little Mermaid Writing Paper, Ariel Writing Paper Writin
ย 
How To Use APA Format Apa Format, Apa Format Ex
How To Use APA Format Apa Format, Apa Format ExHow To Use APA Format Apa Format, Apa Format Ex
How To Use APA Format Apa Format, Apa Format Ex
ย 

Recently uploaded

How to Create and Manage Wizard in Odoo 17
How to Create and Manage Wizard in Odoo 17How to Create and Manage Wizard in Odoo 17
How to Create and Manage Wizard in Odoo 17Celine George
ย 
Holdier Curriculum Vitae (April 2024).pdf
Holdier Curriculum Vitae (April 2024).pdfHoldier Curriculum Vitae (April 2024).pdf
Holdier Curriculum Vitae (April 2024).pdfagholdier
ย 
Magic bus Group work1and 2 (Team 3).pptx
Magic bus Group work1and 2 (Team 3).pptxMagic bus Group work1and 2 (Team 3).pptx
Magic bus Group work1and 2 (Team 3).pptxdhanalakshmis0310
ย 
Russian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in Delhi
Russian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in DelhiRussian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in Delhi
Russian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in Delhikauryashika82
ย 
Activity 01 - Artificial Culture (1).pdf
Activity 01 - Artificial Culture (1).pdfActivity 01 - Artificial Culture (1).pdf
Activity 01 - Artificial Culture (1).pdfciinovamais
ย 
The basics of sentences session 3pptx.pptx
The basics of sentences session 3pptx.pptxThe basics of sentences session 3pptx.pptx
The basics of sentences session 3pptx.pptxheathfieldcps1
ย 
Grant Readiness 101 TechSoup and Remy Consulting
Grant Readiness 101 TechSoup and Remy ConsultingGrant Readiness 101 TechSoup and Remy Consulting
Grant Readiness 101 TechSoup and Remy ConsultingTechSoup
ย 
UGC NET Paper 1 Mathematical Reasoning & Aptitude.pdf
UGC NET Paper 1 Mathematical Reasoning & Aptitude.pdfUGC NET Paper 1 Mathematical Reasoning & Aptitude.pdf
UGC NET Paper 1 Mathematical Reasoning & Aptitude.pdfNirmal Dwivedi
ย 
Key note speaker Neum_Admir Softic_ENG.pdf
Key note speaker Neum_Admir Softic_ENG.pdfKey note speaker Neum_Admir Softic_ENG.pdf
Key note speaker Neum_Admir Softic_ENG.pdfAdmir Softic
ย 
Unit-IV; Professional Sales Representative (PSR).pptx
Unit-IV; Professional Sales Representative (PSR).pptxUnit-IV; Professional Sales Representative (PSR).pptx
Unit-IV; Professional Sales Representative (PSR).pptxVishalSingh1417
ย 
ComPTIA Overview | Comptia Security+ Book SY0-701
ComPTIA Overview | Comptia Security+ Book SY0-701ComPTIA Overview | Comptia Security+ Book SY0-701
ComPTIA Overview | Comptia Security+ Book SY0-701bronxfugly43
ย 
Unit-V; Pricing (Pharma Marketing Management).pptx
Unit-V; Pricing (Pharma Marketing Management).pptxUnit-V; Pricing (Pharma Marketing Management).pptx
Unit-V; Pricing (Pharma Marketing Management).pptxVishalSingh1417
ย 
On National Teacher Day, meet the 2024-25 Kenan Fellows
On National Teacher Day, meet the 2024-25 Kenan FellowsOn National Teacher Day, meet the 2024-25 Kenan Fellows
On National Teacher Day, meet the 2024-25 Kenan FellowsMebane Rash
ย 
Understanding Accommodations and Modifications
Understanding  Accommodations and ModificationsUnderstanding  Accommodations and Modifications
Understanding Accommodations and ModificationsMJDuyan
ย 
Kodo Millet PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...
Kodo Millet  PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...Kodo Millet  PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...
Kodo Millet PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...pradhanghanshyam7136
ย 
This PowerPoint helps students to consider the concept of infinity.
This PowerPoint helps students to consider the concept of infinity.This PowerPoint helps students to consider the concept of infinity.
This PowerPoint helps students to consider the concept of infinity.christianmathematics
ย 
ICT role in 21st century education and it's challenges.
ICT role in 21st century education and it's challenges.ICT role in 21st century education and it's challenges.
ICT role in 21st century education and it's challenges.MaryamAhmad92
ย 
Micro-Scholarship, What it is, How can it help me.pdf
Micro-Scholarship, What it is, How can it help me.pdfMicro-Scholarship, What it is, How can it help me.pdf
Micro-Scholarship, What it is, How can it help me.pdfPoh-Sun Goh
ย 
Python Notes for mca i year students osmania university.docx
Python Notes for mca i year students osmania university.docxPython Notes for mca i year students osmania university.docx
Python Notes for mca i year students osmania university.docxRamakrishna Reddy Bijjam
ย 

Recently uploaded (20)

How to Create and Manage Wizard in Odoo 17
How to Create and Manage Wizard in Odoo 17How to Create and Manage Wizard in Odoo 17
How to Create and Manage Wizard in Odoo 17
ย 
Holdier Curriculum Vitae (April 2024).pdf
Holdier Curriculum Vitae (April 2024).pdfHoldier Curriculum Vitae (April 2024).pdf
Holdier Curriculum Vitae (April 2024).pdf
ย 
Mehran University Newsletter Vol-X, Issue-I, 2024
Mehran University Newsletter Vol-X, Issue-I, 2024Mehran University Newsletter Vol-X, Issue-I, 2024
Mehran University Newsletter Vol-X, Issue-I, 2024
ย 
Magic bus Group work1and 2 (Team 3).pptx
Magic bus Group work1and 2 (Team 3).pptxMagic bus Group work1and 2 (Team 3).pptx
Magic bus Group work1and 2 (Team 3).pptx
ย 
Russian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in Delhi
Russian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in DelhiRussian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in Delhi
Russian Escort Service in Delhi 11k Hotel Foreigner Russian Call Girls in Delhi
ย 
Activity 01 - Artificial Culture (1).pdf
Activity 01 - Artificial Culture (1).pdfActivity 01 - Artificial Culture (1).pdf
Activity 01 - Artificial Culture (1).pdf
ย 
The basics of sentences session 3pptx.pptx
The basics of sentences session 3pptx.pptxThe basics of sentences session 3pptx.pptx
The basics of sentences session 3pptx.pptx
ย 
Grant Readiness 101 TechSoup and Remy Consulting
Grant Readiness 101 TechSoup and Remy ConsultingGrant Readiness 101 TechSoup and Remy Consulting
Grant Readiness 101 TechSoup and Remy Consulting
ย 
UGC NET Paper 1 Mathematical Reasoning & Aptitude.pdf
UGC NET Paper 1 Mathematical Reasoning & Aptitude.pdfUGC NET Paper 1 Mathematical Reasoning & Aptitude.pdf
UGC NET Paper 1 Mathematical Reasoning & Aptitude.pdf
ย 
Key note speaker Neum_Admir Softic_ENG.pdf
Key note speaker Neum_Admir Softic_ENG.pdfKey note speaker Neum_Admir Softic_ENG.pdf
Key note speaker Neum_Admir Softic_ENG.pdf
ย 
Unit-IV; Professional Sales Representative (PSR).pptx
Unit-IV; Professional Sales Representative (PSR).pptxUnit-IV; Professional Sales Representative (PSR).pptx
Unit-IV; Professional Sales Representative (PSR).pptx
ย 
ComPTIA Overview | Comptia Security+ Book SY0-701
ComPTIA Overview | Comptia Security+ Book SY0-701ComPTIA Overview | Comptia Security+ Book SY0-701
ComPTIA Overview | Comptia Security+ Book SY0-701
ย 
Unit-V; Pricing (Pharma Marketing Management).pptx
Unit-V; Pricing (Pharma Marketing Management).pptxUnit-V; Pricing (Pharma Marketing Management).pptx
Unit-V; Pricing (Pharma Marketing Management).pptx
ย 
On National Teacher Day, meet the 2024-25 Kenan Fellows
On National Teacher Day, meet the 2024-25 Kenan FellowsOn National Teacher Day, meet the 2024-25 Kenan Fellows
On National Teacher Day, meet the 2024-25 Kenan Fellows
ย 
Understanding Accommodations and Modifications
Understanding  Accommodations and ModificationsUnderstanding  Accommodations and Modifications
Understanding Accommodations and Modifications
ย 
Kodo Millet PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...
Kodo Millet  PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...Kodo Millet  PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...
Kodo Millet PPT made by Ghanshyam bairwa college of Agriculture kumher bhara...
ย 
This PowerPoint helps students to consider the concept of infinity.
This PowerPoint helps students to consider the concept of infinity.This PowerPoint helps students to consider the concept of infinity.
This PowerPoint helps students to consider the concept of infinity.
ย 
ICT role in 21st century education and it's challenges.
ICT role in 21st century education and it's challenges.ICT role in 21st century education and it's challenges.
ICT role in 21st century education and it's challenges.
ย 
Micro-Scholarship, What it is, How can it help me.pdf
Micro-Scholarship, What it is, How can it help me.pdfMicro-Scholarship, What it is, How can it help me.pdf
Micro-Scholarship, What it is, How can it help me.pdf
ย 
Python Notes for mca i year students osmania university.docx
Python Notes for mca i year students osmania university.docxPython Notes for mca i year students osmania university.docx
Python Notes for mca i year students osmania university.docx
ย 

A MapReduce-Based User Identification Algorithm in Web Usage Mining.pdf

  • 1. DOI: 10.4018/IJITWE.2018040102 International Journal of Information Technology and Web Engineering Volume 13 โ€ข Issue 2 โ€ข April-June 2018 Copyright ยฉ 2018, IGI Global. Copying or distributing in print or electronic forms without written permission of IGI Global is prohibited. 11 A MapReduce-Based User Identification Algorithm in Web Usage Mining Mitali Srivastava, Department of Computer Science, Institute of Science, Banaras Hindu University, Varanasi, India Rakhi Garg, Computer Science Section, Mahila Maha Vidyalaya, Banaras Hindu University, Varanasi, India P.K. Mishra, Department of Computer Science, Institute of Science, Banaras Hindu University, Varanasi, India ABSTRACT This article contends that in the booming era of information, analysing usersโ€™ navigation behaviour is an important task. User identification is considered as one of the important and challenging tasks in the data preprocessing phase of the Web usage mining process. There are three important issues with the reactive strategies of User identification methods that need to be focused: the first is dealing of sharing IP address problem in a proxy server environment, the second is distinguishing users from Web robots, and the third is dealing with huge datasets efficiently. In this article, authors have developed a MapReduce-based User identification algorithm that deals with the above mentioned three issues related to user identification methods. Moreover, the experiment on the real web server log shows the effectiveness and efficiency of the developed algorithm. KEyWoRdS Data Cleaning, Data Preprocessing, Hadoop, MapReduce, User Identification,Web Server Log,Web Usage Mining 1. INTRodUCTIoN Apart from the content and structural information of the Website, server logs have also been considered as one of the valuable sources of information. This information can be used to analyse usersโ€™ navigation behaviour (Pabarskaite & Raudys, 2007). Web usage mining is a class of Web mining to mine server logs to find relevant patterns. These patterns are successfully applied in various applications like restructuring Websites, recommendation of pages and products, personalizing Web contents, and improving server activities like prefetching and caching (Facca & Lanzi, 2005; Kemmar, Lebbah, & Loudni, 2016). Web usage mining process can be divided into three important steps: Data preprocessing, Pattern extraction and Pattern evaluation (Liu, 2007). Due to the unstructured and huge nature of log data, Data preprocessing step has become the essential and time-consuming task in the Web usage mining process. It is a complex task and consumes more than 60% of whole Web usage mining process time (Tanasa & Trousse, 2004). Data preprocessing of server log incorporates several steps: Data fusion, Data cleaning, User identification, Session identification, Path completion, and Data transformation (Cooley, Mobasher, & Srivastava, 1999; Liu, 2007). Among them, User identification is one of the challenging tasks in Data Preprocessing
  • 2. International Journal of Information Technology and Web Engineering Volume 13 โ€ข Issue 2 โ€ข April-June 2018 12 due to the external/local proxy server, shared internet and cache systems (Pabarskaite & Raudys, 2007). This article focuses on User identification, a complex and challenging phase in the Web usage mining process. In User identification phase, users are identified and their activities are grouped and recorded into a user activity file. Several heuristics have been proposed for better identification of the user in last few years. Spiliopoulou et al. have classified user identification methods into two classes namely proactive methods and reactive methods. In proactive methods, users are identified by the previous or current interaction of the user with the Website. Proactive strategies incorporate methods such as user authentication, activation of cookies on the client- side, dynamic pages associated with the browser, etc. (Spiliopoulou, Mobasher, Berendt, & Nakagawa, 2003). However, these proactive approaches are most accurate and reliable methods for identifying users but they raise privacy concerns and purely dependent on usersโ€™ cooperation. In the absence of user authentication approach, the most popular proactive approach to distinguishing unique user is the use of client-side cookies information (Liu, 2007). Whenever a Web user navigates through a Website for the first time, the Web server sends a cookie i.e. a piece of information to the client browser. This information is stored on the client machine in the form of a text file (Facca & Lanzi, 2005). A cookie may contain various information including usersโ€™ unique id. Few researchers have applied the cookie based approach to identify users (Elo-Dean & Viveros, 1997; Ivancsy & Juhasz, 2007; Kamdar & Joshi, 2000). Although this approach is considered as one of the most accurate methods to identify users but cookies are not often recorded on client machine due to browser constraints or usersโ€™ non-cooperation e.g. Some browsers do not support cookies or disable cookies. Sometimes cookies are deleted by the user. On the other hand, in reactive methods, users are identified from existing log records after interaction with the Website. One of the basic approaches in reactive methods is identification by the IP address (Gรฉry & Haddad, 2003). However, this approach is unable to deal with sharing IP address issue in the proxy server. According to Cooley et al., two heuristics can be used to solve this issue: the first heuristic assumes that two log entries having same IP addresses but different User agents may belong to two different users. In the second heuristic, some additional information like Web site topology and referrer log are used to identify users. This heuristic assumes that a user is considered as a new user if requested page is not accessible through hyperlink of previously requested pages of the same IP address (Cooley et al., 1999). Tanasa et al. have used IP address and User agent information to identify users if authentication of the user is not available (Tanasa & Trousse, 2004). Castellano et al. and Suneetha et al. also, have used IP address and user agent information to identifying users (Castellano, Fanelli, & Torsello, 2007; Suneetha & Krishnamoorthi, 2009). Further, researchers have applied the combined approach to identify users. According to their approach, if IP address is same and User agent is different then consider a new user. Further, if both are same and requested resource is not accessible through previously accessed pages then consider a new user (Reddy, Reddy, & Sitaramulu, 2013). However, all above-discussed methods are successfully applied in various applications but they are not suitable for large datasets. In the last few years, MapReduce programming framework has become a popular framework for distributed computation of big data that is executed on a cluster of nodes and Hadoop is an open source implementation of MapReduce framework (Bhandarkar, 2010; Dean & Ghemawat, 2008). Few researchers have focused on scalability issues of Data Preprocessing methods in the Web usage mining process. They have identified Web users by using IP address information in MapReduce framework (Savitha & Vijaya, 2014; Zhang & Zhang, 2013). However, their methods are appropriate for large datasets but are unable to deal with proxy server problem. Huang et al. have given an improved referrer based algorithm for user session identification using MapReduce programming framework. For user identification, they have considered a specific user is under same Asymmetric Digital Subscriber Line (ADSL) and same User agent (Huang, Chen, & Le, 2013). This method is suitable for large datasets however it is not able to distinguish users from Web robots at User identification phase.
  • 3. 11 more pages are available in the full version of this document, which may be purchased using the "Add to Cart" button on the product's webpage: www.igi-global.com/article/a-mapreduce-based-user- identification-algorithm-in-web-usage- mining/198355?camid=4v1 This title is available in InfoSci-Digital Marketing, E-Business, and E-Services eJournal Collection, InfoSci-Networking, Mobile Applications, and Web Technologies eJournal Collection, InfoSci-Journals, InfoSci-Journal Disciplines Computer Science, Security, and Information Technology, InfoSci-Journal Disciplines Engineering, Natural, and Physical Science, InfoSci-Select. Recommend this product to your librarian: www.igi-global.com/e-resources/library- recommendation/?id=162 Related Content A Constraint Programming Approach for Web Log Mining Amina Kemmar, Yahia Lebbah and Samir Loudni (2016). International Journal of Information Technology and Web Engineering (pp. 24-42). www.igi-global.com/article/a-constraint-programming-approach-for-web-log- mining/165524?camid=4v1a What is the Best Technique? Emilia Mendes (2008). Cost Estimation Techniques for Web Projects (pp. 240-274). www.igi-global.com/chapter/best-technique/7167?camid=4v1a
  • 4. Enhancing Interface Understandability as a Means for Better Discovery of Web Services Usama Mahmoud Maabed, Ahmed El-Fatatry and Adel El-Zoghabi (2016). International Journal of Information Technology and Web Engineering (pp. 1-23). www.igi-global.com/article/enhancing-interface-understandability-as-a- means-for-better-discovery-of-web-services/165523?camid=4v1a Ontology-Supported Web Content Management Geun-Sik Jo and Jason J. Jung (2005). Web Engineering: Principles and Techniques (pp. 203-223). www.igi-global.com/chapter/ontology-supported-web-content- management/31114?camid=4v1a