Recommenter
Authored By: Jared Grimm and Jyan Zárate
CS 525 Information Retrieval, WPI Spring 2020, Professor Kyumin Lee
A Youtube Recommendation platform based on comments and transcripts to improve performance for recommending to niche communities. This project is strictly for educational research only.
Documentation
- Project Proposal
- Proposal Presentation
- Final Presentation
- Git Repository
- Video Comment & Transcript Database
- Jaccard Similarity JSON
- NDCG Evaluation Rubrics
- NDCG Evaluation Results
Research Questions
Can we recommend videos solely off of YouTube comments and transcripts through jaccard similarity?
What are the benefits of solely using comments and transcripts?
How will these recommendations compare with YouTube?
How does this system scale?
Research Hypothesis
People of similar communities watch similar content and react with similar comments
Videos that discuss similar things are highly related to each other
Research Conclusions
We successfully can recommend videos solely based off of comments and transcripts
Being able to recommend videos only based off of comments and transcripts avoids any biases toward the user demographics, location, and items they may not be interested in compared to other users. It also avoids content creators from being able to bias their material to be recommended more by adding inaccurate tags. Instead every YouTube channel gains a fair chance of being recommended only based on their videos content and how people react to it.
Our recommendations compared to user generated ground truth had an NDCG score of 0.978 while YouTube scored 0.970
The system was much harder to scale than we anticipated. Gathering video metadata is plausible, however our limited computer resources and time were not enough to generate the scale we anticipated. However once video metadata is extracted and hashing calculations are performed, the system performs well during run time of fetching related videos, relying on query speed of the MongoDB
Accomplishments
- Developed web crawler to scrape YouTube Comments and Transcripts
- Made a database of roughly 3000 videos Comments, Transcripts, upload date channel Id and Video Ids available for future work
- Indexed over 180,000 scored words based on TF-IDF on transcripts and comments
- Computed minhash with a shingle size of 1, 2, 3, and 4 for all 3000 videos comments and transcripts and added the minhash values to Locality Sensitive Hashing Clusters to boost efficiency in comparing hashes for jaccard similarity.
- Computed Jaccard Similarities on all 3000 videos
- Indexed the calculations mentioned above into a MongoDB instance, so we can query for video and get a JSON output of their Jaccard similarities of Comments and Transcripts
- Weighted comments and transcript similarity the same
- Implemented an automated user test ranking rubric through excel and automatic NDCG calculations comparing our results and YouTubes results to user generated ground truth.
- Generated 12 Rubrics to evaluate NDCG and had 4 outside testers help evaluate them
Challenges
- Gathering data is difficult
- Not all videos had comments allowed or transcripts. Oftentimes movie trailers or music videos disabled subtitles on their videos. This would mean no trailers or music videos can be successfully recommended in our current system.
- Github only allows tracking for files up to 100 mbs, so having to un-track our database and manually share and update over Google Drive was an inconvenience.
- Processing speed for comments, it took roughly 1 minute per 1000 comments on videos to download, with some videos having well over 50,0000 comments, drastically slowing down data mining efforts.
- Indexing unique words to build a dictionary for idf weighting took an extremely long time, so having to reindex if changes were made wasted a lot of time.
- Spelling mistakes in the comments or transcripts caused issues as they were seen as unique words and rated very high by idf ratings, yet were meaningless nonsense. TF weighting and restricting our dictionary to include only words that occurred at least 5 times helped.
- Our LSH hashes did not work as planned as videos did not end up in buckets with similar videos, and also prevented actually relevant videos from being compared. For our current results we excluded the LSH stage of the pipeline before computing jaccard similarities. Omitting such a stage is not computationally prohibitive at a scale of 3,000 videos, but would be a necessary component at a larger scale.
- Transcripts that had no meaning such as [Music] or [Applause] caused videos that weren't related at all to be highly related based on transcript
- YouTubes API has flaws for this use-case
- Overall YouTubes API was very limited in meeting our needs for data at scale. We had to construct our own database using a combination of API calls, other libraries, and custom crawlers.
- YouTube quota limit extremely low. With its 10,000 credit per day limit were only able to process roughly 100 videos a day using YouTubes API alone
- YouTube doesnt specify if related videos returned by "relatedToVideoId" which "retrieves a list of videos that are related to the video" are ordered by relevance at all. We assumed they are ranked from most relevant to least relevant in order to calculate NDCG
Future Work
While we successfully demonstrated the feasibility of using video comments and transcripts to target niche communities, we suggest the following improvements or additions to our work
- Optimize Leverage multi-threading and cloud computing resources to offload the data scraping, term scoring, and hashing functions to be able to gather more data much faster
- Graphic Implementation Utilize MonkeyTamper Google chrome extension scripts to replace YouTubes Recommended videos with our Recommenter generated links.
- Learning to Rank Implement pairwise learning to rank in order to better weight the scores of comment similarity vs transcript similarity, or any other added features. We recommend attempting to score words based on Okapi-BM 25 instead of Tf-IDF to factor in the varying length of transcripts and comments. To do this you must store all original comments and transcripts, and use a python package available through pip install rank-bm25
- Request Increased YouTube API Quota Given more time, one can submit a request to YouTube through Google Cloud services to raise the API quota for their project, this would allow for faster data collection
Dependencies
Please download or clone our github repository and run pip3 install -r requirements.txt
Running Instructions
- To gather YouTube video comments and transcripts
- Everything needed to gather new comments and transcripts are stored in the Video Database Generation Folder. Make that folder your working directory with all dependencies installed.
- You will need a google API key For YouTube API v3 in order to run, please put the key on line 30 of DataScraper.py
- Download the "Video and Comment Database" found on this web-page and insure videoInfo.db is in the working directory locally.
- If you'd rather create an empty database and start fresh, the functions to create a new database and table layouts are available in SQL.py to run from command line. Please read the file for more info
- Run DataScraper.py
- Be sure to change the seed video after each run
- You can change the number of videos to scrape for and the category of videos by changing the Seed and Count variables located on the first two lines of code
- To change the threshold of comments that are gathered by Utuby (all comments will be retrieved) vs YouTube API (only 100-500 comments will be retrieved) on line 15 of fetch_comments.py. higher means longer scraping, lower means more chance of rate limit.
- To compute recommendations
- It is assumed that you have a MongoDB server available to you. By default the main.py script will use the one locally available on 127.0.0.1:27017
- You may download the database export linked above and import the databases to your MongoDB server.
- Given that you have a SQLite 3 database containing scraped videos and comments, you can utilize call the jaccardSimilarity() function of the script in order to run the script which will export a JSON file with recommendations and weights for the comments and transcript components.
- If you wish to add new videos, assuming you have a SQLite 3 database containing scraped videos and comments, you may call the populateDatabase() function of the main.py script in order to first populate a word frequency index, create and sort shingles, create and store a MinHash representation of the video, and then store it in an LSH cluster.
- To evaluate NDCG
- You will once again need to place a YouTube API v3 key on line 14 of createRubric.py
- Run createRubric.py with (url to the starting Youtube video) .
- This will generate an excel file populated with 10 links from YouTube recommendations and 10 links from Recommenter recommendations.
- A user will numerically rank in column b from 1-5 their interpretation of how similar the videos are to the input video.
- Once all videos are ranked and the file is saved locally, run ndcg.py.
- The Ground truth will be the top 10 from the union of both sets
- The Ground truth video Ids will be shown, along side the NDCG Scores for Youtube and Recommenter recommendations.