Reddit Analytics Integration Platform Save Abandoned

Project was based on an interest in Data Engineering, ETL pipeline. It also provided a good opportunity to develop skills and experience in a range of tools. As such, project is more complex than required, utilising dbt, airflow, docker and cloud based storage.

Project README

Reddit ETL Pipeline

A data pipeline to extract Reddit data from r/dataengineering.

Output is a Google Data Studio report, providing insight into the Data Engineering official subreddit.

Motivation

Project was based on an interest in Data Engineering and the types of Q&A found on the official subreddit.

It also provided a good opportunity to develop skills and experience in a range of tools. As such, project is more complex than required, utilising dbt, airflow, docker and cloud based storage.

Architecture

Extract data using Reddit API
Load into AWS S3
Copy into AWS Redshift
Transform using dbt
Create PowerBI or Google Data Studio Dashboard
Orchestrate with Airflow in Docker
Create AWS resources with Terraform

Output

Setup

Follow below steps to setup pipeline. I've tried to explain steps where I can. Feel free to make improvements/changes.

NOTE: This was developed using an M1 Macbook Pro. If you're on Windows or Linux, you may need to amend certain components if issues are encountered.

As AWS offer a free tier, this shouldn't cost you anything unless you amend the pipeline to extract large amounts of data, or keep infrastructure running for 2+ months. However, please check AWS free tier limits, as this may change.

First clone the repository into your home directory and follow the steps.

git clone https://github.com/AnMol12499/Reddit-Analytics-Integration-Platform.git

Getting Started

To begin using the project, follow these steps:

More Details

Project Structure: The project's structure includes directories for infrastructure (Terraform), configuration (AWS and Airflow), data extraction (Python scripts), and optional steps like dbt and BI tools integration.

Customization: Feel free to customize the project by modifying configurations, adding new data sources, or integrating additional tools as needed.

Open Source Agenda is not affiliated with "Reddit Analytics Integration Platform" Project. README Source: AnMol12499/Reddit-Analytics-Integration-Platform

Stars

Open Issues

Last Commit

8 months ago

Open Source Agenda Badge

<a href="https://www.opensourceagenda.com/projects/reddit-analytics-integration-platform"><img src="https://www.opensourceagenda.com/projects/reddit-analytics-integration-platform/reviews/badge.svg" alt="Open Source Agenda"></a>

Submit Review Review Your Favorite Project

Submit Resource Articles, Courses, Videos

Submit Article Submit a post to our blog

From the blog

Dec 11, 2022

How to Choose Which Programming Language to Learn First?

From the blog

Dec 11, 2022