Hi, I'm Arjun Reddy Pulugu.

A
Turning Data into Decisions: Boston's Data Science Innovator, Crafting Insights that Drive Success!

About

I am a Data Analytics & AI graduate student at Northeastern University in Boston. I enjoy extracting insights from complex datasets and developing innovative solutions. Always strive to bring 100% to the work I do. I have worked on technologies like Python, SQL, R, PySpark, Tableau, PowerBI, and various machine learning libraries during my studies and professional experience. I have 19 months of work experience as a Systems Data Analyst at Infosys Ltd, which helped me strengthen my skills in data analysis, visualization, and infrastructure optimization. I am passionate about leveraging AI and machine learning techniques to solve real-world problems impacting millions of users and drive business growth.

  • Languages: Python, SQL, R
  • Databases: MySQL, PostgreSQL, MongoDB, ChromaDB, FAISS
  • Libraries: Tidyverse, NumPy, SciPy, Pandas, Matplotlib, Scikit-learn, OpenCV, NLTK
  • Frameworks: PyTorch, Keras, TensorFlow, LangChain, ETL
  • Tools & Technologies: MS Office Suite, Tableau, PowerBI, Git, Docker, AWS, Azure, RAG, Snowflake
  • Certifications: AWS Certified Cloud Practitioner, AWS Certified Solutions Architect - Associate

Seeking a challenging data-focused position that leverages my skills in Data Analysis and Machine Learning. I am eager to contribute to projects that offer professional development, stimulating experiences, and opportunities for personal growth in the field of data science and analytics.

Experience

Graduate Teaching Assistant
  • Provided in-depth guidance on R programming, helping students develop proficiency in data manipulation, statistical analysis, and visualization techniques using real-world datasets.
  • Facilitated weekly discussions and tutoring sessions on fundamental statistical concepts, including probability, measures of central tendency, and variance, ensuring students could apply these principles in their data analysis projects.
  • Offered personalized support during office hours, assisting students in troubleshooting R scripts, interpreting statistical outputs, and applying appropriate analytical methods to solve real-world problems.
  • Evaluated and provided constructive feedback on student assignments, focusing on their ability to use R for exploratory data analysis, implement basic statistical tests, and create compelling data visualizations that support data-driven decision-making.
  • Skills: R, Python, Statistical Analysis, Data visualization, Storytelling
Sep 2024 - Present | Boston, MA
Systems Analyst
  • Extracted time series data from Prometheus using SQL queries for ad-hoc analysis, enhancing data accessibility.
  • Utilized Python and R for statistical analysis and anomaly detection on time series data, identifying critical system issues. This led to a 40% reduction in unplanned downtime and a 25% improvement in mean time to resolution (MTTR).
  • Developed KPIs based on time series analysis, creating a monitoring framework that improved system efficiency by 20% and reduced operational costs by 15%.
  • Designed interactive Tableau dashboards using analyzed time series data, translating complex metrics into actionable insights. This facilitated a 35% increase in cross-team collaboration and supported data-driven decision-making at the executive level.
  • Tools: SQL, Python, R, Tableau
Dec 2021 - July 2023 | Hyderabad, India

Projects

Protein structure modeling
Deep protein structure modeling

Prediction of tertiary structure of a protein from sequences of amino acids.

Accomplishments
  • Standardized datasets like ProteinNet, derived from the CASP competition, are used for training models. Tools like TensorFlow-Keras, along with libraries such as NumPy and Matplotlib, help in building models for predicting protein structures from these datasets.
Explainable-AI-Healthcare
Explainable-AI-Healthcare

Improving accuracy of an explainable AI model’s prediction of patient outcomes by adding stacked generalization

Accomplishments
  • The base learner, an Explainable Boosting Machine (EBM), provides feature importance scores, which are then used to train a meta learner for enhanced prediction. A useful tool for doctors to better understand and trust the model's outcomes.
Intrusion detection system
Intrusion detection system

Exploring the NSL-KDD Dataset: A Comprehensive Analysis About Intrusion Detection System

Accomplishments
  • Built and evaluated intrusion detection models using XGBoost and Logistic Regression, focusing on data cleaning, EDA, preprocessing, and model performance evaluation with metrics like accuracy and feature importance to improve network intrusion detection strategies.
IText2SQL LLM
Text2SQL LLM

Chatbot developed with Langchain and LlamaIndex for natural language to SQL processing, leveraging OpenAI's GPT-3.5 Turbo LLM engine

Accomplishments
  • The project aims to facilitate seamless interaction with IPEDS data, allowing users to generate SQL queries through natural language inputs, thereby enhancing accessibility to educational statistics and insights.
AI Driven Marketing
AI-Driven Marketing

Leveraging AI to enhance customer engagement and reduce churn

Accomplishments
  • Crafting personalized marketing campaigns across various channels (emails, SMS, ads) tailored to customer preferences and behaviors, and segmenting customers (New, Occasional, Regular, Lost) to predict churn and implement targeted retention strategies.
Risk Analysis of tech stocks
Risk Analysis of tech stocks

This analysis provides insights into stock performance, risk-return profiles of tech stocks from 2012 to 2023.

Accomplishments
  • Retrieved data via the yfinance library, calculated daily returns, visualized cumulative returns, computed volatility, performed correlation analysis, and calculated Value at Risk (VaR).
Insurance Fraud
Insurance Fraud

This project aims to predict fraud in insurance policies using various machine learning models.

Accomplishments
  • It focuses on understanding key factors linked to fraud and addressing class imbalance in the target variable. Multiple metrics, such as recall and AUC, are used to evaluate model performance and determine the best approach for predicting fraud effectively.
Online Course Dropouts
Online Course Dropouts

This analysis utilized a LightGBM model to identify the primary indicators of student dropout in online courses.

Accomplishments
  • Key findings revealed that consistent engagement, measured by the "Unique Days" feature, and interaction with video content were significant predictors of student retention, highlighting the importance of fostering a continuous learning experience that keeps students actively involved in their courses.
A/B testing
A/B testing

A/B Testing using Python to determine the most effective marketing campaigns.

Accomplishments
  • The goal is to analyze the impact of each campaign on weekly sales using libraries such as Pandas, Matplotlib, Seaborn, and Scipy to assess which campaign drives the highest sales across the randomly selected outlets.
YOLO Object Detection
YOLO Object Detection

The objective was to accurately identify and localize multiple objects within images

Accomplishments
  • The project involved several key steps, including data collection, annotation, and preprocessing of images to create a robust training dataset. I trained the YOLO model on this dataset, fine-tuning hyperparameters to optimize detection accuracy and speed.
Kafka-Spark
Kafka-Spark

Real time processing and analytics of IoT data using Kafka and Spark

Accomplishments
  • Data from the sensors is streamed to Kafka, processed using PySpark with Spark Streaming, and then analyzed to derive insights such as average temperatures by state, total number of messages processed, and sensor distribution across states.

Skills

Languages and Databases

Python
R
MongoDB
MySQL
PostgreSQL
Shell Scripting

Libraries

NumPy
Pandas
OpenCV
scikit-learn
matplotlib
NLTK

Frameworks

Keras
TensorFlow
PyTorch
LangChain
ETL

Other

Git
AWS
MS Office
Tableau
PowerBI
Docker
Snowflake

Education

Northeastern University

Boston, MA

Degree: Master's in Analytics & AI
CGPA: 3.9/4.0

    Relevant Coursework:

    • Database Management Systems
    • Data Mining
    • Communication and visualization
    • Predictive Analytics
    • Applications of AI
    • Data Management and Big Data

Contact