Data Precog

Projects

Automated Document Categorization Using Machine Learning

  • Developed a machine-learning model for an automated document-categorization system, allowing for documents to be classified into predefined categories based on their content and characteristics. The solution included data preparation, data analysis, feature engineering, model training and evaluation.

  • Causal Analysis and Statistical Modeling of Complex Large-Scale Task-Scheduling Systems

  • Developed a statistical model to identify causal relationships and patterns within a task-scheduling system comprising hundreds of thousands of interdependent, nested tasks.

  • Machine Learning for Detecting Malicious Website Activity and Link Injection

  • Analyzed web activity data, and worked on a project to detect malicious website behavior, including link-injection attacks used for search-engine manipulation. The model analyzed user activity to identify potential SEO spam and support early threat detection.

  • Database Optimization at RTP

  • Analyzed and optimized the database systems supporting satellite communications, programming schedules, and related business-critical operational processes for RTP (Portuguese Television). Through query optimization, indexing improvements, workload analysis, configuration tuning, and infrastructure enhancements, I achieved an approximately 80% improvement in overall database performance and in the systems that used it. These improvements had a direct and measurable impact in the performance of the systems supporting critical satellite-communication and television program scheduling.

  • AWS Data Lake and High-Performance ETL Optimization for Data Science

  • Designed and optimized an AWS-based data lake and ETL environment to support data science and analytics workloads. As part of the initiative, I developed and improved ETL pipelines using AWS Glue, structured and managed data storage in Amazon S3, and optimized querying through Amazon Athena. By converting datasets to columnar Parquet format, implementing adequate partitioning strategies, and selecting only the data required for each analytical workload, I significantly reduced query execution times. These improvements reduced typical AWS Athena query runtimes from approximately seven minutes, down to just one second, substantially improving the efficiency of downstream data science and analytics processes.

  • Multiple Large-Scale Data Analysis for Business Optimization and Performance Improvement

  • Conducted numerous data analysis projects across large-scale database environments to achieve cost-reduction, performance improvements, and support business decision-making such as identifying business oportunities. These projects involved extracting, transforming, and analyzing data from multiple sources; developing metrics and reports; identifying trends and anomalies.