• Does anyone else spend more time updating spreadsheets than making decisions?

    I work in a typical office role and lately I’ve noticed that a huge part of my day is spent collecting data, updating spreadsheets, checking reports, and reconciling numbers from different sources. The strange thing is that everyone talks about being “data-driven,” but it feels like most of the effort goes into preparing information rather(Read More)

    I work in a typical office role and lately I’ve noticed that a huge part of my day is spent collecting data, updating spreadsheets, checking reports, and reconciling numbers from different sources.

    The strange thing is that everyone talks about being “data-driven,” but it feels like most of the effort goes into preparing information rather than actually using it to make decisions.

    By the time the data is cleaned, verified, shared, and discussed, the opportunity to act on it has often passed.

    I’m curious if others are seeing the same thing in their organizations.

    Are teams genuinely becoming more data-driven, or are we just getting better at creating reports and dashboards?

    What has helped your company move from reporting data to actually using it for faster decisions?

  • Are Data Scientists Becoming AI Supervisors?

    With the rapid adoption of agentic AI, automated feature engineering, AutoML, and AI-assisted analytics, I’m starting to wonder whether the role of a data scientist is changing faster than many expected. Tasks that once required hours of manual work—data cleaning, exploratory analysis, feature selection, model tuning, and even insight generation—can now be partially automated by(Read More)

    With the rapid adoption of agentic AI, automated feature engineering, AutoML, and AI-assisted analytics, I’m starting to wonder whether the role of a data scientist is changing faster than many expected.

    Tasks that once required hours of manual work—data cleaning, exploratory analysis, feature selection, model tuning, and even insight generation—can now be partially automated by AI systems.

    A recent trend highlighted by industry leaders and platforms like Databricks, OpenAI, and Snowflake suggests that data professionals may spend less time building models and more time validating outputs, governing AI systems, and translating results into business decisions.

    Does this mean the future data scientist will look more like an AI supervisor and strategist than a traditional model builder?

    Or do you think deep statistical and machine learning expertise will remain the primary differentiator despite advances in AI tooling?

    Curious to hear how others see the role evolving over the next few years.

  • Has synthetic data become the most important breakthrough in data science?

    As AI models become more data-hungry, many organizations are running into the same problem: obtaining high-quality, diverse, and privacy-compliant data at scale. That’s why synthetic data is gaining so much attention. Instead of relying solely on real-world datasets, teams can generate artificial data that preserves statistical patterns while reducing privacy concerns and addressing data scarcity.(Read More)

    As AI models become more data-hungry, many organizations are running into the same problem: obtaining high-quality, diverse, and privacy-compliant data at scale.

    That’s why synthetic data is gaining so much attention.

    Instead of relying solely on real-world datasets, teams can generate artificial data that preserves statistical patterns while reducing privacy concerns and addressing data scarcity.

    Supporters argue it could unlock innovation in healthcare, finance, autonomous systems, and other industries where data access is limited.

    Critics argue that models trained on synthetic data may inherit biases, amplify errors, or drift away from real-world conditions.

    I’m curious where the community stands:

    Is synthetic data a game-changing breakthrough for data science, or are we overestimating its long-term impact?

    What use cases have you seen where synthetic data genuinely outperformed traditional approaches?

  • Where do you draw the line between feature engineering and in scikit-learn pipeline?

    As machine learning pipelines become more complex, I find myself struggling with a design question rather than a coding one. When you’re working with scikit-learn’s Pipeline and ColumnTransformer, where do you place custom feature creation logic? For example: Creating interaction features Extracting date-based features Combining multiple columns into a new feature Domain-specific transformations Some practitioners(Read More)

    As machine learning pipelines become more complex, I find myself struggling with a design question rather than a coding one.

    When you’re working with scikit-learn’s Pipeline and ColumnTransformer, where do you place custom feature creation logic?

    For example:

    • Creating interaction features
    • Extracting date-based features
    • Combining multiple columns into a new feature
    • Domain-specific transformations

    Some practitioners add these steps before the ColumnTransformer, while others treat feature engineering as part of preprocessing and keep everything inside a single pipeline.

    On one hand, keeping everything in the pipeline improves reproducibility and prevents training-serving skew. On the other hand, deeply nested transformers can become difficult to debug and maintain.

    I’m curious how experienced ML engineers structure their workflows:

    • Do you separate feature engineering from preprocessing?
    • Do you use custom transformers extensively?
    • How do you keep pipelines both reproducible and understandable as projects grow?

    Interested in hearing real-world approaches, especially from teams managing large production ML workflows.

  • Why am I getting SettingWithCopyWarning in Pandas?

    Hi everyone, I’m new to working with Python and Pandas at my job, and I’ve started seeing the SettingWithCopyWarning while cleaning data. The confusing part is that my code still runs, so I’m not sure whether this is something I can ignore or if it’s actually causing problems. I’ve read a few explanations online, but(Read More)

    Hi everyone,

    I’m new to working with Python and Pandas at my job, and I’ve started seeing the SettingWithCopyWarning while cleaning data. The confusing part is that my code still runs, so I’m not sure whether this is something I can ignore or if it’s actually causing problems.

    I’ve read a few explanations online, but I’m still struggling to understand what this warning really means. Is it telling me that I’m modifying a copy instead of the original DataFrame? If so, what’s the recommended way to avoid this warning and make sure my changes are applied correctly?

    I’d appreciate a beginner-friendly explanation, especially if someone can explain why this warning exists and the best practices for handling it in real projects. 

Loading more threads