✨作者主页:IT毕设梦工厂✨
个人简介:曾从事计算机专业培训教学,擅长Java、Python、PHP、.NET、Node.js、GO、微信小程序、安卓Android等项目实战。接项目定制开发、代码讲解、答辩教学、文档编写、降重等。
☑文末获取源码☑
精彩专栏推荐⬇⬇⬇
Java项目
Python项目
安卓项目
微信小程序项目
文章目录
- 一、前言
- 二、开发环境
- 三、系统界面展示
- 四、部分代码设计
- 五、论文参考
- 六、系统视频
- 结语
一、前言
《基于大数据的合成航空公司乘客和航班数据分析与可视化》这个系统主要围绕合成航空公司乘客信息与航班运行数据做分析和可视化。数据侧用 Hadoop 做存储、Spark 做计算,通过 Spark SQL、Pandas、NumPy 对航班、乘客、航司、航线、预订、服务体验、延误等数据进行清洗、统计和指标计算,结果落到 MySQL 供后端查询。后端提供 Python+Django 和 Java+Spring Boot 两个版本,前端用 Vue+ElementUI+Echarts 展示图表和大屏。功能包括系统首页、大屏可视化、用户管理、合成航司信息、航司运营分析、航线结构分析、客群画像分析、预订行为分析、服务体验分析、延误风险分析、出行模式分析、个人信息和修改密码。用户可以从不同维度查看航班量、准点率、延误分布、客群特征、预订偏好、服务评价和出行规律,适合作为计算机专业毕设中大数据方向的可视化分析项目。
二、开发环境
大数据框架:Hadoop+Spark(本次没用Hive,支持定制)
开发语言:Python+Java(两个版本都支持)
后端框架:Django+Spring Boot(Spring+SpringMVC+Mybatis)(两个版本都支持)
前端:Vue+ElementUI+Echarts+HTML+CSS+JavaScript+jQuery
详细技术点:Hadoop、HDFS、Spark、Spark SQL、Pandas、NumPy
数据库:MySQL
三、系统界面展示
- 基于大数据的合成航空公司乘客和航班数据分析与可视化系统界面展示:
四、部分代码设计
- 项目实战-代码参考:
spark=SparkSession.builder.appName("SyntheticAirlinePassengerFlightAnalysis").master("local[*]").config("spark.sql.shuffle.partitions","4").getOrCreate()defanalyze_airline_operation():flight_df=spark.read.format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","flight").option("user","root").option("password","123456").load()airline_df=spark.read.format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","airline").option("user","root").option("password","123456").load()passenger_df=spark.read.format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","passenger").option("user","root").option("password","123456").load()flight_df.createOrReplaceTempView("flight")airline_df.createOrReplaceTempView("airline")passenger_df.createOrReplaceTempView("passenger")airline_stat=spark.sql("select a.airline_id,a.airline_name,count(f.flight_id) as flight_count,sum(f.passenger_count) as passenger_total,avg(f.delay_minutes) as avg_delay,sum(case when f.delay_minutes<=15 then 1 else 0 end)/count(f.flight_id) as on_time_rate from flight f join airline a on f.airline_id=a.airline_id group by a.airline_id,a.airline_name")route_stat=spark.sql("select airline_id,departure_city,arrival_city,count(flight_id) as route_flight_count,avg(delay_minutes) as route_avg_delay from flight group by airline_id,departure_city,arrival_city")result=airline_stat.join(route_stat,airline_stat.airline_id==route_stat.airline_id,"left").drop(route_stat.airline_id)result=result.withColumn("on_time_rate",result["on_time_rate"]*100)result=result.orderBy(result["flight_count"].desc(),result["passenger_total"].desc())result.show(50,False)result.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","airline_operation_result").option("user","root").option("password","123456").save()returnresult.toPandas().to_dict("records")defanalyze_delay_risk():flight_df=spark.read.format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","flight").option("user","root").option("password","123456").load()flight_df.createOrReplaceTempView("flight")delay_base=spark.sql("select airline_id,departure_city,arrival_city,flight_date,departure_time,delay_minutes,case when delay_minutes>=60 then '高风险' when delay_minutes>=30 then '中风险' else '低风险' end as risk_level from flight")delay_base.createOrReplaceTempView("delay_base")risk_by_route=spark.sql("select departure_city,arrival_city,count(*) as flight_count,sum(case when risk_level='高风险' then 1 else 0 end) as high_risk_count,avg(delay_minutes) as avg_delay from delay_base group by departure_city,arrival_city")risk_by_airline=spark.sql("select airline_id,count(*) as flight_count,sum(case when risk_level='高风险' then 1 else 0 end) as high_risk_count,avg(delay_minutes) as avg_delay from delay_base group by airline_id")risk_by_hour=spark.sql("select hour(departure_time) as depart_hour,count(*) as flight_count,sum(case when risk_level='高风险' then 1 else 0 end) as high_risk_count,avg(delay_minutes) as avg_delay from delay_base group by hour(departure_time)")route_result=risk_by_route.withColumn("high_risk_rate",risk_by_route["high_risk_count"]/risk_by_route["flight_count"]*100)airline_result=risk_by_airline.withColumn("high_risk_rate",risk_by_airline["high_risk_count"]/risk_by_airline["flight_count"]*100)hour_result=risk_by_hour.withColumn("high_risk_rate",risk_by_hour["high_risk_count"]/risk_by_hour["flight_count"]*100)route_result=route_result.orderBy(route_result["high_risk_rate"].desc())airline_result=airline_result.orderBy(airline_result["high_risk_rate"].desc())hour_result=hour_result.orderBy(hour_result["high_risk_rate"].desc())route_result.show(30,False)airline_result.show(30,False)hour_result.show(24,False)route_result.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","delay_risk_route_result").option("user","root").option("password","123456").save()airline_result.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","delay_risk_airline_result").option("user","root").option("password","123456").save()hour_result.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","delay_risk_hour_result").option("user","root").option("password","123456").save()return{"route":route_result.toPandas().to_dict("records"),"airline":airline_result.toPandas().to_dict("records"),"hour":hour_result.toPandas().to_dict("records")}defanalyze_passenger_profile():passenger_df=spark.read.format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","passenger").option("user","root").option("password","123456").load()booking_df=spark.read.format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","booking").option("user","root").option("password","123456").load()passenger_df.createOrReplaceTempView("passenger")booking_df.createOrReplaceTempView("booking")profile=spark.sql("select p.passenger_id,p.gender,p.age,p.city,p.member_level,count(b.booking_id) as booking_count,sum(b.ticket_price) as total_amount,avg(b.ticket_price) as avg_ticket_price from passenger p left join booking b on p.passenger_id=b.passenger_id group by p.passenger_id,p.gender,p.age,p.city,p.member_level")profile.createOrReplaceTempView("profile")gender_stat=spark.sql("select gender,count(*) as passenger_count,avg(booking_count) as avg_booking_count,avg(total_amount) as avg_total_amount from profile group by gender")age_stat=spark.sql("select case when age<25 then '青年' when age<40 then '中青年' when age<55 then '中年' else '中老年' end as age_group,count(*) as passenger_count,avg(booking_count) as avg_booking_count,avg(total_amount) as avg_total_amount from profile group by case when age<25 then '青年' when age<40 then '中青年' when age<55 then '中年' else '中老年' end")city_stat=spark.sql("select city,count(*) as passenger_count,avg(booking_count) as avg_booking_count,avg(total_amount) as avg_total_amount from profile group by city order by passenger_count desc")member_stat=spark.sql("select member_level,count(*) as passenger_count,avg(booking_count) as avg_booking_count,avg(total_amount) as avg_total_amount from profile group by member_level")gender_stat.show(10,False)age_stat.show(10,False)city_stat.show(20,False)member_stat.show(10,False)gender_stat.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","passenger_gender_profile").option("user","root").option("password","123456").save()age_stat.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","passenger_age_profile").option("user","root").option("password","123456").save()city_stat.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","passenger_city_profile").option("user","root").option("password","123456").save()member_stat.write.mode("overwrite").format("jdbc").option("url","jdbc:mysql://localhost:3306/airline_bigdata").option("dbtable","passenger_member_profile").option("user","root").option("password","123456").save()return{"gender":gender_stat.toPandas().to_dict("records"),"age":age_stat.toPandas().to_dict("records"),"city":city_stat.toPandas().to_dict("records"),"member":member_stat.toPandas().to_dict("records")}五、论文参考
- 计算机毕业设计选题推荐-基于大数据的合成航空公司乘客和航班数据分析与可视化系统-论文参考:
六、系统视频
- 基于大数据的合成航空公司乘客和航班数据分析与可视化系统-项目视频:
项目演示视频
结语
计算机毕业设计选题推荐:基于大数据的合成航空公司乘客和航班数据分析与可视化|毕业设计选题|计算机毕设|选题推荐|毕设指导|项目定制|源码|高质量项目
大家可以帮忙点赞、收藏、关注、评论啦~
源码获取:⬇⬇⬇
精彩专栏推荐⬇⬇⬇
Java项目
Python项目
安卓项目
微信小程序项目